An authentication whole-process efficiency improvement method

By building a structured certification knowledge base and intelligent decision-making, the system automatically analyzes evidence materials, formulates certification plans, and generates audit conclusions, solving the problems of low efficiency, high cost, and poor consistency in the traditional certification process, and achieving efficient, accurate, and traceable automation of the certification process.

CN121365986BActive Publication Date: 2026-04-17CHINA ELECTRONICS STANDARDIZATION INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA ELECTRONICS STANDARDIZATION INST
Filing Date
2025-12-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional authentication processes rely on manual operation, resulting in long cycles, low efficiency, high costs, and difficulty in achieving consistency and compliance of authentication results. Existing technologies cannot effectively solve the problems of deep semantic understanding and complex cognitive tasks of unstructured documents.

Method used

A structured certification knowledge base is built, which automates the certification process end-to-end by parsing evidence materials, developing certification plans, associating inspection items, and generating audit conclusions. Optical character recognition, semantic reasoning, and rule engines are used for information extraction and decision-making.

Benefits of technology

It significantly improves certification efficiency, reduces labor costs, ensures consistency and accuracy of audits, actively retrieves evidence fragments through semantic vector indexing, eliminates individual subjective differences, and improves the traceability and quality of certification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365986B_ABST
    Figure CN121365986B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent processing in compliance certification, and discloses a method for improving the efficiency of the entire certification process. The method includes: receiving certification evidence materials; extracting structured information from the certification evidence materials based on a certification knowledge base; formulating a certification plan based on the structured information; associating and matching the certification plan with structured expert experience in the certification knowledge base to determine a set of items to be audited; and determining the compliance status of the certification evidence materials relative to the requirements of the audit items based on the set of items to be audited, and generating an audit conclusion for use in auditing the verification form. By constructing and utilizing a structured certification knowledge base to achieve end-to-end automation of the certification process, the method significantly improves certification efficiency, reduces labor costs, and ensures the consistency and accuracy of the audit process and results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent processing of compliance certification, and in particular to a method for improving the efficiency of the entire certification process. Background Technology

[0002] Traditional certification processes generally rely on a linear workflow driven entirely by humans. From project initiation to certificate issuance, the entire process encompasses at least a dozen core steps, including information collection, planning, on-site audits, checklist preparation, and review, each requiring manual operation and decision-making by certification personnel. This highly human-dependent model results in long certification cycles, slow response times, and labor costs that increase linearly with workload, severely restricting the scalability of services. Particularly when handling unstructured or semi-structured documents such as laws and regulations, industry standards, and enterprise qualification certificates, existing methods rely entirely on manual reading, understanding, and key information extraction, leading to inefficiency and susceptibility to omissions or misjudgments. This constitutes a fundamental bottleneck to achieving full automation of the certification process. Furthermore, because key aspects such as audit plan development, on-site audit focus, and report writing heavily depend on the auditors' personal experience and subjective judgment, inconsistencies in standard application among different auditors can easily arise, directly affecting the consistency of certification results and introducing potential compliance risks. To improve efficiency, the industry has applied robotic process automation (RPA) technology in some stages. However, its essence is the automation of structured tasks that execute predefined rules, and it cannot handle complex cognitive tasks that require deep semantic understanding, contextual reasoning, and experience-based judgment, such as intelligent deconstruction of lengthy regulations and generation of audit strategies based on multi-source information. Meanwhile, some vertical platforms using artificial intelligence to assist compliance management have emerged in the market, but their core positioning is to serve internal compliance self-inspection and audit preparation for enterprises, and they do not provide third-party certification bodies with end-to-end solutions covering the entire certification execution process, capable of dynamic task orchestration and intelligent decision-making. Therefore, the existing technological system struggles to truly achieve deep automation and intelligence throughout the entire certification process while ensuring certification quality and standard consistency; efficiency bottlenecks and cost issues have not been fundamentally resolved. Summary of the Invention

[0003] The purpose of this invention is to provide a method for improving the efficiency of the entire certification process. By constructing and utilizing a structured certification knowledge base, the end-to-end automation of the certification process is achieved, thereby significantly improving certification efficiency, reducing labor costs, and ensuring the consistency and accuracy of the audit process and results.

[0004] To address the aforementioned technical problems, a first aspect of this invention provides a method for improving the efficiency of the entire authentication process, comprising the following steps:

[0005] Receive authentication evidence materials and extract structured information from the authentication evidence materials based on the authentication knowledge base;

[0006] Based on the structured information, an authentication plan is developed;

[0007] The certification plan is matched with the structured expert experience in the certification knowledge base to determine the set of items to be audited;

[0008] Based on the set of inspection items to be audited, the certification evidence materials are assessed for compliance to determine their compliance status with the inspection item requirements, and an audit conclusion is generated for reviewing the verification form.

[0009] Furthermore, the extraction of structured information from the authentication evidence materials based on the authentication knowledge base includes:

[0010] The authentication evidence materials are parsed to obtain their text content;

[0011] Based on the preset information extraction mode in the authentication knowledge base, the text fragments corresponding to the target information field set are identified and located from the text content. The preset information extraction mode includes the target information field set to be extracted required by the authentication process.

[0012] Using the extraction rules in the authentication knowledge base, information extraction is performed on the text content to extract entity information or attribute values ​​from the text fragments and fill them into the corresponding target information fields;

[0013] The filled target information field set is integrated to obtain the structured information of the authentication evidence material.

[0014] Further, the step of parsing the authentication evidence materials to obtain the text content of the authentication evidence materials includes:

[0015] Obtain key document elements related to compliance proof from the certification evidence materials. These key document elements include system architecture diagrams, configuration interface screenshots, operation log records, and compliance declaration forms.

[0016] Text information in the system architecture diagram and configuration interface screenshot is extracted using optical character recognition technology, and the position coordinates of the text information in the original image are recorded.

[0017] Identify the timestamp format and event type of the operation log records, and reorganize the log events in a structured manner according to the time sequence;

[0018] The row and column structure of the compliance declaration form is parsed to establish a mapping relationship between the clause requirements and the corresponding cell content, and the cross-referenced cell content is stored in association.

[0019] The extracted text information, structured log events, and associated table content are integrated according to the evidence logic required for certification audit to generate complete text content for subsequent compliance determination.

[0020] Furthermore, the step of formulating an authentication plan based on the structured information includes:

[0021] Obtain the enterprise attributes contained in the structured information, including industry classification, certification history, and target certification standards;

[0022] The enterprise attributes are matched with the applicable conditions of the pre-stored certification plan templates in the certification knowledge base, and the corresponding target certification plan template is selected.

[0023] Extract enterprise feature data corresponding to the placeholders in the target certification plan template from the structured information, and fill the placeholders with the enterprise feature data;

[0024] Based on the target certification plan template after the enterprise characteristic data is filled in, a complete certification plan is generated, which includes the audit stage sequence, time nodes and audit type definitions.

[0025] Furthermore, the step of matching the enterprise attributes with the applicable conditions of the pre-stored certification plan templates in the certification knowledge base includes:

[0026] Based on the preset matching rules in the certification knowledge base, logical judgments are made on the enterprise attributes to generate preliminary template matching results;

[0027] When the preliminary template matching result has multiple candidate templates or cannot be determined, the enterprise attribute is identified and its features are extracted based on the preset decision prompt words in the authentication knowledge base.

[0028] Based on the results of authentication scenario identification and feature extraction, the priority of the main business and compliance risks in the enterprise attributes is determined;

[0029] Based on the aforementioned main business and compliance risk priorities, the final target certification plan template is selected from the candidate templates.

[0030] Furthermore, the step of associating and matching the certification plan with the structured expert experience in the certification knowledge base to determine the set of items to be audited includes:

[0031] Extract the certification standard identifier from the certification plan;

[0032] Based on the certification standard identifier, all check items corresponding to the certification standard identifier are retrieved from the certification knowledge base, and the check items are a component of the structured expert experience;

[0033] Combine all the aforementioned check items to form an initial check item set;

[0034] Based on the audit scope defined in the certification plan, check items related to the audit scope are selected from the initial check item set to form the check item set to be audited.

[0035] Furthermore, the conformity determination of the certification evidence materials based on the set of items to be audited includes:

[0036] For each inspection item in the set of inspection items to be audited, retrieve the evidence content associated with the requirements of the inspection item from the certification evidence materials;

[0037] The content of the evidence is compared with the textual requirements of the inspection items for consistency.

[0038] Based on the preset evaluation criteria corresponding to the inspection item in the certification knowledge base, the result of the consistency comparison is judged to determine the compliance status.

[0039] Based on the determination of the compliance status of all inspection items in the set of inspection items to be audited, an audit conclusion on the compliance determination is generated.

[0040] Further, retrieving evidence content related to the inspection item requirement from the certified evidence materials includes:

[0041] A semantic vector index is constructed for the authentication evidence materials, and the text content, table content and image OCR results are uniformly represented as high-dimensional vectors;

[0042] The text requirements for the inspection items are also converted into query vectors;

[0043] A similarity search is performed in the semantic vector index to find the multiple evidence fragments most relevant to the query vector;

[0044] The retrieved evidence fragments are sorted based on semantic relevance, and the top K most relevant evidence fragments are returned as the search results.

[0045] Furthermore, before receiving the authentication evidence materials and extracting the structured information of the authentication evidence materials based on the authentication knowledge base, the process further includes:

[0046] The layout analysis and content reconstruction of the certification regulations and standards documents are performed to convert them into standardized documents with hierarchical heading structures.

[0047] Identify cross-reference relationships in the standardized documents and construct a cross-reference knowledge graph containing information on the association of referenced clauses;

[0048] Based on natural language processing technology, the chapter range where the inspection item is located is identified in the standardized document;

[0049] Based on the structural characteristics of the chapter scope, the rules for extracting inspection item types, inspection item content, evaluation content, and evaluation criteria are determined.

[0050] Based on the extraction rules, the structured inspection item types, inspection item contents, evaluation contents, and evaluation criteria in the chapter scope are extracted to obtain the structured expert experience;

[0051] The standardized documents, the cross-referenced knowledge graph, and the structured expert experience are integrated to construct the certification knowledge base.

[0052] Furthermore, based on the extraction rules, the structured inspection item types, inspection item content, evaluation content, and evaluation criteria within the chapter scope are extracted to obtain the structured expert experience, including:

[0053] Based on the inspection item type index defined in the extraction rules, the corresponding target heading level is determined from the hierarchical heading structure of the standardized document, and all subheading path sequences belonging to the target heading level are taken as inspection item types.

[0054] Based on the positioning identifiers defined in the extraction rules, text fragments associated with the positioning identifiers are located and extracted from the text content of the chapter range, respectively, to form structured inspection items, evaluation content and evaluation criteria.

[0055] During the extraction process, if a cross-reference relationship is identified within the processed text fragment, the cross-reference knowledge graph is queried to obtain the full text of the corresponding referenced clause, and the obtained full text content is dynamically embedded into the structured entry currently being constructed.

[0056] The structured expert experience is obtained by structurally encapsulating and associating the inspection item type, the inspection item content with completed context embedding, the evaluation content and the evaluation criteria with the inspection item as the basic unit.

[0057] Accordingly, a second aspect of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described authentication process efficiency improvement method.

[0058] Accordingly, a third aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described method for improving the efficiency of the entire authentication process.

[0059] The above-described technical solutions of the embodiments of the present invention have the following beneficial technical effects:

[0060] 1. By automating the parsing of regulatory and standard documents and building an certification knowledge base containing cross-references, unstructured expert experience that traditionally relied on manual interpretation and memorization is transformed into structured knowledge that can be processed and correlated by machines. This fundamentally solves the bottleneck of unstructured information processing, enabling certification activities to be based on a unified, accurate, and context-complete knowledge foundation, and significantly improving the completeness and standardization of the audit.

[0061] 2. With intelligent agent technology as the core, it automatically formulates plans, associates check items and generates notifications based on structured information, realizing full-process automation from project initiation to report generation; by integrating rule engine and semantic reasoning, it can handle various decisions from clear rules to complex scenarios, freeing human resources from repetitive paperwork and allowing them to focus on high-value professional judgment, thereby significantly shortening the certification cycle and reducing operating costs while ensuring decision quality.

[0062] 3. Based on the constructed semantic vector index and dynamic evidence chain, it can actively retrieve evidence fragments related to the inspection items from multi-source heterogeneous materials, and perform traceable compliance reasoning according to preset evaluation criteria. This mode of transforming the "experience verification" of human auditors into the "evidence verification" of the system not only greatly improves the coverage depth and accuracy of the audit, but also effectively eliminates individual subjective differences through the unified standard execution, ensuring high consistency of certification results. Attached Figure Description

[0063] Figure 1 This is a flowchart of the authentication process efficiency improvement method provided in this embodiment of the invention;

[0064] Figure 2 This is an architecture diagram of the authentication process efficiency improvement method provided in this embodiment of the invention;

[0065] Figure 3 This is a flowchart of the technology for automatically developing a large-scale authentication plan based on a template, provided in an embodiment of the present invention.

[0066] Figure 4 This is a flowchart illustrating the process of parsing regulatory standard PDFs into Markdown documents, as provided in this embodiment of the invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0068] Please refer to Figure 1 and Figure 2 The first aspect of this invention provides a method for improving the efficiency of the entire authentication process, comprising the following steps:

[0069] Step S100: Receive authentication evidence materials and extract structured information from the authentication evidence materials based on the authentication knowledge base.

[0070] We receive various original evidence materials submitted by companies applying for certification. These materials are typically multi-source and heterogeneous, including scanned PDFs of company qualification certificates, schematic diagrams describing system architecture, screenshots of software configuration interfaces, system operation log files, and compliance statement forms declaring compliance with specific clauses. To achieve automated information extraction, we first perform targeted analysis on different types of materials: For scanned or image-formatted materials, we use optical character recognition technology to convert them into text and record the position coordinates of key text in the original image to preserve contextual relationships; for log-type time-series data, we identify and parse its timestamp format and event type, and then reorganize it according to the time sequence to form a structured operation event flow; for tabular materials, we analyze its row and column logic to establish a precise mapping between the clause requirements in the table and the content of specific evidence cells, and we associate and store the content of cells with cross-references. After initial parsing, based on the information extraction patterns and extraction rules pre-set in the certification knowledge base, the target information fields related to the certification process (such as company name, organizational structure, business scope, core system name, implemented control measures, etc.) are identified and located from the integrated text content. Then, specific entity information and attribute values ​​are extracted through named entity recognition or rule matching and filled into the corresponding structured fields. The final output is a machine-readable and semantically clear structured dataset of enterprise information.

[0071] Step S200: Develop an authentication plan based on structured information.

[0072] Based on the structured information extracted in step S100, a complete, compliant, and actionable certification execution plan is automatically generated. First, key enterprise attributes are extracted from the structured information, including its industry classification code, past certification history, and the target certification standard system for this application. Then, these attributes are matched against the applicable conditions of a series of pre-stored certification plan templates in the certification knowledge base. The matching process employs a hierarchical decision-making mechanism: first, a pre-set deterministic rule engine performs rapid filtering; for example, if the enterprise belongs to the automotive manufacturing industry and is applying for IATF 16949 certification for the first time, the corresponding dedicated template is directly matched. When the enterprise's situation is complex, attributes overlap, or rules cannot be uniquely determined (e.g., the enterprise is involved in both medical devices and cloud computing services, requiring simultaneous compliance with ISO 13485 and ISO 27001), a reasoning mechanism based on a large language model is activated. This mechanism guides the model to perform chain-like reasoning based on pre-set decision prompts, analyzing the priority of different business functions and compliance risks, and then selecting the most suitable target template from the candidate templates. After selecting a template, the system further extracts enterprise characteristic data (such as company size, product name, and existing key processes) corresponding to the reserved placeholders in the template from the structured information and automatically fills them into the template, ultimately generating a detailed certification plan. This plan clarifies the sequence arrangement, specific time nodes, and the type and core focus of each audit stage (such as preliminary audit, surveillance audit, and recertification).

[0073] Step S300: The certification plan is matched with the structured expert experience in the certification knowledge base to determine the set of items to be audited.

[0074] The certification plan, generated in the previous step and tailored to a specific project, is precisely correlated with the universally applicable structured expert experience contained in the certification knowledge base to derive the specific set of inspection items that need to be verified in this audit. First, the certification plan is parsed to extract the identifier of the certification standard explicitly referred to (e.g., "ISO / IEC 27001:2022"). Then, using this identifier as the query key, all structured inspection items corresponding to that standard are retrieved from the certification knowledge base. These inspection items are structured expert experience formed offline through intelligent parsing, chapter location, and rule extraction of the original text of the certification standard. Each inspection item typically includes its unique clause number, inspection item content, assessment implementation method, and conformity assessment criteria. All retrieved inspection items are summarized to form the initial complete set of inspection items. Next, based on the audit scope specifically defined in the certification plan (for example, this audit only covers the two control domains of "access control" and "password management"), the initial complete set is filtered to remove check items outside the audit scope. Finally, the "set of check items to be audited" that is actually required for this certification task is accurately located and output, thereby ensuring that the subsequent audit work fully covers the standard requirements while closely focusing on the scope agreed upon in the plan.

[0075] Step S400: Based on the set of inspection items to be audited, conduct a conformity assessment of the certification evidence materials, determine the conformity status of the certification evidence materials relative to the requirements of the inspection items, and generate an audit conclusion for auditing the verification form.

[0076] Based on the evidence materials, a compliance conclusion is made for each item to be reviewed. For each item in the checklist, a deep search is performed on the complete set of authentication evidence materials received and preprocessed in step S100. To achieve accurate retrieval, a unified semantic vector index is pre-constructed for all evidence materials (including text, tables, and image OCR content). The text requirements of the check items are converted into query vectors. Through similarity calculation, the top K most semantically relevant evidence fragments are retrieved from the vector index. These fragments may come from different sources such as program files, screenshot descriptions, interview records, or log entries, and together they constitute a dynamic chain of evidence supporting or refuting the requirements of the check item. Subsequently, the retrieved evidence content is carefully compared with the text requirements of the check items for consistency. And strictly according to the pre-set, clear judgment criteria for the check item in the authentication knowledge base (e.g., "if A exists and B exists, it is judged as compliant; if A is missing or C exists, it is judged as non-compliant"), the results of the consistency comparison are automatically and logically judged to obtain a clear state of "compliant", "non-compliant", or "partially compliant". After processing all pending audit items, the judgment results for each item and the key evidence references cited are summarized. Following the standardized report format, a draft audit checklist with a clear structure and traceable evidence is automatically generated. This draft contains the core audit conclusions of this certification activity.

[0077] Through the continuous automated execution of the above steps, this invention transforms the traditional authentication process, which heavily relies on manual reading, understanding, and judgment, into a standardized process driven by a structured knowledge base and centered on intelligent matching and reasoning. This method automates the parsing and information extraction of unstructured regulations and evidence materials, replacing manual information organization and entry. Intelligent decision-making combining rules and models generates accurate and personalized authentication plans, replacing manual experience-based planning. Precise correlation and filtering of check items ensure the completeness and relevance of the audit tasks. Finally, through automated conformity determination based on semantic retrieval and evidence chains, a systematic verification of massive amounts of evidence materials is completed, generating preliminary audit conclusions. The entire process significantly reduces reliance on repetitive manual labor and individual experience in the authentication process. While improving processing efficiency and shortening the authentication cycle, it significantly enhances the standard consistency of the audit process, the objectivity of conclusions, and the traceability of evidence, thereby improving the overall quality, reliability, and scalability of authentication services.

[0078] Specifically, step S100, which involves extracting structured information from authentication evidence materials based on the authentication knowledge base, includes:

[0079] Step S110: Analyze the authentication evidence materials to obtain the text content of the authentication evidence materials.

[0080] Because the evidence submitted by enterprises is multi-source and heterogeneous, differentiated analysis strategies are required for materials of different formats and types. For image-based evidence such as system architecture diagrams, network topology diagrams, or screenshots of configuration management interfaces, optical character recognition (OCR) technology is first used to extract all the text information contained therein, and the coordinate bounding box of each text block in the original image is recorded to preserve its spatial layout and contextual relationships. For various types of plain text or semi-structured log files such as operation logs, audit logs, or transaction records, their specific timestamp formats, event level tags, and message body structures are identified and parsed. Then, all log entries are reordered and combined according to the time series to form a coherent, time-axis-organized structured event flow. For documents such as compliance declaration forms or self-assessment reports submitted by enterprises, layout analysis technology is used to identify their table framework, cell merging relationships, and table header structure. This establishes a precise mapping relationship between the clause numbers and control requirements declared in the table and the specific descriptions used as supporting evidence. Furthermore, cross-references that may exist in the table pointing to other chapters or external documents are identified and associated with the data. Finally, the text information extracted from various forms such as images, logs, and tables, the structured event sequences, and the associated declarations are integrated and arranged according to the logical evidence chain required for certification and auditing, outputting a unified and complete set of text content that can be processed in the subsequent information extraction process.

[0081] Step S120: Based on the preset information extraction mode in the authentication knowledge base, identify and locate the text fragments corresponding to the target information field set from the text content. The preset information extraction mode contains the target information field set to be extracted required by the authentication process.

[0082] Based on domain knowledge, the scope of information closely related to the certification decision is initially defined from the integrated text content generated in step S110. The certification knowledge base has a pre-defined "information extraction pattern," which is essentially a standardized "information requirements list" or "question set" defined for specific certification scenarios (such as ISO 27001 information security system certification). This pattern clearly lists the set of enterprise information fields that must be obtained to complete the certification process. These fields typically include, but are not limited to: the enterprise's standard industry classification code, the range of total employees, the main types and quantities of information systems, existing management system certification certificates, the specific scope and exclusion statements of this certification application, key business process descriptions, and a core data asset list. This pattern is applied to the integrated text content, and text fragments associated with each target information field description are scanned and located throughout the entire text using rule-based keyword matching, regular expression pattern recognition, or lightweight semantic fragment boundary detection technology. For example, the text searches for sentences or paragraphs containing suggestive language such as "company size approximately," "total number of employees," "main systems include," "certified," and "scope of this certification," and marks these as candidate text fragments that may contain information corresponding to target fields such as "company size," "core systems," "historical certifications," and "certification scope." This step does not perform precise information extraction but rather completes a preliminary "information detection" and "scope positioning," providing a clear target for the next step of precise extraction.

[0083] Step S130: Using the extraction rules in the authentication knowledge base, information extraction is performed on the text content, entity information or attribute values ​​are extracted from the text fragments and filled into the corresponding target information fields.

[0084] Based on the candidate text fragments located in step S120, this step performs refined information extraction and structured assignment. The authentication knowledge base is configured with corresponding "extraction rules" for each target information field. These rules may be text parsing rules based on specific contextual patterns, or they may be instructions to call a pre-trained natural language processing model (such as a named entity recognition model). Each text fragment located in step S120, along with its corresponding target field information, is input into the corresponding extraction rule for processing. For example, for the "company size" field, the rule might be defined as: in sentences containing keywords such as "employees" and "number of people," extract the nearest numerical entity and its unit (e.g., "approximately 300 people"), and normalize it to the pure number "300" to fill in the field. For the "industry classification" field, the rule might contain a mapping table of standard industry classification codes and industry keywords, determining the corresponding standard code (e.g., I6510) by matching keywords in the text fragment (e.g., "software development," "medical device manufacturing"). For more complex, unstructured descriptive fields, such as "core business processes," rules may guide a large language model to summarize and extract key steps and inputs / outputs from lengthy descriptions. Executing these rules precisely extracts specific entity names, numerical values, dates, codes, or general descriptions from text fragments and populates these values ​​into predefined target information fields with clear data types and semantic definitions. This process completes the transformation from unstructured text to discrete, typed data points.

[0085] Step S140: Integrate the filled target information field set to obtain the structured information of the authentication evidence materials.

[0086] The discrete data points generated in the preceding steps are collected, verified, and encapsulated to form a structured data object that can ultimately drive subsequent automated processes. All target information fields filled in step S130 are assembled according to their logical relationships defined in the preset information extraction mode. This typically involves combining related fields into higher-level structures; for example, combining fields such as "Company Name," "Unified Social Credit Code," and "Registered Address" into a "Corporate Entity" information block; and combining fields such as "Number of Servers," "Database System Type," and "Network Boundary Devices" into an "IT Infrastructure Overview" information block. During this process, basic data consistency checks are performed, such as checking whether numeric fields are within reasonable ranges or whether all required fields have obtained valid values. Finally, a complete, machine-readable structured information object is generated. This object typically exists in the form of JSON, XML, or a specific protocol buffer, with a clear hierarchical structure, where each leaf node corresponds to a filled target information field and its value. This structured information object comprehensively and accurately represents the enterprise entity attributes and environmental characteristics that are crucial to the certification process, extracted from the original certification evidence materials. It provides a solid and reliable data input foundation for subsequent automated decision-making processes such as intelligent certification plan formulation and precise correlation of inspection items.

[0087] Further, step S110 involves parsing the authentication evidence materials to obtain their text content, including:

[0088] Step S111: Obtain key document elements related to compliance proof from the certification evidence materials. Key document elements include system architecture diagram, configuration interface screenshots, operation log records, and compliance declaration forms.

[0089] In the context of certification audits, the collection of evidentiary materials is usually comprehensive, while the core evidentiary materials directly used to prove that a system or product meets specific standard requirements are specifically manifested in several typical document elements. For example, in information security certification, system architecture diagrams are used to show network boundaries, security zone divisions, and critical asset deployments, and are the foundation for understanding the control environment; configuration interface screenshots (such as firewall policy configuration pages and operating system security settings interfaces) directly provide visual proof of the current configuration status of the system; operation log records (such as access logs, change audit logs, and error logs) objectively record the actual operating behavior and events of the system in chronological order, and are key dynamic evidence to verify the effective operation of control measures; while compliance declaration forms are the applicant's self-statements and evidence indexes based on standard clauses, usually in a structured or semi-structured form, clearly listing the declared controls, corresponding clauses, and their supporting positions in the evidentiary materials.

[0090] Step S112: Extract text information from the system architecture diagram and configuration interface screenshot using optical character recognition technology, and record the position coordinates of the text information in the original image.

[0091] The visual text information is converted into digital text that can be processed and analyzed, while preserving its spatial layout semantics. For the system architecture diagram, device icon labels (such as "Core Switch FW-01", "Database Server DB-SVR"), network link labels (such as "Gigabit Fiber", "DMZ Area"), and topology annotation text all contain key information for understanding the system's structure and relationships. For configuration interface screenshots, the titles, tab names, checkbox status text, parameter input box values ​​(such as "Password Complexity Requirement: Enabled", "Session Timeout: 15 minutes"), and button labels are direct evidence of whether the system configuration complies with security policies. This step uses optical character recognition (OCR) technology to preprocess the input image, perform text detection, and character recognition. While recognizing the text content, the bounding box position of each recognized text segment in the original image's pixel coordinate system is precisely recorded. For example, the text "Minimum password length is 8" is recognized, and its specific coordinate range in the upper left corner of the screenshot is recorded.

[0092] Step S113: Identify the timestamp format and event type of the operation log records, and restructure the log events in a structured manner according to the time sequence.

[0093] The process involves parsing and reconstructing raw, typically text-based operation logs into a structured data sequence with a clear time dimension and event semantics. Different systems generate logs with varying formats, such as Linux system auth.log, Windows security event logs, and application server custom logs, but they generally contain two core elements: timestamps and event descriptions. First, regular expression matching, pattern recognition, or a pre-trained log parsing model are used to automatically identify and parse the timestamp portion of each log record, converting it to a standard time format (such as ISO 8601). Simultaneously, the event type described by the log entry is identified, such as "user login successful," "file modified," "permission change request," or "system alarm." This is typically achieved by matching keywords or event IDs in the log message or using a classification model. After parsing each log entry, all entries are sorted in ascending order according to the converted timestamps, thus reconstructing what might have been a scattered, disorganized log stream into a coherent, structured event sequence arranged along a single timeline.

[0094] Step S114: Parse the row and column structure of the compliance declaration form, establish the mapping relationship between the clause requirements and the corresponding cell content, and store the cross-referenced cell content together.

[0095] Such tables typically use rows to represent specific control requirements or standard clauses, and columns to represent different attributes (such as clause number, control description, implementation status, evidence document index, remarks, etc.). First, layout analysis techniques are used to identify the physical structure of the table, including determining the header area, data area, and analyzing cell merging and splitting. Based on this, a mapping relationship is established between the "row-column" coordinates of the table data area and the semantics of the header. This allows for understanding the standard clause corresponding to each row of data (e.g., ISO 27001 A.9.2.1), and the attributes represented by the content of each column cell under that row (e.g., "Implemented" under the "Implementation Status" column, "Section 3.2 of the Access Control Procedure" under the "Evidence Index" column). This mapping relationship enables the precise extraction of knowledge pairs such as "Clause X corresponds to Evidence Y". Specifically, the system detects whether there are cross-reference marks in the cell content pointing to other parts of the table, other documents, or external standards (such as "see Section 4.5" or "according to Article 3 of the Specification"). Once such references are detected, the association between the source cell (the referencer) and the target identifier (the referenced) is recorded and stored independently in an association index.

[0096] Step S115: The extracted text information, structured log events, and associated table content are integrated according to the evidence logic required for certification audit to generate complete text content for subsequent compliance determination.

[0097] The scattered and diverse intermediate results obtained from the previous independent sub-steps are integrated into a unified, coherent, and logically ordered comprehensive text content, according to the cognitive logic of certification and auditing and the needs of evidence chain organization. Specifically, this is not simply a matter of piecing together all the identified text, but rather following a specific integration strategy. For example, the "clause-evidence" mapping relationship parsed from the compliance declaration form can be used as a framework and guide: for the evidence "Access Control Policy Document" cited in "Clause A.9.1.1" of the declaration form, the relevant network boundary and access path description text extracted from the system architecture diagram will be located and inserted; simultaneously, event records related to user access authentication (such as login success / failure logs) from the structured log event sequence will be selected in chronological order and embedded in the corresponding positions as dynamic evidence of policy execution. Text extracted from configuration interface screenshots, such as access control list configurations, will be associated with the corresponding control measure descriptions.

[0098] For details, please refer to Figure 3Step S200, which involves developing an authentication plan based on structured information, includes:

[0099] Step S210: Obtain the enterprise attributes contained in the structured information. The enterprise attributes include industry classification, certification history and target certification standards.

[0100] From the structured data object of enterprise information generated in step S100, several key attributes that play a decisive role in the macro framework of the certification plan are precisely located and extracted. These attributes constitute the basic dimensions for subsequent intelligent decision-making. Among them, "industry classification" is usually a standardized code (such as the National Economic Industry Classification GB / T 4754), used to determine the specific regulations and industry practices followed by the enterprise's operations. For example, code C3660 represents "manufacturing of automotive parts and accessories," and its applicable core certification standard (such as IATF 16949) is completely different from that of a software company (code I6510). "Certification history" refers to the management system certifications that the enterprise has obtained in the past that are related to this certification, such as whether it has passed the basic ISO 9001 quality management system certification. This determines whether this certification is a "first certification" or an "upgrade or extension certification based on the existing system," directly affecting the complexity and starting point of the plan. "Target certification standard" is the specific certification system that the enterprise is clearly applying for this time, such as ISO 27001 information security management system or ISO 13485 medical device quality management system. By accessing the corresponding predefined fields in the structured data object, the values ​​of these attributes can be read directly.

[0101] Step S220: Match the enterprise attributes with the applicable conditions of the pre-stored certification plan templates in the certification knowledge base, and select the corresponding target certification plan template.

[0102] The certification knowledge base pre-stores "certification plan templates" designed for different typical scenarios, each template accompanied by a clear description of "applicable conditions." The matching process employs a hierarchical, collaborative intelligent decision-making mechanism. First, a rule engine based on deterministic logic is activated to quickly compare the enterprise attributes obtained in step S210 with the applicable conditions of each template in the template library. The rule engine consists of a series of "IF-THEN" logical statements, handling scenarios with clear conditions and well-defined boundaries. For example, a rule could be: "IF (Industry Classification IN ['Automotive Manufacturing Related Code']) AND (Certification History Does Not Include IATF 16949) AND (Target Certification Standard == 'IATF 16949') THEN Select Template 'T-IATF-Initial Certification-Three-Year Plan'". For such scenarios, the rule engine can quickly and accurately output a uniquely matching target template. However, when a company's attributes are complex, overlapping, or ambiguous (for example, a company's business involves both "smart hardware manufacturing" (C code) and "online health services" (I code), and it is simultaneously applying for dual certifications of "ISO 13485" and "ISO 27001"), simple rules may not provide a definitive conclusion or generate multiple candidate templates. In such cases, a semantic reasoning mechanism based on a large language model is employed. This mechanism guides the model to perform chain-like reasoning through carefully designed "decision prompts": first, it analyzes the core compliance priorities and risk levels of each target certification standard; then, considering the company's industry composition, it determines its main business and the priority of compliance risks; finally, based on these analyses, it evaluates the completeness of each candidate template's coverage of high-risk requirements and its ability to integrate secondary standards, and outputs a recommended template ID and detailed reasoning justification.

[0103] Step S230: Extract enterprise feature data from the structured information that corresponds to the placeholders in the target certification plan template, and fill the placeholders with the enterprise feature data.

[0104] The generic template is "instantiated" into a personalized plan draft for a specific company. The certification plan template is a document blueprint containing a fixed structural framework and variable placeholders. Placeholders mark the locations of content that needs to be filled in based on the specific company's situation; they come in various types, such as {CompanyName}, {EmployeeScale}, {CoreProductDescription}, {InitialAuditYear}, {FirstSurveillanceAuditMonth}, etc. It is necessary to accurately locate and extract the company feature data corresponding to the semantics of each placeholder from the more detailed structured information generated in step S100. This requires establishing a mapping relationship from placeholder labels to structured information fields. Some mappings are direct, such as {CompanyName} mapping to the "Company Name" field. Some calculations or derivations are required. For example, `{FirstSurveillanceAuditMonth}` might need to be calculated based on company size, industry risk level, and standard requirements using pre-defined logical rules (e.g., "For medium-sized manufacturing companies, the first surveillance audit will be conducted in the 12th month after the initial audit"). `{CoreProductDescription}` might need to extract key information from the "Product and Service" description field using text summarization technology. This mapping and extraction process involves filling the obtained numerical values, text, or dates into the corresponding placeholder positions in the template. For example, for a medical software development company with 80 employees, `{EmployeeScale}` might be filled with "80 people (small business)," `{CoreProductDescription}` with "Software as a Medical Device (SaMD), for remote patient monitoring," and the calculated `{FirstSurveillanceAuditMonth}` with "10th month." This step completes the transformation from a general template to a specific company's initial plan.

[0105] Step S240: Based on the target certification plan template after filling in the enterprise characteristic data, generate a complete certification plan that includes the audit stage sequence, time nodes and audit type definitions.

[0106] The completed template is then rendered and formatted to generate a formal certification plan document that can be manually reviewed, confirmed, and implemented. All placeholders in the template have been replaced with specific, company-related data. Based on the document structure and logic defined in the template, a complete plan is generated. This plan clearly outlines the overall roadmap for the certification project, with core content including: a sequence of audit phases, clearly listing all stages from project initiation to certificate issuance and subsequent maintenance, such as "Gap Analysis and Preparation Phase," "System Document Establishment and Release Phase," "Internal Audit and Management Review Phase," "Preliminary Audit by Certification Body Phase," "First Surveillance Audit Phase," "Second Surveillance Audit Phase," and "Recertification Planning Phase." Timeframes are specified for each major stage and key activity (such as on-site audit date, document submission deadline, and management review meeting time), with specific planned dates or relative times (e.g., "within 10-12 months after the preliminary audit"). These timeframes are determined by the entered characteristic data (such as company size and product complexity) and pre-defined scheduling logic. The audit type definition clarifies the nature of each external audit performed by the certification body, such as "initial certification audit," "surveillance audit (first time)," or "recertification audit," and may include a description of the key areas of focus for that audit. Further, step S220 involves matching the company attributes with the applicable conditions of pre-stored certification plan templates in the certification knowledge base, including:

[0107] Step S221: Based on the preset matching rules in the certification knowledge base, perform logical judgment on the enterprise attributes to generate preliminary template matching results.

[0108] This step is the first-level rapid filtering mechanism in the template matching process. Its core is to execute a set of predefined, deterministic logical rules to achieve efficient and accurate processing of common or typical enterprise certification scenarios. The preset matching rules stored in the certification knowledge base are a collection of "condition-action" statements, refined by domain experts based on historical certification projects, standards, and best practices. The rules directly apply to the enterprise attributes (industry classification, certification history, target certification standards, etc.) extracted in step S210, using logical operators (such as AND, OR, NOT) and comparison operators for combined judgment. The input enterprise attributes are iterated through all preset rules; when all conditions are met, the corresponding action is triggered, marking the certification plan template pointed to by that rule as a match. This process is fast and provides deterministic results, capable of directly handling a large number of well-structured and clearly defined matching requests. For example, for a pure software development company applying for ISO 9001 certification for the first time, the rule engine can quickly match the "ISO 9001 First-Time Certification (Software Industry Applicable Version)" template and output this single result as the initial matching result.

[0109] Step S222: When the preliminary template matching result has multiple candidate templates or cannot be determined, the enterprise attributes are identified and feature extracted based on the preset decision prompt words in the authentication knowledge base.

[0110] When the rule engine in step S221 fails to produce a unique match (e.g., returning multiple candidate templates, or failing to trigger any rules due to unique attribute combinations), a deep analysis process based on a large language model will be initiated. At this point, "decision prompts" pre-set in the certification knowledge base for this type of complex decision-making scenario are invoked. These prompts are structured text instructions designed to guide the large language model to act as a seasoned certification expert, conducting in-depth analysis of enterprise attributes. Their content not only includes instructions for the model to perform tasks but also typically provides the framework and steps for the analysis.

[0111] Step S223: Based on the results of authentication scenario identification and feature extraction, determine the priority of main business and compliance risks in the enterprise attributes.

[0112] The features extracted in step S222 may include: the company has multiple business segments (such as "smart hardware manufacturing" and "online health services"), or it applies for multiple certification standards simultaneously (such as "ISO 13485" and "ISO 27001"). This step requires interpreting and weighing these features. Based on the embedded domain knowledge logic or the model's own reasoning ability, first determine which business is the company's "core business". The criteria for judgment may include: the proportion of each business description, the revenue share implied (if included in the attributes), its core position in the industry classification, and its historical correlation with the target certification standard. Secondly, and more importantly, assess the "compliance risk priority" corresponding to different target certification standards. This is usually based on factors such as the mandatory nature of regulations, the severity of the consequences of violations, and the intensity of industry supervision. Combining the core business with the compliance risk priority analysis forms the focus of decision-making: for companies with multiple businesses, the plan should prioritize ensuring that the certification requirements of high-risk core businesses (such as medical devices) are covered, and on this basis, reasonably integrate or arrange the audit of other standards (such as information security). The output of this step is a clear weighting guide, such as: "Main business: Medical device software; Highest compliance risk priority: ISO13485; Secondary priority: ISO 27001".

[0113] Step S224: Based on the priority of main business and compliance risks, select the final target certification plan template from the candidate templates.

[0114] Based on the explicit weighting guidelines generated in step S223, the multiple candidate templates generated in step S221 are finally evaluated and selected. At this point, selection is not random, but rather based on the priority of core business and compliance risks, assessing the degree to which each candidate template meets the core requirements. The evaluation process may involve comparing the template's "applicable conditions" description, the structure of the template content, or conducting comparative analysis between templates by using the priority guidelines as secondary query input to a large model. For example, candidate templates may include "TMPL_13485_FIRST" (focusing on initial ISO 13485 certification), "TMPL_27001_FIRST" (focusing on initial ISO 27001 certification), and "TMPL_13485_27001_INTEGRATED" (a template for integrated certification of the two standards). Based on the guideline of "prioritizing coverage of ISO 13485," "TMPL_13485_FIRST" or "TMPL_13485_27001_INTEGRATED" would be preferred. Next, further assessment will be made: if the company's business is highly integrated and resources allow for parallel audits, then the "integrated certification template" may be better, as it can coordinate audit activities and reduce duplicate work; if the company wishes to implement in stages, it may choose the pure ISO 13485 template first. Taking into account these subtle considerations, the template that best fits the main business, most efficiently manages the highest compliance risks, and also takes into account the company's actual situation and the ease of operation for the certification body will be selected as the "final target certification plan template," and its unique identifier will be output and passed to the subsequent template filling steps.

[0115] Further, in step S300, the certification plan is correlated and matched with the structured expert experience in the certification knowledge base to determine the set of items to be audited, including:

[0116] Step S310: Extract the certification standard identifier of the certification plan.

[0117] From the complete certification plan document generated in step S200, which already includes specific audit stages and timelines, the core identifiers of one or more authoritative standard systems directly upon which this certification task is based are automatically identified and extracted. The certification plan will explicitly state the standards covered by this audit; this information typically exists in the "Audit Basis" or "Certification Scope" section of the plan in the form of standard names, standard codes, and version numbers. For example, a plan might state, "The audit will be conducted in accordance with ISO / IEC 27001:2022 'Information Technology Security Technology - Information Security Management System Requirements' and ISO / IEC 27002:2022 'Control Guidelines'." Such standard statements are located and extracted from the certification plan text using predefined parsing rules or natural language processing models. The extracted "certification standard identifiers" are typically normalized, for example, uniformly mapped to unique codes used in the internal knowledge base. For instance, "ISO / IEC 27001:2022" is mapped to the identifier "ISMS-27001-2022", or "IATF 16949" is mapped to "IATF-16949". When a project involves multiple standards (such as integrating quality, environmental, and safety systems), all relevant identifiers are extracted to form an identifier set, such as ['QMS-9001-2015', 'EMS-14001-2015', 'OHSMS-45001-2018'], laying the foundation for accurately locating relevant clauses from a vast amount of knowledge.

[0118] Step S320: Based on the certification standard identifier, retrieve all check items corresponding to the certification standard identifier from the certification knowledge base. The check items are components of structured expert experience.

[0119] After obtaining a clear certification standard identifier, this step performs an efficient knowledge base query to acquire all audit knowledge units corresponding to the target standard that can be directly processed by the machine. During the offline construction phase, the certification knowledge base has already intelligently parsed, converted, and stored various regulatory and standard documents (such as full texts of ISO standards and industry-specific specifications) into structured "expert experience." This experience is organized in the basic form of "checklist items," each checklist item being a well-encapsulated data object, typically containing several key fields, such as: the unique identifier of the standard to which it belongs, the clause number (e.g., "A.5.1"), the checklist item content text, a description of the assessment implementation method, compliance assessment criteria, and possible recommendations on associated risk levels or evidence types. Using the certification standard identifier extracted in step S310 as the query key, a precise match is performed in the knowledge base index to quickly retrieve all checklist records whose "belonging standard identifier" field matches. For example, entering the identifier "ISMS-27001-2022" will return hundreds of structured checklist items under that standard, from "A.5 Leadership and Commitment" to "A.10 Improvement." This retrieval process enables a rapid mapping from macro standards to micro audit points, ensuring the comprehensiveness of knowledge retrieval. These checklists form a complete knowledge base for subsequent audit tasks.

[0120] Step S330: Merge all inspection items to form an initial inspection item set.

[0121] After step S310 extracts multiple certification standard identifiers and step S320 retrieves multiple independent sets of inspection items from the knowledge base, these inspection items from different standard systems but potentially serving the same certification project need to be physically merged and logically integrated. A new set (i.e., the "initial inspection item set") is created, and all retrieved inspection item objects are added to this set. During this process, potentially duplicated or highly similar inspection items need to be intelligently handled. For example, ISO 9001 (quality management) and IATF 16949 (automotive industry quality) may have many similar clauses in basic requirements such as "document control" and "record management." Based on semantic similarity calculations of the inspection item content or predefined mapping rules, these essentially identical inspection items are identified, and deduplication or association is performed during merging to avoid repeated evaluation of the same facts in subsequent audits. The resulting "initial inspection item set" is a complete list containing all standards and clause requirements involved in this certification, representing the theoretically maximum audit scope and providing a comprehensive candidate pool for precise selection based on actual project needs in the next step.

[0122] Step S340: Based on the audit scope defined in the certification plan, select the audit items related to the audit scope from the initial set of audit items to form the audit item set.

[0123] The "defined audit scope" in the certification program is typically described in natural language, such as: "This audit scope covers the information security management system of the company's R&D center and data center located in City A, excluding its marketing department in City B." Furthermore, the scope may be clarified through limiting clauses (such as "Audit only in control domains A.6 to A.9 of ISO 27001 Annex A") or enumeration of products / services / locations. Based on natural language understanding technology, this scope description is parsed to extract key limiting elements, such as physical location (R&D center in City A), organizational boundaries (excluding the marketing department), business processes (R&D, data hosting), and standard subdomains (A.6-A.9). Subsequently, each check item in the "initial checklist" is traversed to determine its relevance to the audit scope. The judgment logic could be: whether the check item's "applicable scenario" metadata includes "data center"; whether its clause number falls within the range of A.6 to A.9; or whether semantic analysis is used to determine whether the business activities described by the check item fall within the "R&D" category. This matching process filters out all inspection items relevant to the defined scope, while eliminating those explicitly excluded or irrelevant, ultimately outputting a precisely sized and targeted "set of inspection items to be reviewed." This set serves as the direct and complete task list for the subsequent step S400, which automates the conformity assessment of specific evidence materials.

[0124] Furthermore, step S400, based on the set of items to be audited, determines the conformity of the certification evidence materials, including:

[0125] Step S410: For each inspection item in the set of inspection items to be audited, retrieve the evidence content related to the inspection item requirements from the certification evidence materials.

[0126] For each specific inspection item requirement, relevant evidentiary information fragments are proactively and precisely located and extracted from all certified evidence materials. To achieve efficient and accurate cross-modal retrieval, a unified semantic vector index has been constructed for all certified evidence materials during the preprocessing stage. The index construction process is as follows: for text materials (such as program files and interview records), their text content is directly extracted; for tables, their row and column structures are parsed and converted into descriptive text sequences; for image materials (such as configuration screenshots), text information is extracted using OCR technology. All extracted text content is converted into a fixed-dimensional high-dimensional vector representation by a pre-trained semantic encoding model (such as Transformer-based Sentence-BERT or similar models) and stored in the vector database along with its metadata (such as source file name, page number, and coordinate position). When evidence retrieval is required for a specific inspection item (e.g., "IS 5.4: Ensure that the information security management system complies with the requirements of relevant legal documents A"), the text requirement of that inspection item (possibly along with its associated regulatory context) is first input into the same semantic encoding model to generate a "query vector". Subsequently, an approximate nearest neighbor search is performed in the vector database to calculate the cosine similarity or other semantic similarity measures between the query vector and all evidence vectors, quickly identifying the original evidence fragments corresponding to the Top-K vectors with the highest similarity. These fragments are initially considered the "evidence content" most relevant to the inspection requirements. They can be a statement in a program file, a dialogue in an interview transcript, a text description in a system configuration screenshot, or a log record. This process achieves intelligent association from "inspection requirements" to "multi-source heterogeneous evidence."

[0127] Step S420: Compare the content of the evidence with the text requirements of the inspection items for consistency.

[0128] After obtaining the relevant evidence, this step requires a detailed consistency analysis of the facts or states stated in the evidence with the normative requirements of the inspection item. This comparison is not a simple keyword matching, but rather a semantic understanding and logical relationship judgment. It is necessary to identify whether the evidence content "supports," "satisfies," "partially satisfies," "does not satisfy," or "refutes" the requirements of the inspection item. For example, the inspection item requires that "information security risk assessments should be conducted regularly (at least annually)," while the retrieved evidence content may include: "Article 3.2 of the Risk Assessment Management Procedure stipulates that 'a company-wide risk assessment should be completed in the first quarter of each year'" (directly supporting); "The 2023 annual risk assessment report was released on March 15, 2023" (supported by specific facts); "The company has not yet established a formal risk assessment process" (directly refuting); or "We have relevant awareness, but the implementation time is not fixed" (partially satisfied but with deficiencies). This in-depth semantic relationship judgment is performed on each piece of relevant evidence content and the requirements of the inspection item through natural language reasoning or textual implication recognition technology. For more complex situations, such as when the evidence describes specific control measures (e.g., "adopt two-factor authentication to log in to the core system"), while the inspection requirement is a higher-level principle (e.g., "strong identity authentication should be performed for access"), it is necessary to determine that the former is a specific implementation instance of the latter, thus confirming consistency support. The output of this step is a preliminary qualitative judgment set on the consistency relationship between each relevant piece of evidence and the inspection requirement.

[0129] Step S430: Based on the preset evaluation criteria corresponding to the inspection items in the certification knowledge base, determine the consistency comparison results and identify the compliance status.

[0130] Based on established, objective rules, the consistency relationships of the multiple pieces of evidence generated in step S420 are comprehensively adjudicated, outputting a clear status of "compliant," "non-compliant," or "partially compliant." In the certification knowledge base, each structured check item is associated with one or more "pre-defined evaluation criteria." These criteria define specific logical rules for drawing conclusions based on evidence, and their form may be logical expressions, decision trees, or rules described in natural language. For example, for a certain check item, its evaluation criteria might be described as: "If there is a formally published procedure document that explicitly stipulates the requirement, and at least one record of effective execution (such as a report or log) is provided, it is judged as 'compliant'; if there is only a procedure document but no effective execution record, it is judged as 'partially compliant'; if there is neither a procedure document nor an execution record, it is judged as 'non-compliant'." The consistency judgment results of each piece of evidence in step S420 (such as "Evidence A: Procedure document, supports"; "Evidence B: Execution record, supports"; "Evidence C: Interview record, indicating execution flaws") are used as input and substituted into the pre-defined evaluation criteria for that check item for logical calculation. It is necessary to comprehensively weigh the strength, directness, and reliability of supporting and rebuttal evidence. For example, even if supporting procedural documents exist, if strong rebuttal evidence (such as non-compliance reports found during audits) is also present, the result may be determined as "partially compliant" or "non-compliant" based on the standards. This process simulates the evidence weighing logic of audit experts, ultimately generating a definite compliance status and a brief summary of the reasons for the determination for each inspection item.

[0131] Step S440: Based on the determination of the conformity status of all inspection items in the set of inspection items to be audited, generate an audit conclusion on conformity determination.

[0132] The results of all individual inspection items are summarized, organized, and formatted to generate a structured draft audit conclusion for the certification body to use. The entire "set of inspection items to be audited" is traversed, collecting the final compliance status (compliant / non-compliant / partially compliant) for each item, the key evidence cited for the judgment (such as document name, location, relevant excerpts), and a brief reasoning from step S430. Subsequently, this information is organized into a complete audit findings summary report according to industry-standard reporting formats or predefined templates. The report is typically grouped by standard clauses or control areas, clearly listing the item number and content, judgment status, corresponding index of objective evidence, and any further explanations required (especially for "non-compliant" and "partially compliant" items). Furthermore, a high-level conclusive summary may be generated based on the overall compliance of all inspection items, indicating areas where the system performs well and areas with systemic weaknesses. This generated "audit conclusion" is no longer a simple list of conclusions, but a preliminary audit report draft with a clear chain of evidence, well-founded judgments, and standardized format. It provides a comprehensive, accurate, and structured foundation for certification teachers to conduct final professional review, write formal reports, and communicate audit findings with enterprises.

[0133] Furthermore, step S410, retrieving evidence content related to the inspection requirements from the certified evidence materials, includes:

[0134] Step S411: Construct a semantic vector index for the authentication evidence materials, and uniformly represent the text content, table content and image OCR results as high-dimensional vectors.

[0135] After the initial analysis and integration of the evidence materials, these multi-source, heterogeneous original contents need to be deeply processed to transform them into a unified mathematical representation that machines can understand and efficiently compute. To this end, a deep semantic coding model pre-trained on a large-scale text corpus is used. This model can map text input of arbitrary length into a fixed-dimensional, high-dimensional dense vector (i.e., a semantic vector). This vector can effectively capture the deep semantic information of the text, not just surface keywords. The processing flow is as follows: For text-based evidence materials (such as program files and interview transcripts), their plain text content is directly extracted; for table-based evidence, their structural information (such as "row header: access control policy, column header: review cycle, cell value: every six months") is converted into a descriptive natural language text; text information extracted from system architecture diagrams and configuration interface screenshots via OCR is also used as text input. All text fragments, along with their key metadata identifiers (such as source file name, chapter, and coordinates in the image), are fed into the semantic coding model in batches. The model generates a unique semantic vector for each text fragment. Subsequently, all generated vectors, along with their corresponding metadata, are stored and indexed in a dedicated vector database. This process essentially constructs a "semantic map" covering all evidence materials, where the position of each evidence point (vector) in the semantic space represents its meaning. Semantically similar evidence points are also close to each other in the vector space, laying the foundation for subsequent semantic similarity-based retrieval.

[0136] Step S412: The text requirements of the inspection items are also converted into query vectors.

[0137] The textual requirements of the inspection item are transformed into a mathematical representation that is in the same semantic space as the evidence materials and can be directly compared. The text content of the inspection item to be retrieved is extracted, which typically includes the descriptive requirements of the inspection item itself, such as "Important business data should be backed up regularly, and the backup cycle should not exceed 24 hours." To enhance the context and accuracy of the query, more contextual information associated with the inspection item can be included, such as the title of its superior clause or relevant excerpts of regulations from the certification knowledge base. This combined query text is fed into the same deep semantic encoding model used in step S411 when building the index. This model encodes the query text with the exact same algorithm and parameters, outputting a fixed-dimensional "query vector." This query vector is mathematically homogeneous with the evidence vectors in the vector database and is embedded in the same high-dimensional semantic space. Therefore, the semantic relevance between the inspection item requirements and potential evidence can be quantitatively measured by calculating the distance or similarity (such as cosine similarity) between the query vector and the evidence vector. This step ensures that the retrieval is based on deep semantic matching rather than shallow literal matching, enabling the system to identify evidence that is expressed differently but has the same meaning. For example, the textual evidence of "performing a full data backup every day" can be correctly associated with the check requirement that "the backup cycle must not exceed 24 hours".

[0138] Step S413: Perform similarity retrieval in the semantic vector index to find the multiple evidence fragments most relevant to the query vector.

[0139] Once the query vector corresponding to the inspection item is obtained, it is used as search input and submitted to a vector database with a pre-built semantic vector index. The vector database employs an efficient approximate nearest neighbor search algorithm, traversing all evidence vectors in the index within milliseconds and calculating a semantic similarity score (usually cosine similarity, with a value range of [-1, 1], where higher values ​​indicate greater semantic similarity) between each evidence vector and the query vector. The algorithm aims to quickly find the set of evidence vectors with the highest similarity to the query vector. Instead of finding only a single best match, it sets a threshold or a return quantity K to find the top K evidence vectors with the highest similarity. For example, K might be set to 10 or 20. The original text fragments (and their metadata) corresponding to these retrieved evidence vectors are considered the "evidence fragments" most likely relevant to the current inspection item requirement. This vector similarity-based retrieval overcomes the problem of keyword mismatch, rapidly narrowing the scope from massive amounts of evidence to locate a set of materials semantically relevant to the core.

[0140] Step S414: Sort the retrieved evidence fragments based on semantic relevance and return the top K most relevant evidence fragments as the search results.

[0141] After initially retrieving multiple relevant evidence fragments in step S413, this step refines and filters these fragments to output the highest quality and most relevant final result set. Based on the original similarity score between each evidence fragment returned by the vector database and the query vector, they are sorted in descending order to form a preliminary list from high to low semantic relevance. However, more complex sorting logic may be introduced to optimize result quality. For example, the authority of the evidence source can be considered, assigning higher weight to procedural documents and formal policy evidence than to interview transcripts and temporary notes; or, considering the length and completeness of the retrieved fragments, excessively short fragmented texts may be appropriately downweighted. The sorted list is then truncated, and the top K evidence fragments are selected as the official output of this retrieval. This result set is encapsulated as a structured list, where each entry contains the text content of the evidence fragment, the source document, the specific location (such as page number, coordinates), and its calculated similarity score to the requirements of the inspection item.

[0142] Furthermore, before receiving the authentication evidence materials in step S100 and extracting the structured information of the authentication evidence materials based on the authentication knowledge base, the process also includes:

[0143] Step S101: Perform layout analysis and content reconstruction on the certification regulatory standard document, and convert the certification regulatory standard document into a standardized document with a hierarchical heading structure.

[0144] Please refer to Figure 4 The input certification regulations and standards documents are typically in PDF format, which may be a normal PDF containing selectable Chinese text or a scanned image PDF. First, the nature of the PDF is determined using document type detection technology. For scanned or photocopied PDFs, optical character recognition (OCR) technology is used to convert them into an editable text stream, preserving the original positional information of the characters. Then, a pre-trained layout analysis model, such as a deep learning-based document layout recognition model, is invoked to perform semantic segmentation of the page, accurately identifying and labeling various elements and their bounding boxes, including headers, footers, body paragraphs, headings at all levels, tables, images, formulas, etc. Based on the obtained layout elements, the content is reconstructed according to its visual layout and logical order: irrelevant elements such as headers and footers are filtered out; for multi-column layouts, the content is rearranged according to the reading order from top to bottom and left to right, converting it into a coherent single-column text sequence. Finally, the reconstructed content is output as a Markdown document with a clearly defined hierarchical heading structure. During this process, the heading styles in the original text (such as font, font size, bold) will be identified and automatically converted into Markdown headings of different levels through regular expression matching or style analysis, forming a complete tree-like document structure from chapters, sections, articles to clauses, providing standardized input for subsequent knowledge extraction.

[0145] Step S102: Identify cross-reference relationships in standardized documents and construct a cross-reference knowledge graph containing information on the association of referenced clauses.

[0146] Based on the standardized Markdown document generated in step S101, natural language processing techniques, particularly entity linking and relation extraction, are used to automatically scan the entire document and identify all explicit cross-reference expressions. These expressions typically have specific patterns, such as "see Clause 5.2.1," "shall comply with the requirements of Clause 8.1 of GB / T 22080-2016," and "in accordance with the provisions of Appendix A." Through pattern matching combined with contextual analysis, key identifiers of the referenced objects, such as target standard numbers, chapter numbers, clause numbers, or appendix numbers, are accurately extracted from these statements. Subsequently, a node is created in the knowledge base for the currently processed document, and for each identified cross-reference relationship, a directed edge is established between the citation source (current clause) and the referenced target (the clause or document it refers to). All nodes and edges together constitute a cross-reference knowledge graph. Attribute information can be attached to each edge, such as citation type ("see," "in accordance with," "shall comply with"), and the context in which the citation occurs.

[0147] Step S103: Based on natural language processing technology, locate the chapter range where the check item is located in the standardized document.

[0148] The goal of this step is to automatically identify the core text areas containing specific audit requirements (i.e., inspection items) from the entire regulatory standard document, thus defining the scope for subsequent refined extraction. Inspection items are typically scattered throughout specific chapters of the standard, such as the "Assessment Unit" section in the "Information Security Technology: Network Security Level Protection Assessment Requirements." This localization is achieved by combining multiple technologies. First, the document's hierarchical heading structure (Markdown's H1-H6 headings) can be used as prior navigation. More importantly, a natural language processing model is used to understand the semantics of the content. Specifically, a specific query instruction can be constructed and input into the large language model, describing the text features to be located, such as: "Please find all chapters in the document that contain specific assessment requirements, assessment implementation steps, or unit judgments." Simultaneously, the complete chapter structure information (heading tree) of the document is provided to the model as context. The large language model, by understanding the query intent and analyzing the semantics of chapter titles and content, outputs one or more consecutive lists of chapter scope identifiers.

[0149] Step S104: Based on the structural characteristics of the chapter scope, determine the extraction rules for the inspection item type, inspection item content, evaluation content, and evaluation criteria.

[0150] Because the expression format of inspection items may differ between different standards, and even between different chapters of the same standard, it is necessary to adaptively determine the extraction rules. The text structure and language features within the identified chapter scope are analyzed. For inspection item types, they are usually associated with heading levels. By analyzing the structure of the chapter heading tree, it is possible to identify which level of heading directly contains the specific inspection item description, thus determining the specific level index (e.g., level 3 and level 4 headings) corresponding to the "inspection item type". For inspection item content, evaluation content, and judgment criteria, it is necessary to analyze the textual patterns within paragraphs. A large language model can be used again, using typical text fragments and task instructions from the chapter as input. An example instruction is: "Please analyze the following text and find the field names corresponding to the content following 'evaluation indicators', 'evaluation implementation', and 'unit judgment' respectively." These rules essentially define the mapping relationship from unstructured text to structured fields, serving as templates for subsequent information decomposition.

[0151] Step S105: Based on the extraction rules, extract the structured inspection item types, inspection item content, evaluation content and evaluation criteria in the chapter scope to obtain structured expert experience.

[0152] Based on the extraction rules determined in step S104, the text within the located chapters is subjected to batch structured information extraction to form "expert experience" data units that can be directly used by computers. The processing is rule-driven. First, according to the title level index defined in the "check item type" extraction rule, all title paths that match the level are extracted from the document title tree as the "type" information for each check item. Then, the text content within the chapter range is processed segment by segment. For each segment of text that may contain a complete check item, extraction rules for fields such as "check item content," "evaluation content," and "judgment criteria" are applied. For example, if the rule is defined as "the content following the evaluation indicator is the check item," then the pattern "evaluation indicator:" is located in the text segment, and the subsequent text fragments are extracted as the value of the "check item content" field in the structured object.

[0153] Step S106 involves integrating standardized documents, cross-referenced knowledge graphs, and structured expert experience to construct a certification knowledge base.

[0154] The system integrates all outputs from the preceding steps. The core of this integration is organization by "certification standards." For each completed regulatory standard, its corresponding "standardized document" (Markdown format) is stored as the raw text archive. A "cross-reference knowledge graph" is stored as a semantic network revealing the internal and external relationships of the standard. "Structured expert experience" (i.e., a set of checkpoint objects) is stored as directly accessible audit knowledge. These three are linked and indexed using unified identifiers (such as standard number and version number). The resulting certification knowledge base is a hierarchical data system: the top layer is the standard metadata index; the middle layer is the structured expert experience database and relationship graph corresponding to each standard; and the bottom layer is linked to the standardized raw text. This knowledge base provides rich query interfaces; for example, all checkpoints can be retrieved by standard ID, and all cross-reference relationships can be queried by clause number.

[0155] Furthermore, in step S105, based on extraction rules, the structured inspection item types, inspection item content, evaluation content, and evaluation criteria within the chapter scope are extracted to obtain structured expert experience, including:

[0156] Step S1051: Based on the check item type index defined in the extraction rules, determine the corresponding target heading level from the hierarchical heading structure of the standardized document, and take all subheading path sequences belonging to the target heading level as check item types.

[0157] The "Index of Check Item Type" in the extraction rule is an explicit indicator that defines which level of heading in the standard document's heading tree directly represents the logical classification of the check item. For example, a rule might specify an index of [3, 4], which means that in the document tree consisting of Markdown H1 to H6 headings, the paths of the third and fourth level headings together define the check item type. First, the complete heading hierarchy of the standardized document is accessed, which is a tree-like data structure starting from the root node. Then, the target heading level is located based on the index. Next, the document is traversed, collecting all heading nodes located at that target level. For each such node, instead of just taking its heading text, the complete heading path sequence from the root node to that node is obtained. For example, in the "Network Security Level Protection Assessment Requirements," a complete path sequence might be: "6 Level 1 Assessment Requirements -> 6.1 General Security Assessment Requirements -> 6.1.1 Secure Physical Environment -> 6.1.1.1 Physical Access Control". This complete path sequence, rather than isolated "physical access control," is used as the "check item type" for all sub-check items belonging to that node.

[0158] Step S1052: Based on the various positioning identifiers defined in the extraction rules, locate and extract the text fragments associated with the positioning identifiers from the text content of each chapter range, so as to form structured inspection items, evaluation content and evaluation criteria respectively.

[0159] The extraction rules define "location identifiers" for fields such as "Inspection Item Content," "Evaluation Content," and "Judgment Criteria." These identifiers are typically key phrases or tags with fixed patterns within paragraphs. For example, in a standard "Evaluation Unit" description, common location identifiers include "Evaluation Indicator:," "Evaluation Implementation:," and "Unit Judgment:." The text within the section located in step S103 is scanned. For each paragraph of text to be processed (usually corresponding to one evaluation unit), the occurrence position of each location identifier is sequentially searched according to the rules. Once found, the text segment starting from that identifier and continuing until the next location identifier appears or the paragraph ends is extracted. For example, if the text "Evaluation Indicator: The entrance and exit of the computer room should be staffed by designated personnel or equipped with an electronic access control system." is encountered, "The entrance and exit of the computer room should be staffed by designated personnel or equipped with an electronic access control system." will be extracted as the value of the "Inspection Item Content" field. Through this rule-based pattern matching, structured key information can be extracted stably and in batches from complex documents.

[0160] In step S1053, during the extraction process, if a cross-reference relationship is identified within the processed text fragment, the cross-reference knowledge graph is queried to obtain the full text of the corresponding cited clause, and the obtained full text content is dynamically embedded into the structured entry currently being constructed.

[0161] When extracting text fragments in step S1052, these fragments may contain cross-reference markers identified in step S102, such as "according to Clause 3.2 of Document A" or "see Clause 5.1.3". If the raw cited text is stored directly, the auditor may still need to manually search for the cited content when using the check item, disrupting the automated process. To address this issue, after extracting the text of each field, it is immediately scanned to check for known cross-reference patterns. Once such a reference is identified, it is used as the query key to query the "cross-reference knowledge graph" constructed in step S102. The knowledge graph returns the complete, standardized text content of the cited clause. An "embedding" operation is then performed: the returned full text of the cited clause is inserted or appended to the corresponding field of the currently constructed structured check item in a specific format (e.g., as a citation or as an additional explanatory field). For example, if the check item content cites another clause, the full text of that clause is embedded as supplementary information in the "check item content" field as "citation context". In this way, the final generated structured check item object is no longer just an isolated sentence, but a self-sufficient knowledge unit containing all the necessary original text references.

[0162] Step S1054: The inspection item type, the inspection item content with completed context embedding, the evaluation content and the evaluation criteria are encapsulated and stored in a structured manner with the inspection item as the basic unit to obtain structured expert experience.

[0163] Having completed the preceding steps, all the constituent elements have been prepared for each original "assessment unit" or similar structure: the complete type path from step S1051, and the core field content from step S1052, which has already undergone context embedding in step S1053. This step encapsulates these discrete elements according to a predefined, unified pattern. A new data structure is created, typically a JSON object or database record, with its fields corresponding one-to-one with the extracted elements. Thousands of such objects are organized in units of standards and stored in batches in a dedicated storage area of ​​the certification knowledge base (such as a document database or a node set of a graph database). During storage, efficient indexes are built using fields such as standard IDs and clause paths to ensure fast and accurate retrieval and association during the online phase. Ultimately, the collection of all these stored structured check item objects constitutes the "structured expert experience" knowledge base that drives the entire intelligent certification process.

[0164] Accordingly, a second aspect of the present invention provides an electronic device, including: at least one processor and a memory connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the at least one processor to perform the aforementioned authentication process efficiency improvement method.

[0165] Accordingly, a third aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described method for improving the efficiency of the entire authentication process.

[0166] The embodiments of the present invention aim to protect a method for improving the efficiency of the entire authentication process, which has the following effects:

[0167] 1. By automating the parsing of regulatory and standard documents and building an certification knowledge base containing cross-references, unstructured expert experience that traditionally relied on manual interpretation and memorization is transformed into structured knowledge that can be processed and correlated by machines. This fundamentally solves the bottleneck of unstructured information processing, enabling certification activities to be based on a unified, accurate, and context-complete knowledge foundation, and significantly improving the completeness and standardization of the audit.

[0168] 2. With intelligent agent technology as the core, it automatically formulates plans, associates check items and generates notifications based on structured information, realizing full-process automation from project initiation to report generation; by integrating rule engine and semantic reasoning, it can handle various decisions from clear rules to complex scenarios, freeing human resources from repetitive paperwork and allowing them to focus on high-value professional judgment, thereby significantly shortening the certification cycle and reducing operating costs while ensuring decision quality.

[0169] 3. Based on the constructed semantic vector index and dynamic evidence chain, it can actively retrieve evidence fragments related to the inspection items from multi-source heterogeneous materials, and perform traceable compliance reasoning according to preset evaluation criteria. This mode of transforming the "experience verification" of human auditors into the "evidence verification" of the system not only greatly improves the coverage depth and accuracy of the audit, but also effectively eliminates individual subjective differences through the unified standard execution, ensuring high consistency of certification results.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for improving efficiency throughout the entire authentication process, characterized in that, Includes the following steps: Receive authentication evidence materials and extract structured information from the authentication evidence materials based on the authentication knowledge base; Based on the structured information, an authentication plan is developed; The certification plan is matched with the structured expert experience in the certification knowledge base to determine the set of items to be audited; Based on the set of inspection items to be audited, the certification evidence materials are assessed for compliance to determine their compliance status with the inspection item requirements, and an audit conclusion is generated for auditing the verification form. The extraction of structured information from the authentication evidence materials based on the authentication knowledge base includes: The authentication evidence materials are parsed to obtain their text content; Based on the preset information extraction mode in the authentication knowledge base, the text fragments corresponding to the target information field set are identified and located from the text content. The preset information extraction mode includes the target information field set to be extracted required by the authentication process. Using the extraction rules in the authentication knowledge base, information extraction is performed on the text content to extract entity information or attribute values ​​from the text fragments and fill them into the corresponding target information fields; The filled target information field set is integrated to obtain the structured information of the authentication evidence material; The process of developing an authentication plan based on the structured information includes: Obtain the enterprise attributes contained in the structured information, including industry classification, certification history, and target certification standards; The enterprise attributes are matched with the applicable conditions of the pre-stored certification plan templates in the certification knowledge base, and the corresponding target certification plan template is selected. Extract enterprise feature data corresponding to the placeholders in the target certification plan template from the structured information, and fill the placeholders with the enterprise feature data; Based on the target certification plan template after the enterprise characteristic data is filled in, a complete certification plan is generated, which includes the audit stage sequence, time nodes and audit type definitions.

2. The method for improving efficiency throughout the entire certification process according to claim 1, characterized in that, The step of parsing the authentication evidence materials to obtain the text content of the authentication evidence materials includes: Obtain key document elements related to compliance proof from the certification evidence materials. These key document elements include system architecture diagrams, configuration interface screenshots, operation log records, and compliance declaration forms. Text information in the system architecture diagram and configuration interface screenshot is extracted using optical character recognition technology, and the position coordinates of the text information in the original image are recorded. Identify the timestamp format and event type of the operation log records, and reorganize the log events in a structured manner according to the time sequence; The row and column structure of the compliance declaration form is parsed to establish a mapping relationship between the clause requirements and the corresponding cell content, and the cross-referenced cell content is stored in association. The extracted text information, structured log events, and associated table content are integrated according to the evidence logic required for certification audit to generate complete text content for subsequent compliance determination.

3. The method for improving efficiency throughout the entire certification process according to claim 1, characterized in that, The step of matching the enterprise attributes with the applicable conditions of the pre-stored certification plan templates in the certification knowledge base includes: Based on the preset matching rules in the certification knowledge base, logical judgments are made on the enterprise attributes to generate preliminary template matching results; When the preliminary template matching result has multiple candidate templates or cannot be determined, the enterprise attribute is identified and its features are extracted based on the preset decision prompt words in the authentication knowledge base. Based on the results of authentication scenario identification and feature extraction, the priority of the main business and compliance risks in the enterprise attributes is determined; Based on the aforementioned main business and compliance risk priorities, the final target certification plan template is selected from the candidate templates.

4. The method for improving efficiency throughout the entire certification process according to claim 1, characterized in that, The step of associating and matching the certification plan with the structured expert experience in the certification knowledge base to determine the set of items to be audited includes: Extract the certification standard identifier from the certification plan; Based on the certification standard identifier, all check items corresponding to the certification standard identifier are retrieved from the certification knowledge base, and the check items are a component of the structured expert experience; Combine all the aforementioned check items to form an initial check item set; Based on the audit scope defined in the certification plan, check items related to the audit scope are selected from the initial check item set to form the check item set to be audited.

5. The method for improving efficiency throughout the entire certification process according to claim 1, characterized in that, The process of determining the conformity of the certification evidence materials based on the set of items to be audited includes: For each inspection item in the set of inspection items to be audited, retrieve the evidence content associated with the requirements of the inspection item from the certification evidence materials; The content of the evidence is compared with the textual requirements of the inspection items for consistency. Based on the preset evaluation criteria corresponding to the inspection item in the certification knowledge base, the result of the consistency comparison is judged to determine the compliance status. Based on the determination of the compliance status of all inspection items in the set of inspection items to be audited, an audit conclusion on the compliance determination is generated.

6. The method for improving efficiency throughout the entire certification process according to claim 5, characterized in that, The step of retrieving evidence content related to the inspection item requirements from the certified evidence materials includes: A semantic vector index is constructed for the authentication evidence materials, and the text content, table content and image OCR results are uniformly represented as high-dimensional vectors; The text requirements for the inspection items are also converted into query vectors; A similarity search is performed in the semantic vector index to find the multiple evidence fragments most relevant to the query vector; The retrieved evidence fragments are sorted based on semantic relevance, and the top K most relevant evidence fragments are returned as the search results.

7. The method for improving efficiency throughout the entire certification process according to claim 1, characterized in that, Before receiving the authentication evidence materials and extracting the structured information of the authentication evidence materials based on the authentication knowledge base, the process also includes: The layout analysis and content reconstruction of the certification regulations and standards documents are performed to convert them into standardized documents with hierarchical heading structures. Identify cross-reference relationships in the standardized documents and construct a cross-reference knowledge graph containing information on the association of referenced clauses; Based on natural language processing technology, the chapter range where the inspection item is located is identified in the standardized document; Based on the structural characteristics of the chapter scope, the rules for extracting inspection item types, inspection item content, evaluation content, and evaluation criteria are determined. Based on the extraction rules, the structured inspection item types, inspection item contents, evaluation contents, and evaluation criteria in the chapter scope are extracted to obtain the structured expert experience; The standardized documents, the cross-referenced knowledge graph, and the structured expert experience are integrated to construct the certification knowledge base.

8. The method for improving efficiency throughout the entire certification process according to claim 7, characterized in that, Based on the extraction rules, the structured inspection item types, inspection item content, evaluation content, and evaluation criteria within the chapter scope are extracted to obtain the structured expert experience, including: Based on the inspection item type index defined in the extraction rules, the corresponding target heading level is determined from the hierarchical heading structure of the standardized document, and all subheading path sequences belonging to the target heading level are taken as inspection item types. Based on the positioning identifiers defined in the extraction rules, text fragments associated with the positioning identifiers are located and extracted from the text content of the chapter range, respectively, to form structured inspection items, evaluation content and evaluation criteria. During the extraction process, if a cross-reference relationship is identified within the processed text fragment, the cross-reference knowledge graph is queried to obtain the full text of the corresponding referenced clause, and the obtained full text content is dynamically embedded into the structured entry currently being constructed. The structured expert experience is obtained by structurally encapsulating and associating the inspection item type, the inspection item content with completed context embedding, the evaluation content and the evaluation criteria with the inspection item as the basic unit.

Citation Information

Patent Citations

  • Intelligent auditing authentication management method, system, equipment and medium

    CN118261564A

  • Metering system authentication management system and method

    CN121010389A