Cross-system integrated audit path optimization method and system based on industry knowledge graph

CN122264494BActive Publication Date: 2026-09-11BEIJING HUAXIA MEIXING QUALITY CERTIFICATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610299294.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-12
Publication Date
2026-09-11
Estimated Expiration
2046-03-12

AI Technical Summary

Technical Problem

由于各体系标准由不同机构制定,可能导致实际审核流程复杂、效率低下

Benefits of technology

采集目标行业多源异构数据并进行知识抽取、冲突消解,构建行业知识图谱;实现多源数据的规范归集,消除冲突冗余,建立跨体系与生产流程的关联,形成结构化知识体系,为后续审核提供可靠知识支撑;基于行业知识图谱对企业待审资料进行向量嵌入和语义匹配,剔除重复冲突内容得到一体化审核文档库;规范待审资料格式,提升规整度与可用性,为审核路径生成提供标准化数据基础;基于一体化审核文档库提取相关数据生成初步审核路径,经区域离散化、代价修正后优化得到一体化现场审核路径;实现审核路径的生成与优化,提升现场审核效率与规范性;将不符合项信息输入知识图谱溯源形成关联簇,结合历史案例匹配得到根源整改方案;实现不符合项分类溯源与案例复用,提升整改针对性与规范性;将审核全量数据反馈至知识图谱完成增量更新,复盘路径得到优化约束条件;实现数据闭环复用与知识图谱演进,为后续审核路径动态优化提供依据,发挥数据持续价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264494B_ABST
    Figure CN122264494B_ABST
Patent Text Reader

Abstract

The application provides a cross-system integrated audit path optimization method and system based on an industry knowledge graph, relates to the technical field of data processing and artificial intelligence, and comprises the following steps: based on an integrated field audit path, inputting non-compliance item information into the industry knowledge graph for association and tracing, obtaining an association non-compliance item cluster, and performing similarity matching based on a historical rectification case library to obtain a root cause rectification scheme; taking the root cause rectification scheme and the path time consumption, non-compliance item distribution and rectification effect of the current audit process as total data, feeding back to the industry knowledge graph for incremental updating to obtain an evolved knowledge graph; and using the evolved knowledge graph to review the integrated field audit path to obtain path optimization constraint conditions for subsequent dynamic optimization of the audit path. The application is based on the industry knowledge graph, realizes the dynamic evolution of the audit path through multi-source data fusion and path optimization, and improves the audit efficiency and standardization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data processing and artificial intelligence technology, and in particular to a cross-system integrated audit path optimization method and system based on industry knowledge graphs. Background Technology

[0002] Automotive parts manufacturers typically need to meet the certification requirements of four management systems simultaneously: ISO 9001, ISO 14001, ISO 45001, and IATF 16949. Because these standards are developed by different organizations, the actual audit process can be complex and inefficient.

[0003] At the data processing level, the following shortcomings exist: Regarding the collection and integration of multi-source heterogeneous data, enterprises need to manage a wide variety of audit materials, including quality manuals and work instructions. These materials have different formats, covering structured, semi-structured, and unstructured data. Due to the lack of a unified collection and standardized processing mechanism, data cannot be automatically aligned and formatted, requiring manual cleaning and terminology standardization, which may lead to version confusion and data omissions. Regarding cross-system knowledge extraction and correlation, there are complex cross-correlationships between the standard clauses of various systems, such as the high correlation between ISO 9001 process control and IATF 16949 production part approval procedures. Due to the lack of entity identification and relationship extraction capabilities, these correlations cannot be automatically identified, potentially leading to fragmented storage of data from different systems and difficulties in controlling the same process. The requirements are inconsistent and even conflicting across different documents; in terms of in-depth data mining, there is a lack of correlation analysis capabilities for non-conformity data; for example, if environmental non-conformity such as excessive dust in the workshop and occupational health non-conformity such as employees not wearing masks originate from the same source (unmaintained dust removal facilities), multiple non-conformities are handled independently due to the inability to identify this causal relationship, resulting in repeated rectification; at the same time, data from each system is stored independently, making it impossible to uniformly summarize and perform correlation analysis; finally, in terms of audit data feedback and reuse, there is a lack of systematic collection and feature extraction of historical audit data, and data such as path time and non-conformity distribution are not effectively structured and stored, making it impossible to form a reusable feature dataset. Subsequent audit path generation still relies on static templates, which may make it difficult to achieve dynamic optimization based on historical data. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a cross-system integrated review path optimization method and system based on industry knowledge graph. Based on industry knowledge graph, through multi-source data fusion and path optimization, the review path can be dynamically evolved, thereby improving review efficiency and standardization.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a cross-system integrated audit path optimization method based on industry knowledge graphs, the method comprising: Collect multi-source heterogeneous data from the target industry, extract knowledge and resolve conflicts, and construct an industry knowledge graph covering cross-system standard associations and production process flows; Based on industry knowledge graphs, vector embedding and semantic matching are performed on the audit materials of target companies. By identifying and eliminating duplicate and conflicting content, an integrated audit document library is obtained. Based on the integrated audit document library, cross-system clause association paths and historical audit data are extracted from the industry knowledge graph to obtain the preliminary audit path; based on the preliminary audit path, key audit areas are delineated and spatially discretized to obtain audit feature data; based on the audit feature data, the path cost correction coefficient is calculated, and the dwell time and inspection order of the preliminary audit path are optimized and adjusted to obtain the integrated on-site audit path. Based on the integrated on-site audit path, non-conformity information is input into the industry knowledge graph for correlation and source tracing to obtain clusters of related non-conformities. Then, based on the historical rectification case library, similarity matching is performed to obtain root cause rectification solutions. The root cause rectification plan, the path time, non-compliance distribution, and rectification effect of this audit process are used as full data and fed back to the industry knowledge graph for incremental updates, resulting in an evolved knowledge graph. The evolved knowledge graph is then used to review the integrated on-site audit path and obtain path optimization constraints for dynamic optimization of subsequent audit paths.

[0006] Secondly, a cross-system integrated audit path optimization system based on industry knowledge graphs includes: The graph construction module is used to collect multi-source heterogeneous data from the target industry, extract knowledge and resolve conflicts, and build an industry knowledge graph that covers cross-system standard associations and production process flows. The document library module is used to perform vector embedding and semantic matching on the documents to be reviewed of the target company based on industry knowledge graphs. By identifying and removing duplicate and conflicting content, an integrated review document library is obtained. The path planning module is used to extract cross-system clause association paths and historical audit data from the industry knowledge graph based on the integrated audit document library to obtain the preliminary audit path; based on the preliminary audit path, key audit areas are delineated and spatially discretized to obtain audit feature data; based on the audit feature data, the path cost correction coefficient is calculated, and the dwell time and inspection order of the preliminary audit path are optimized and adjusted to obtain the integrated on-site audit path; The non-compliance processing module is used to input non-compliance information into the industry knowledge graph for correlation and source tracing based on the integrated on-site audit path, obtain related non-compliance clusters, and perform similarity matching based on the historical rectification case library to obtain the root cause rectification plan; The feedback optimization module is used to feed the root cause rectification plan and the path time, non-conformity distribution and rectification effect of this audit process as full data to the industry knowledge graph for incremental updates, resulting in an evolved knowledge graph. The evolved knowledge graph is used to review the integrated on-site audit path to obtain path optimization constraints, which are used for dynamic optimization of the audit path in the future.

[0007] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0008] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0009] The above-described solution of the present invention has at least the following beneficial effects: This process involves collecting heterogeneous data from multiple sources within the target industry, extracting knowledge, resolving conflicts, and constructing an industry knowledge graph. It achieves standardized aggregation of multi-source data, eliminates redundancy and conflicts, establishes cross-system and production process connections, and forms a structured knowledge system to provide reliable knowledge support for subsequent audits. Based on the industry knowledge graph, it performs vector embedding and semantic matching on enterprise audit materials, eliminating duplicate and conflicting content to obtain an integrated audit document library. It standardizes the format of audit materials, improving regularity and usability, and providing a standardized data foundation for audit path generation. Based on the integrated audit document library, it extracts relevant data to generate preliminary audit paths, which are then optimized through regional discretization and cost correction to obtain an integrated on-site audit path. This process enables the generation and optimization of audit paths, improving the efficiency and standardization of on-site audits. It inputs non-conformity information into the knowledge graph for source tracing to form association clusters, and combines this with historical case matching to obtain root cause rectification solutions. It achieves non-conformity classification and source tracing and case reuse, improving the targeting and standardization of rectification. It feeds back all audit data to the knowledge graph for incremental updates, and optimizes the constraints of the review path. This process achieves closed-loop data reuse and knowledge graph evolution, providing a basis for dynamic optimization of subsequent audit paths and leveraging the continuous value of data. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating the cross-system integrated audit path optimization method based on industry knowledge graphs provided in an embodiment of the present invention.

[0011] Figure 2 This is a schematic diagram of a cross-system integrated audit path optimization system based on industry knowledge graph, provided by an embodiment of the present invention. Detailed Implementation

[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0013] like Figure 1 As shown, embodiments of the present invention propose a cross-system integrated audit path optimization method based on industry knowledge graphs, the method comprising the following steps: Step 100: Collect multi-source heterogeneous data from the target industry, extract knowledge and resolve conflicts, and construct an industry knowledge graph covering cross-system standard associations and production process flows. Step 200: Based on the industry knowledge graph, vector embedding and semantic matching are performed on the materials to be reviewed of the target enterprise. By identifying and removing duplicate and conflicting content, an integrated review document library is obtained. Step 300: Based on the integrated audit document library, extract cross-system clause association paths and historical audit data from the industry knowledge graph to obtain a preliminary audit path; based on the preliminary audit path, delineate key audit areas and perform spatial discretization processing to obtain audit feature data; based on the audit feature data, calculate the path cost correction coefficient, and optimize and adjust the dwell time and inspection order of the preliminary audit path to obtain the integrated on-site audit path; Step 400: Based on the integrated on-site audit path, input the non-conformity information into the industry knowledge graph for correlation and source tracing to obtain the cluster of related non-conformities, and perform similarity matching based on the historical rectification case library to obtain the root cause rectification plan; Step 500: The root cause rectification plan and the path time, non-conformity distribution, and rectification effect of this audit process are used as full data and fed back to the industry knowledge graph for incremental updates to obtain an evolved knowledge graph. The evolved knowledge graph is then used to review the integrated on-site audit path to obtain path optimization constraints, which are used for dynamic optimization of the audit path in the future.

[0014] In this embodiment of the invention, unified structured processing of multi-source heterogeneous data is achieved, knowledge conflict resolution and standard association modeling are completed, and a unified knowledge representation covering system standards and process flows is formed, improving the efficiency of cross-system knowledge organization and retrieval. Automatic deduplication and conflict verification of audit materials are achieved through vector embedding and semantic matching, forming a standardized integrated audit document library, reducing data redundancy and improving the consistency and usability of audit documents. Quantitative modeling and discretization of audit paths are completed based on knowledge graphs, and adaptive optimization of dwell time and inspection order is achieved through path cost correction coefficients, improving the rationality and execution efficiency of on-site audit paths. The association tracing and clustering analysis of non-conformities are completed based on knowledge graphs, and accurate matching of root cause rectification solutions is achieved by combining historical rectification cases, improving the systematic nature of problem location and rectification governance. Incremental iterative updates of the knowledge graph are achieved through full audit data, and dynamic optimization constraints are formed based on review results, supporting the continuous evolution of subsequent audit paths and improving the adaptability and long-term stability of the method.

[0015] In a preferred embodiment of the present invention, step 100 includes: Step 101: Collect multi-source heterogeneous data from the target industry. This multi-source heterogeneous data includes at least standard clause texts from multiple management systems, industry production specifications, real-time enterprise production data, and historical audit case reports. The target industry is preferably the automotive parts manufacturing industry. Specifically, this involves comprehensively collecting multi-source heterogeneous data related to cross-system audits and production processes within the automotive parts manufacturing industry. The collected multi-source heterogeneous data specifically covers ISO 9001 quality management system, ISO 14001 environmental management system, ISO 45001 occupational health and safety management system, and IATF... The document includes the complete text of all standard clauses of the 16949 automotive industry-specific quality management system; general production specifications and specific process requirements for the automotive parts manufacturing industry; real-time production data generated during the production process of automotive parts manufacturing enterprises, including raw material procurement data, precision machining parameter data, surface treatment data, finished product testing data, and warehousing and transportation data; and various cross-system audit case reports conducted in the industry in the past, including audit plans, audit records, non-conformity lists, and rectification reports. This ensures that the collected data can comprehensively cover cross-system audit standards and automotive parts production processes, providing data support that is relevant to the actual situation of the industry.

[0016] Step 102 involves preprocessing the multi-source heterogeneous data, including removing irrelevant characters, correcting formatting errors, and standardizing terminology to obtain cleaned data. Specifically, this includes: preprocessing various types of text data by removing irrelevant characters, redundant characters, special symbols, and invalid information unrelated to cross-system audit standards, industry production specifications, and enterprise production audits; correcting formatting errors in different formats by standardizing the formats of structured, semi-structured, and unstructured data to resolve inconsistencies; and standardizing terminology in various system standards, industry specifications, and enterprise documents according to the general terminology standards of the automotive parts manufacturing industry and the standardized terminology of each management system to eliminate terminology differences. The final result is cleaned data with a unified format, standardized terminology, and no redundant or invalid information, laying the foundation for subsequent knowledge extraction.

[0017] Step 103: Based on the cleaned data, named entity recognition technology is used to identify entities in the text, resulting in an entity set. Specifically, this includes: based on the cleaned data, named entity recognition technology is used to extract entities from various types of text data. For management system standard clause texts, entities such as standard clause numbers, control requirements, and responsible entities are identified; for industry production specifications, entities such as process names, control parameters, and operating requirements are identified; for enterprise real-time production data, entities such as production links, equipment names, raw material types, and testing indicators are identified; for historical audit case reports, entities such as audit items, non-conformity types, and corrective measures are identified. All identified entities are summarized and organized, duplicate entities are removed, and a complete and structured entity set is formed, providing processing objects for subsequent entity relationship extraction and attribute labeling.

[0018] Step 104 involves extracting relationships from the entities in the entity set to determine the semantic associations between entities and obtain entity-relationship triples. Specifically, this includes: using relation extraction technology to perform semantic association analysis on various types of entities in the entity set, and combining the production process logic of the automotive parts manufacturing industry with the inherent associations of cross-system audit standards to determine the semantic relationships between different entities; for example, identifying the association between the process control entity in the ISO9001 system and the production part approval procedure entity in the IATF16949 system, identifying the correspondence between production link entities and control parameter entities, and identifying the association between non-conformity entities and audit item entities; organizing each group of semantically related entities and the relationship between them into entity-relationship triples, i.e., a standardized representation of subject, relation, and object, to achieve a structured presentation of semantic associations between entities and eliminate the current situation of fragmented data from different systems.

[0019] Step 105 involves attribute labeling of entities in the entity set to supplement their descriptive information, resulting in entities containing attribute information. Specifically, this includes: performing comprehensive attribute labeling on each entity in the entity set, combining relevant information from the cleaned data to supplement detailed descriptions of each entity; for example, labeling standard clause entities with attributes such as clause effective time, scope of application, and control level; labeling production process entities with attributes such as operation procedures, required equipment, and quality requirements; and labeling non-conformity entities with attributes such as occurrence scenario, severity, and associated production processes. By improving the descriptive dimensions of each entity through attribute labeling, entities containing complete attribute information are obtained, enhancing the completeness of entity representation and providing knowledge support for subsequent knowledge unit merging and knowledge graph construction.

[0020] Step 106: Merge the entity relation triples with entities containing attribute information, and perform conflict detection and resolution. The conflict detection includes at least entity conflicts, attribute conflicts, relation conflicts, and semantic conflicts to obtain resolved knowledge units. Specifically, this includes: integrating and merging the entity relation triples with entities containing attribute information to form a preliminary knowledge unit set; performing comprehensive conflict detection on the preliminary knowledge unit set, with the detection content including at least entity conflicts, attribute conflicts, relation conflicts, and semantic conflicts. Entity conflicts refer to conflicts between entities with different names but referring to the same object in different data sources; attribute conflicts refer to conflicts between the same attribute of the same entity with different values; relation conflicts refer to conflicts between different semantic relationships between the same group of entities; and semantic conflicts refer to conflicts between inconsistent semantic expressions of the same concept in different data sources. For the various types of conflicts detected, the principle of prioritizing industry norms, adhering to standard clauses, and adapting to actual production is adopted to resolve conflicts, delete redundant conflict information, correct erroneous information, and unify inconsistent information, and finally obtain conflict-free and highly consistent resolved knowledge units.

[0021] Step 107 involves standardizing the concepts of the resolved knowledge units, unifying different expressions of the same concept from different data sources, and constructing a hierarchical conceptual structure to obtain standardized knowledge units. Specifically, this includes: standardizing the concepts of the resolved knowledge units; addressing different expressions of the same concept from different data sources; and unifying the names and semantics of concepts by combining general standards of the automotive parts manufacturing industry with standard expressions of various management systems. For example, unifying expressions such as "precision machining" and "fine machining" referring to the same process in different materials as "precision machining." Simultaneously, based on the production process logic of the automotive parts manufacturing industry and the hierarchical relationship of cross-system standards, a hierarchical conceptual structure is constructed, classifying and grading similar concepts to form a hierarchical system from general industry concepts to specific sub-concepts. This achieves standardized organization of knowledge units, resulting in standardized knowledge units and improving the standardization and scalability of the knowledge system.

[0022] Step 108: Based on standardized knowledge units, construct an industry knowledge graph covering cross-system standard associations and production process flows. Specifically, this includes: constructing the industry knowledge graph using graph structure modeling based on standardized knowledge units; determining the graph structure components of the industry knowledge graph; using various entities in the standardized knowledge units as nodes in the knowledge graph, including: entities related to cross-system standard clauses, entities related to automotive parts production process flows, entities related to historical audit cases, and entities related to enterprise production data; using the semantic relationships between entities recorded in entity relationship triples as connecting edges between nodes in the knowledge graph; simultaneously associating all attribute information contained in the entities after attribute annotation to each entity node as attribute parameters for each node, realizing structured association of entities, relationships, and attributes; further, establishing cross-system standard clause associations, including those from the four management system standards: ISO 9001, ISO 14001, ISO 45001, and IATF 16949. The knowledge graph identifies the inherent, subordinate, and complementary relationships between different aspects of the automotive parts manufacturing process, as well as the connections between various stages, including raw material procurement, precision machining, surface treatment, finished product testing, warehousing, and transportation. It determines the sequential connections, mutual influences, and collaborative control relationships between these stages and integrates these relationships into the graph structure. The visualization and structured presentation of these relationships are achieved through the correspondence between nodes and edges. During the knowledge graph construction process, the coverage is simultaneously verified to ensure that the constructed industry knowledge graph comprehensively covers all core content related to cross-system standards and the automotive parts manufacturing process, without any omissions of key knowledge, ultimately forming a unified, complete, and structured industry knowledge system. This knowledge system enables the centralized integration and efficient association of cross-system knowledge, industry production knowledge, and audit-related knowledge, providing unified knowledge support for subsequent steps such as the construction of an integrated audit document library, the optimization of integrated on-site audit paths, and the tracing of non-conformities.

[0023] In this embodiment of the invention, multi-source heterogeneous core data of the target industry are collected, with a preferred focus on the automotive parts manufacturing industry, to achieve comprehensive collection of multi-source data and provide diverse data support for subsequent knowledge graph construction. The multi-source heterogeneous data is preprocessed to obtain cleaned data, eliminating data noise and format differences, improving data quality, and laying a reliable foundation for knowledge extraction. Based on the cleaned data, named entity recognition technology is used to extract entities, obtaining an entity set and accurately extracting core entities, providing targeted objects for subsequent relation extraction and attribute annotation. Relationship extraction is performed on the entities to obtain entity relation triples, mining the inherent associations between entities, and constructing a structured relation network. The system strengthens the foundation for knowledge graph association; it annotates entities with attributes to obtain entities containing attribute information, expands entity features, improves entity information, and enhances the accuracy and usability of knowledge units; it merges related knowledge units and performs multi-type conflict detection and resolution to eliminate various data conflicts, ensuring the consistency and accuracy of knowledge units and forming standardized knowledge units; it standardizes the concepts of knowledge units and constructs a concept hierarchy, unifies concept expressions, and sorts out hierarchical relationships to improve the systematicness and standardization of knowledge units; based on standardized knowledge units, it constructs an industry knowledge graph, integrates knowledge resources, realizes cross-system and production process knowledge association, and provides core support for subsequent audits.

[0024] In a preferred embodiment of the present invention, step 200 includes: Step 201: Obtain the audit materials from the target company and perform format conversion and content extraction on the audit materials to obtain the text data to be processed. Specifically, this includes: obtaining all audit materials from the target automotive parts manufacturing company for cross-system audits. These audit materials cover various documents prepared by the company to meet the audit requirements of the four management systems: ISO 9001, ISO 14001, ISO 45001, and IATF 16949, including quality manuals, work instructions, production process documents, testing records, environmental control reports, occupational health and safety operating procedures, etc. These audit materials cover structured data, semi-structured data, and unstructured data, exhibiting diverse formats. A comprehensive format conversion process is performed on the obtained audit materials, uniformly converting the different formats into a preset standard text format. The preset standard text format is a universal plain text format compatible with structured, semi-structured, and unstructured data. The pre-set text format process involves determining the encoding format, character encoding standard, and line break rules of the pre-set text based on the characteristics of cross-system audit materials in the automotive parts manufacturing industry and the subsequent vector embedding and semantic matching processing requirements. It also references the general expression standards of audit materials from the four management systems: ISO 9001, ISO 14001, ISO 45001, and IATF 16949, and pre-sets the basic layout requirements of the pre-set text content. This ensures that the pre-set standard text format can adapt to the format conversion of various audit materials and meet the needs of subsequent data processing, resolving the problem of inconsistent audit material formats. Simultaneously, content extraction is performed on the converted materials, removing redundant content such as headers and footers, blank paragraphs, and duplicate attachments irrelevant to the audit, and extracting core text information relevant to the cross-system audit. The final result is a uniformly formatted and content-focused text data to be processed, providing a standardized processing foundation for subsequent vector embedding and semantic matching.

[0025] Step 202: Based on the industry knowledge graph, a vector embedding model is used to vectorize the text data to be processed, obtaining the embedding vectors of the audited materials. Specifically, this includes: using the automotive parts manufacturing industry knowledge graph as a foundation, calling a preset vector embedding model. The construction process of this vector embedding model involves combining the text characteristics of cross-system audits in the automotive parts manufacturing industry, customizing the construction based on the basic vector embedding framework, incorporating entity, relationship, and attribute features from the industry knowledge graph, and determining the network structure, input / output dimensions, and core parameters of the vector embedding model. The training process of the vector embedding model involves selecting a large amount of cross-system standard text, enterprise audit materials, production process documents, and historical audit case texts within the automotive parts manufacturing industry as training datasets. After cleaning, analyzing, and standardizing the training datasets, they are input into the vector embedding model for iterative training. The training error of the vector embedding model is optimized until the semantic representation accuracy of the vector embedding model reaches the preset requirements, thus completing the training of the vector embedding model. The implementation process of the vector embedding model is as follows: the trained vector embedding model is called and the trained parameters are loaded. The text data to be processed is received and semantically parsed. Then, the text data to be processed is subjected to comprehensive vectorization mapping. In the process of vectorization mapping, combined with the existing entity, relationship and attribute information in the industry knowledge graph, each core content in the text data to be processed is mapped into a fixed-dimensional embedding vector. This realizes the transformation of unstructured text data into a computable and comparable vector form, while retaining the semantic information of the text content. This ensures that the embedding vector can represent the core semantics of the data to be reviewed, providing a quantifiable comparison object for subsequent semantic matching with cross-system standard knowledge, and improving the accuracy and efficiency of semantic matching.

[0026] Step 203 involves extracting the embedding vectors of cross-system standard knowledge from the industry knowledge graph as a reference benchmark for semantic matching. Specifically, this includes extracting cross-system standard knowledge related to the four management systems—ISO 9001, ISO 14001, ISO 45001, and IATF 16949—from the industry knowledge graph. This includes the core content such as the standard clause texts, clause relationships, and specification requirements of each system. Using the same vector embedding model as in Step 202, the extracted cross-system standard knowledge is vectorized to obtain the embedding vectors of the cross-system standard knowledge. These embedding vectors are then used as the reference benchmark for subsequent semantic matching to determine the comparison standard for semantic matching. This ensures that the subsequent matching process closely aligns with the standard requirements of cross-system audits, guaranteeing the relevance and standardization of the semantic matching results.

[0027] Step 204: Calculate the similarity between the embedding vector of the data to be reviewed and the embedding vector of the cross-system standard knowledge to obtain a matching value. Then, convert the matching value into a matching label based on a preset threshold. The matching label includes at least three categories: duplication, conflict, and consistency. Specifically, this involves: calculating the similarity between the embedding vector of the data to be reviewed and the embedding vector of the cross-system standard knowledge one by one to obtain a matching value between each pair of embedding vectors. This matching value characterizes the semantic association between the content of the data to be reviewed and the content of the cross-system standard knowledge. A reasonable matching threshold range is preset. The preset threshold is a specific numerical value used to define the semantic association state between the data to be reviewed and the cross-system standard knowledge, including a preset high threshold, a preset medium threshold, and a preset low threshold. The preset high threshold is 0.85, the preset medium threshold is 0.65, and the preset low threshold is 0.35. The preset threshold setting process is as follows: To meet the actual needs of cross-system audits in the automotive parts manufacturing industry, this study references a large amount of semantic matching test data between cross-system standard texts and enterprise audit materials. It combines the semantic representation accuracy of vector embedding models with the accuracy and efficiency of redundancy removal and integration of audit materials. After multiple tests and verifications, specific values ​​for each threshold were determined. Based on the threshold range of the calculated matching degree value, each segment of audit material content was converted into a corresponding matching tag. These matching tags include at least three types: repetition, conflict, and consistency. A matching degree value of 0.85 or higher is marked as consistent; a matching degree value between 0.65 and 0.85 with semantic overlap is marked as repetition; and a matching degree value below 0.35 with contradictory semantics is marked as conflict. These matching tags clearly define the association between the audit materials and cross-system standard knowledge, providing data support for subsequent redundancy removal and integration work.

[0028] Step 205: Based on the matching tags, the content marked as duplicates and conflicts is eliminated and integrated according to preset merging rules and priority strategies to obtain the deredundant audit document content. Specifically, this includes: based on the matching tags, systematically eliminating and integrating the audit materials marked as duplicates and conflicts according to preset merging rules and priority strategies; the preset merging rules are rules used to determine the criteria for selecting duplicate content in the audit materials. The core content is to prioritize retaining content that is highly consistent with the automotive parts manufacturing process and meets the latest system standards, while eliminating redundant and repetitive statements. The process of setting these merging rules involves combining the actual needs of cross-system audits in the automotive parts manufacturing industry, referring to the characteristics of audit materials from a large number of automotive parts companies, the updates to cross-system standards, and production process specifications, while balancing the practicality and standardization of the audit materials; and combining historical experience data on handling duplicate content in audits, after multiple tests and verifications. The specific content of the merger rules is defined; the preset priority strategy is a strategy used to determine the principles for handling conflicting content in the audit materials. Its core content is to prioritize the normative requirements in cross-system standard knowledge, combine them with the actual production situation of the target enterprise, reasonably integrate the conflicting content, correct contradictory statements, and retain the core information that conforms to industry norms and the actual production of the enterprise. The preset process of this priority strategy is as follows: combining the core requirements of cross-system audits in the automotive parts manufacturing industry, referring to the standard priority of the four management systems ISO9001, ISO14001, ISO45001 and IATF16949, combining the actual production characteristics of the target enterprise and the experience in handling conflicting content in historical audits, taking into account the compliance of the audit and the feasibility of the enterprise's production, and determining the specific content of the priority strategy after multiple tests and verifications, finally obtaining the deredundant audit document content that is free of duplication and conflict and has standardized content, thus solving the problems of redundancy and conflicting statements in the audit materials.

[0029] Step 206 involves structuring and standardizing the format of the deredundant audit documents to obtain an integrated audit document library. This includes: comprehensively structuring the deredundant audit document content, dividing it according to the core categories of cross-system audits into standard clause content, production process audit content, environmental control audit content, occupational health and safety audit content, and testing record-related content, determining the hierarchical relationships and content boundaries of each part; simultaneously, standardizing the format of the structured audit document content, unifying the font, line spacing, heading levels, and numbering rules to ensure the entire audit document library is formatted correctly, neatly, and orderly, ultimately forming an integrated audit document library. This achieves centralized integration and standardized management of audit documents, improves the efficiency of audit document retrieval, and provides efficient and unified data support for the subsequent construction of an integrated on-site audit path.

[0030] In this embodiment of the invention, target enterprise's pending review materials are obtained, and their format is converted and content extracted to obtain text data to be processed. This achieves preliminary organization and core content extraction of the pending review materials, eliminating format barriers and providing a unified and usable text data foundation for subsequent vectorization processing. Based on an industry knowledge graph, a vector embedding model is used to vectorize the processed text data, obtaining the embedded vectors of the pending review materials. This transforms unstructured text data into a computable vector form, achieving quantitative representation of the pending review materials and providing data support for subsequent semantic matching. Embedded vectors of cross-system standard knowledge are extracted from the industry knowledge graph as semantic matching reference benchmarks, determining the core reference basis for semantic matching, ensuring that the matching process conforms to cross-system standard requirements, and improving the accuracy of the matching. The process involves: 1) Calculating the similarity between the materials to be reviewed and the cross-system standard knowledge embedding vectors to obtain a matching score. This score is then converted into matching tags based on a preset threshold, enabling precise comparison between the materials to be reviewed and the standard knowledge, clearly defining the content matching status, and providing a basis for subsequent redundancy removal and integration. Based on the matching tags, duplicate and conflicting content is removed and integrated according to preset merging rules and priority strategies to obtain the deredundant review document content. This achieves redundancy optimization of the materials to be reviewed, eliminating content redundancy and conflicts, and improving the accuracy and conciseness of the document content. Finally, the deredundant review document content is structured and formatted to obtain an integrated review document library, achieving standardized organization of review documents and forming a unified and orderly document system, providing standardized data support for the generation of subsequent review paths.

[0031] In a preferred embodiment of the present invention, step 300 includes: Step 301: Based on the integrated audit document library, extract cross-system clause association path data from the industry knowledge graph. This association path data includes: nodes and their corresponding audit clauses, production processes, and the association edges between nodes. Each association edge records the two nodes connected and their logical relationship type. Specifically, this involves using the integrated audit document library as the data foundation and relying on the automotive parts manufacturing industry knowledge graph to extract cross-system clause association path data. This association path data specifically includes various audit nodes, each corresponding to ISO9001, ISO14001, ISO45001, and IATF1694. The document outlines the specific audit clauses within the four management systems, as well as the corresponding production stages in the automotive parts manufacturing process. It also includes the connecting edges between each audit node, with each edge clearly recording the specific information of the two connected audit nodes and the logical relationship type between them, including subordinate, related, and sequential relationships. The document counts the number of nodes corresponding to each system's audit clause, identifies the production stage number corresponding to each node, and records the node number and logical relationship type connected by each connecting edge, forming a structured association path data table. This enables the structured extraction of cross-system clause relationships, eliminating the fragmented nature of the clauses across different systems.

[0032] Step 302: Extract historical review data from the industry knowledge graph. This historical review data includes at least the historical average time spent on each review node and the historical non-compliance rate. Specifically, this includes: extracting historical review data related to the review nodes from the industry knowledge graph. This historical review data covers at least two types of core data: first, the historical average time spent on each review node in past review processes; and second, the historical non-compliance rate of each review node. The calculation process for the historical average time is as follows: collect the specific time used for each review of the review node in the past three years, add up the times of all single reviews to obtain the total time, and divide the total time by the number of reviews to obtain the historical average time. The average time taken to review each node; the historical non-compliance rate is calculated by counting the total number of reviews for that node in the past three years, and simultaneously counting the number of non-compliance items that occurred during those reviews. The historical non-compliance rate is obtained by dividing the number of non-compliance items by the total number of reviews, which is the ratio of the number of non-compliance items that occurred during past reviews to the total number of reviews. The extracted historical review data is organized and standardized, outlier data is removed, and missing data is added to ensure the completeness and accuracy of the data. This provides real and effective historical data support for the subsequent assignment of path cost weights and the optimization of review paths, solving the problem of the inability to effectively reuse historical review data.

[0033] Step 303: Based on the nodes and their relationships in the cross-system clause association path data, and combining the historical average time spent and non-compliance occurrence rate of each node, assign a comprehensive cost weight to the association edge corresponding to each node. The comprehensive cost weight is calculated by weighting the historical average time spent and non-compliance occurrence rate of the node. Specifically, this includes: taking the audit nodes and their relationships in the cross-system clause association path data as the core, and combining the historical average time spent and historical non-compliance occurrence rate of each audit node, assigning a comprehensive cost weight to the association edge corresponding to each audit node; the comprehensive cost weight is calculated by weighting the historical average time spent and historical non-compliance occurrence rate of the node. The occurrence rates are calculated by weighting the factors, where the weighting coefficients are set based on the actual needs of cross-system audits of automotive parts and historical audit experience. The weighting coefficient corresponding to the historical average time is set to 0.6, and the weighting coefficient corresponding to the historical non-conformity occurrence rate is set to 0.4. The specific calculation process is as follows: the historical average time of the node is multiplied by 0.6 to obtain the first weighting value, and the historical non-conformity occurrence rate of the node is multiplied by 0.4 to obtain the second weighting value. The first weighting value and the second weighting value are added together to obtain the comprehensive cost weight of the associated edge of the node. This ensures that the comprehensive cost weight can reflect the audit cost corresponding to the associated edge, thereby achieving a quantitative representation of the audit cost.

[0034] Step 304: Using the audit start node as the search starting point and the audit end node as the target, traverse all reachable paths and calculate the sum of the comprehensive cost weights of each edge on each path as the total cost of the path. Specifically, this includes: determining the start and end nodes of the cross-system audit of automotive parts, where the start node is set as the audit node corresponding to the enterprise's production entry point and the end node is set as the audit node corresponding to the finished product inspection qualification; using the audit start node as the search starting point and the audit end node as the target, use a path search algorithm to traverse all reachable audit paths from the start node to the end node, and calculate the sum of the comprehensive cost weights of each associated edge on each reachable path; the specific calculation process is as follows: extract the comprehensive cost weight of each associated edge on the path in sequence, starting from the weight of the first associated edge, and add the weight of each subsequent associated edge in sequence, continuing to accumulate until the weights of all associated edges on the path are accumulated, and use the sum obtained as the total cost of the audit path. Perform the same accumulation calculation operation on each reachable path to achieve a quantitative evaluation of the cost of each reachable path.

[0035] Step 305: Select the path with the minimum total cost as the preliminary review path. The preliminary review path consists of a sequence of review nodes arranged in order. Specifically, it includes: comparing the total cost of all reachable paths one by one. The comparison process is as follows: extract the total cost of the first path as the benchmark value, compare the total cost of the second path with the benchmark value, and retain the path with the smaller total cost as the new benchmark value. Then compare the total cost of the third path with the new benchmark value and continue to retain the path with the smaller total cost, and so on, until the comparison of the total cost of all reachable paths is completed, and the path with the minimum total cost is selected as the preliminary review path. The preliminary review path consists of a sequence of review nodes arranged in order. Each node corresponds to a clear review clause and production link. The arrangement order of the node sequence conforms to the automotive parts production process and cross-system review logic. The number of nodes in the preliminary review path and the weight of the associated edges corresponding to each node are counted to achieve preliminary optimization of the review path, taking into account both review efficiency and cost control.

[0036] Step 306: Based on the key process nodes of the enterprise's production site corresponding to the audit nodes involved in the preliminary audit path, map the location information of each node in physical space. Specifically, this includes: associating all audit nodes involved in the preliminary audit path with the key process nodes of the automotive parts enterprise's production site, where each audit node corresponds to a specific process node in the production site; using a preset location mapping rule, which is a rule used to map audit nodes to the physical space location information of their corresponding process nodes in the production site, the core content of which is to pre-establish a correspondence table between process nodes in the production site and physical space coordinates, and by querying the process nodes corresponding to the audit nodes, extracting the spatial coordinates of the process nodes based on the correspondence table and determining their production area; the preset process of this location mapping rule is to combine the actual layout of the automotive parts enterprise's production site and sort out the production... The specific physical locations of all process flow nodes on-site are recorded, including the spatial coordinates and production area of ​​each node, forming a correspondence table between process flow nodes and their physical locations. Simultaneously, based on the association between audit nodes and process flow nodes, the query, extraction, and recording procedures are determined. After multiple tests to ensure mapping accuracy, preset location mapping rules are established, mapping each audit node to the specific physical location information of its corresponding production site process flow node, including spatial coordinates and its production area. Specifically, the mapping process involves querying the corresponding process flow node for each audit node according to the preset location mapping rules, extracting the spatial coordinates of that process flow node from the correspondence table, determining its production area, and recording the spatial coordinates and production area information for each audit node. This achieves a precise correspondence between audit nodes and the physical location of the enterprise's production site, providing spatial data support for the delineation of key audit areas.

[0037] Step 307: Based on the position information of each node in physical space, perform convex hull calculation, select the set of vertices that constitute the boundary of the convex hull and connect them in counterclockwise order to form the smallest convex polygon containing all nodes. The area enclosed by the smallest convex polygon is taken as the spatial range of the key review area. Specifically, this includes: collecting the position information of each review node in physical space and extracting the spatial coordinate data of each review node; performing convex hull calculation based on these spatial coordinate data. The specific calculation process is as follows: selecting the nodes with the smallest and largest x-coordinates and the smallest and largest y-coordinates among all spatial coordinates as the initial convex hull vertices, and then sequentially judging each of the remaining nodes... If a node is inside the polygon formed by the initial convex hull vertices, it is added to the convex hull vertex set. At the same time, the original convex hull vertices that are obscured by the node are removed. This judgment process is repeated until all nodes have been judged. The convex hull algorithm is used to select the set of vertices that can form the boundary of the convex hull. These vertices are connected in counterclockwise order to form the smallest convex polygon that can contain all the review nodes. The length of each side and the coordinates of each vertex of the smallest convex polygon are measured. The spatial area enclosed by the smallest convex polygon is determined as the spatial range of the key review area, delineating the core review area and reducing the invalid review range.

[0038] Step 308: Based on the spatial range of the key audit area, set the granularity parameters for grid division; perform gridded spatial discretization processing on the key audit area according to the granularity parameters, dividing the spatial range into multiple sub-region units; record the spatial coordinate range and boundary information of each sub-region unit to obtain structured spatial discretized grid data. Specifically, this includes: based on the spatial range of the key audit area, measuring the length and width of the area, and combining the actual layout of the automotive parts production site, audit accuracy requirements, and audit efficiency needs, setting the granularity parameters for grid division; the granularity parameters are the size and number of grids, and the specific setting process is to determine the grid edges based on the length and width of the key audit area. The number of horizontal grids is obtained by dividing the length of the key review area by the grid side length, and the number of vertical grids is obtained by dividing the width of the key review area by the grid side length. The grid side length is set to 1 meter. The number of horizontal and vertical grids is calculated and determined according to the actual area size. The key review area is spatially discretized into grids according to the set granularity parameters. The entire key review area is uniformly divided into multiple sub-region units of uniform size. The spatial coordinate range, boundary coordinates, and area of ​​each sub-region unit are measured one by one. The number of each sub-region unit and its corresponding boundary information are recorded. Finally, structured spatial discretized grid data is obtained, realizing the structured organization of the spatial data of the key review area.

[0039] Step 309: Based on structured spatial discretized grid data, extract historical audit data corresponding to each sub-region unit from the industry knowledge graph. Statistically analyze the distribution density of historical non-compliance items, the concentration of audit points, and the pre-defined risk level within each sub-region unit to obtain audit characteristic data for each sub-region. Specifically, this includes: using structured spatial discretized grid data as a basis, extracting historical audit data corresponding to each sub-region unit from the industry knowledge graph, and performing statistical analysis on the historical audit data within each sub-region unit. The specific statistical process involves: counting the number of historical non-compliance items in the sub-region unit over the past three years; measuring the area of ​​the sub-region unit; dividing the number of non-compliance items by the area of ​​the sub-region to obtain the distribution density of historical non-compliance items within the sub-region unit, i.e., the ratio of the number of non-compliance items to the area of ​​the sub-region; counting the number of audit points contained within the sub-region unit; recording the specific location of each audit point; determining the distribution density of audit points to obtain the concentration of audit points, i.e., the number and distribution of audit points contained within the region. Furthermore, based on the non-conformity distribution density and audit point concentration of the sub-region, and combined with historical audit rectification data, the risk level of the sub-region is pre-defined. The pre-determining process for the risk level involves, in conjunction with the actual needs of cross-system audits of automotive parts, referring to a large amount of historical audit data, determining the high, medium, and low standards for non-conformity distribution density and the standards for audit point concentration, and combining the difficulty, rectification cycle, and rectification effect of historical audits to determine the risk level classification hierarchy and the corresponding judgment conditions for each level. The risk level is divided into three levels: Level 1 is defined as high non-conformity distribution density and concentrated audit points; Level 2 is defined as medium non-conformity distribution density and relatively concentrated audit points; and Level 3 is defined as low non-conformity distribution density and dispersed audit points. After multiple tests and verifications to ensure that the risk level labeling fits the actual audit scenario and that the judgment standards are clear and executable, the pre-determining of the risk level is completed. The statistically obtained information is integrated into the audit feature data of each sub-region, providing comprehensive feature data support for the subsequent calculation of the path cost correction coefficient.

[0040] Step 310: Based on the historical audit data of each sub-region, calculate the weights of each indicator to determine the weight coefficient corresponding to each indicator, thus obtaining the weight coefficients based on historical data. Specifically, this includes: using the entropy weight method to calculate the weights of each indicator in the audit feature data based on the historical audit data of each sub-region. The audit feature indicators include: non-conformity distribution density, audit point concentration, and risk level. The specific calculation process is as follows: collect the specific values ​​of each indicator in all sub-regions, standardize the values ​​of each indicator, calculate the proportion of the value of that indicator in each sub-region to the total value of that indicator in all sub-regions; then calculate the information entropy of each indicator based on this proportion, subtract the information entropy of the indicator from 1 to obtain the difference coefficient of the indicator, add the difference coefficients of all indicators to obtain the sum of the difference coefficients, and divide the difference coefficient of each indicator by the sum of the difference coefficients to obtain the weight coefficient corresponding to the indicator. The smaller the information entropy, the greater the weight of the indicator. The weight calculation of the three indicators is completed in sequence to finally obtain the weight coefficients based on historical data, ensuring that the weight allocation fits the actual audit scenario and improving the accuracy of subsequent correction calculations.

[0041] Step 311: Based on the audit feature data of each sub-region, normalize each indicator in the audit feature data to obtain the dimensionless normalized value of each indicator. Specifically, this includes: normalizing each indicator in the audit feature data based on the audit feature data of each sub-region; using a preset normalization method, the preset process of which combines the type characteristics of each indicator in the audit feature data, refers to the industry standard for indicator processing in cross-system audits of automotive parts, and takes into account the convenience and accuracy of subsequent weighted summation calculations, and formulates corresponding normalization methods for different types of indicators; for continuous numerical indicators such as non-conformity distribution density and audit point concentration, extreme value normalization method is adopted, and the operation process of calculating the difference through the maximum and minimum values ​​to obtain the normalized value is determined; for discrete level indicators such as risk level, level calibration normalization method is adopted, and the level is set according to the actual impact of the risk level. The dimensionless values ​​corresponding to the risk levels were verified through multiple tests to ensure that the normalization process eliminated dimensional differences and achieved uniform comparability for all indicators. After confirming that the normalization method met the requirements of subsequent data processing, a pre-defined normalization method was determined. Specifically, for the indicators of non-compliance distribution density and audit point concentration, the maximum and minimum values ​​of these indicators for all sub-regions were collected. The minimum value of each indicator in each sub-region was subtracted from its value to obtain the difference. This difference was then divided by the difference between the maximum and minimum values ​​to obtain the dimensionless normalized value of the indicator. For the risk level indicators, Level 1 risk was assigned a value of 1, Level 2 risk 0.5, and Level 3 risk 0.2. The risk levels were directly converted into their corresponding dimensionless normalized values. Through this process, the values ​​of all indicators were transformed into a pre-defined uniform value range, eliminating dimensional differences between different indicators and ensuring uniform comparability. This laid a standardized data foundation for the subsequent weighted summation calculation of all indicators.

[0042] Step 312: Based on the weighting coefficients calibrated from historical data, the dimensionless normalized values ​​of each indicator are weighted and summed to calculate the path cost correction coefficient corresponding to each sub-regional unit. Specifically, this includes: using the weighting coefficients calibrated from historical data as a basis, the dimensionless normalized values ​​of each indicator are weighted and summed. The specific calculation process is as follows: first, the dimensionless normalized values ​​of each indicator in each sub-regional unit and the corresponding weighting coefficients are extracted; the normalized value of the non-compliance distribution density is multiplied by its corresponding weighting coefficient to obtain the weighted value of the non-compliance distribution density; the normalized value of the concentration of audit points is multiplied by its corresponding weighting coefficient to obtain the weighted value of the concentration of audit points; the normalized value of the risk level is multiplied by its corresponding weighting coefficient to obtain the weighted value of the risk level; the weighted values ​​of the three indicators are added together to obtain the path cost correction coefficient corresponding to each sub-regional unit. The same calculation operation is performed on each sub-regional unit to achieve a quantitative representation of the path cost correction coefficient, providing a basis for the optimization and adjustment of the preliminary audit path.

[0043] Step 313: Based on the path cost correction coefficient corresponding to each sub-region unit and the sequence information of the audit nodes in the preliminary audit path, determine the sub-region unit to which each audit node belongs and establish the correspondence between nodes and sub-regions. Specifically, this includes: combining the path cost correction coefficient corresponding to each sub-region unit and the sequence information of the audit nodes in the preliminary audit path, determining the sub-region unit to which each audit node belongs through spatial coordinate matching; the specific matching process is as follows: extract the physical spatial coordinates of each audit node, and simultaneously extract the spatial coordinate range of each sub-region unit, and determine whether the physical spatial coordinates of each audit node are within the spatial coordinate range of a certain sub-region unit; if they are, determine that the audit node belongs to that sub-region unit; if not, recheck the coordinate information until each audit node finds a corresponding sub-region unit, thereby establishing a one-to-one correspondence between audit nodes and sub-region units; record the sub-region unit number and path cost correction coefficient corresponding to each audit node, realizing the association and binding of node information with spatial features and path cost correction coefficients, providing clear data association support for subsequent audit path adjustments.

[0044] Step 314: Based on the correspondence between nodes and sub-regions, dynamically adjust the dwell time of corresponding nodes in the preliminary review path according to the path cost correction coefficient of the sub-region to which each node belongs, to obtain the adjusted dwell time of each node; simultaneously, according to the relationship between the path cost correction coefficients of the sub-regions to which adjacent nodes belong, locally rearrange the inspection order of adjacent nodes in the preliminary review path to obtain the adjusted review node sequence. Specifically, this includes: obtaining the path cost correction coefficient of the sub-region to which each review node belongs based on the correspondence between nodes and sub-regions; dynamically adjusting the dwell time of corresponding nodes in the preliminary review path according to the magnitude of the path cost correction coefficient. The specific adjustment process is as follows: the basic dwell time of each review node is preset to 15 minutes, and the basic dwell time is multiplied by the sub-region to which the node belongs. The path cost correction coefficient is used to obtain the adjusted dwell time of the node. The larger the path cost correction coefficient, the longer the dwell time of the node; the smaller the correction coefficient, the shorter the dwell time. The same adjustment operation is performed on each review node to obtain the adjusted dwell time of each node. At the same time, based on the relationship between the path cost correction coefficients of the sub-regions to which two adjacent review nodes belong, the inspection order of adjacent nodes in the preliminary review path is locally rearranged. The specific rearrangement process is as follows: the path cost correction coefficients of the sub-regions to which two adjacent nodes belong are compared, and the nodes corresponding to the sub-regions with larger correction coefficients are checked first; if the path cost correction coefficients are equal, they are arranged according to the original order in the preliminary review path. The same comparison and rearrangement operation is performed on adjacent nodes one by one to obtain the adjusted review node sequence, thereby optimizing the preliminary review path.

[0045] Step 315: Integrate the adjusted audit node sequence with the corresponding adjusted dwell time to form an integrated on-site audit path. Specifically, this includes: integrating the adjusted audit node sequence with the corresponding adjusted dwell time. The integration process involves matching the adjusted dwell time to each audit node according to the adjusted audit node sequence order, recording the audit sequence number, audit node name, and adjusted dwell time for each node, and determining the audit sequence and time used for each node. Simultaneously, supplement the audit clause number, production process name, and physical coordinate information corresponding to each node; calculate the total number of integrated audit nodes and the total dwell time; verify the completeness of information for each node to ensure no information omissions or logical contradictions, forming a complete and standardized integrated on-site audit path. This achieves the standardization and integration of the audit path, providing on-site auditors with efficient and accurate execution basis and improving audit efficiency.

[0046] In this embodiment of the invention, cross-system clause association path data is extracted to determine audit nodes, corresponding clauses, and production stages, and to clarify the logical relationships between nodes, providing structured and traceable association data support for the construction of the initial audit path; key historical audit data is extracted, and the historical time consumption and non-compliance occurrence rate of each audit node are integrated to provide real and effective data reference for subsequent path cost weighting and path optimization; combined with the historical data of nodes, a comprehensive cost weight is assigned to the associated edges to achieve a quantitative representation of the cost weight, making the path cost calculation more targeted and improving the rationality of path selection; all reachable paths are traversed and the total cost is calculated to realize the path A comprehensive quantitative assessment of path costs provides a quantitative basis for selecting the initial audit path; the path with the minimum total cost is selected as the initial audit path to achieve preliminary optimization of the audit path, balancing audit efficiency and cost control, and forming an orderly sequence of audit nodes; audit nodes are mapped to physical spatial location information to achieve precise correspondence between audit nodes and the enterprise's production site, providing spatial data support for the delineation of key audit areas; the spatial range of key audit areas is delineated through convex hull calculation, accurately delineating the core audit area, reducing invalid audit areas, and improving the rationality of audit area division; the key audit areas are discretized into a grid to divide the spatial range. To standardize sub-regional units, achieve structured organization of spatial data, and improve the processability of spatial data; extract historical audit data corresponding to sub-regions and statistically analyze audit features to achieve accurate quantification of audit features, providing feature data support for subsequent path cost correction; calculate the weight coefficients of each indicator, calibrate the weights based on historical data, ensure that the weight allocation fits the actual audit scenario, and improve the accuracy of subsequent correction calculations; normalize the audit feature indicators to eliminate the differences in the dimensions of different indicators, achieve unified comparability of indicators, and lay a standardized data foundation for weighted summation calculation; calculate the path cost correction coefficient of sub-regions through weighted summation, and realize... The quantitative representation of the path cost correction coefficient provides a quantitative basis for optimizing and adjusting the preliminary review path; the establishment of the correspondence between review nodes and sub-regions enables the association and binding of node information with spatial characteristics and path cost correction coefficients, providing clear data association support for subsequent path adjustments; the dynamic adjustment of node dwell time and partial rearrangement of inspection order based on the path cost correction coefficient achieves refined optimization of the preliminary review path, balancing review targeting and efficiency; the integration of the adjusted node sequence and dwell time forms an integrated on-site review path, achieving standardization and integrated integration of the review path, providing an efficient and accurate execution basis for on-site review.

[0047] In a preferred embodiment of the present invention, step 400 includes: Step 401: During the implementation of the integrated on-site audit path, collect information on identified nonconformities. This information includes at least a description of the nonconformity, its location, and the relevant standard clauses. Specifically, during the implementation of the integrated on-site audit path, conduct on-site audits according to the adjusted audit node sequence and dwell time, collecting all identified nonconformity information in real time. This information includes at least: a description of the nonconformity, its location, and the relevant standard clauses. The description of the nonconformity must record the specific manifestations, occurrence scenarios, and related details of the nonconformity. The location must accurately correspond to the specific sub-area unit and spatial coordinates of the production site. The relevant standard clauses must specify the specific clause number and content of the four management systems: ISO9001, ISO14001, ISO45001, and IATF16949. In this data collection process, a conical lateral surface area algorithm is introduced. This algorithm calculates the area of ​​the unfolded lateral surface of a cone and combines this with the dimensional parameters of conical equipment on the production site to assist in verification. The algorithm for accurately determining the relative position of non-conformities to conical equipment is based on calculating the actual lateral surface area of ​​the cone using its base radius and generatrix length. This area is then used to define the coverage range of the conical equipment, thus determining whether the collected non-conformity location falls within the equipment's influence range and assisting in verifying the accuracy of the location coordinates. A specific application example is a conical dust collector in a production site. The base radius and generatrix length of the dust collector are measured beforehand, and its lateral surface area is calculated using the conical lateral surface area algorithm. Combined with the spatial coordinates of the equipment's installation location, the influence range boundary of the dust collector is defined based on the lateral surface area, identifying the area around the equipment prone to dust-related non-conformities. The spatial coordinates of the dust-related non-conformity location are collected, and combined with the equipment's influence range calculated using the conical lateral surface area algorithm, it is determined whether the non-conformity location is within the dust collector's influence range. If it is, the correlation between the non-conformity and the equipment is further confirmed; if not, the auditors are prompted to verify the location coordinates or investigate other influencing factors.

[0048] The collected non-conformity information was initially organized, invalid information was removed, missing information was supplemented, the total number of non-conformities was counted, and the completeness of the three core pieces of information for each non-conformity was verified. The information completeness rate was calculated by dividing the number of complete non-conformities by the total number of non-conformities collected, ensuring that the information completeness rate reached 100%. At the same time, the verification results of the cone lateral surface area algorithm were combined to confirm the accuracy of the association between the location of the non-conformity and the relevant conical equipment and the precision of the coordinates. This ensured the completeness and accuracy of the non-conformity information, providing a complete and reliable data foundation for subsequent non-conformity tracing and the generation of root cause rectification plans.

[0049] Step 402 involves inputting the non-compliance information into the industry knowledge graph. Starting with the key entities in the non-compliance description text, a correlation search is performed to obtain a set of related entities. Specifically, this includes: inputting all collected and organized non-compliance information into the automotive parts manufacturing industry knowledge graph; performing semantic parsing on the non-compliance description text to break down keywords and extract key entities; key entities include: the type of automotive parts involved in the non-compliance, production process, equipment name, control requirements, etc.; counting the number of extracted key entities; classifying and labeling each key entity; and using the extracted key entity as a starting point, performing a multi-dimensional correlation search in the industry knowledge graph, traversing all entities and relationships related to that key entity; setting the search depth to 3 levels to ensure a comprehensive and non-redundant search scope; collecting all related entity information; organizing it into a set of related entities; and counting the number of entities in the set of related entities; simultaneously recording the relationship type between each related entity and the key entity, achieving accurate association between non-compliance and related entities in the industry knowledge graph, providing structured correlation data support for subsequent root cause tracing of non-compliance.

[0050] Step 403: Based on the set of associated entities, extract corresponding historical non-compliance records from the industry knowledge graph. Aggregate non-compliance items with common root cause characteristics to form a cluster of related non-compliance items. Specifically, this includes: based on the set of associated entities, extracting historical non-compliance records corresponding to each associated entity from the industry knowledge graph; the historical non-compliance records include: descriptions, locations, relevant standard clauses, and rectification status of all non-compliance items related to the associated entity in past audits; counting the total number of extracted historical non-compliance records; sorting and analyzing the extracted historical non-compliance records; and selecting those with common root cause characteristics. The non-compliance items identified share common root cause characteristics, including involvement in the same production process, equipment, control standards, or rectification root causes. The number of non-compliance items corresponding to each root cause characteristic is statistically analyzed. These non-compliance items with common root cause characteristics are then aggregated and grouped, forming a cluster of related non-compliance items. The number of these clusters is counted, and each cluster is labeled with a core root cause characteristic tag. The proportion of non-compliance items within each cluster to the total number of historical non-compliance records is calculated. This achieves classification, aggregation, and root cause focus for non-compliance items, avoiding the problem of multiple non-compliance items with the same origin being handled independently.

[0051] Step 404: Extract feature information from clusters of non-conforming items and encode the feature information into feature vectors to obtain the feature vector of the current non-conforming item cluster. Specifically, this includes: comprehensively extracting feature information for each cluster of non-conforming items; the feature information includes: common root features of non-conforming items within the cluster, the set of relevant standard clauses, the distribution of occurrence locations, and the specific manifestation type and quantity of non-conforming items; standardizing the extracted feature information, unifying the feature representation format, eliminating redundant features, and retaining core features; counting the number of core features for each cluster, determining the fixed dimension of the feature vector, ensuring the number of dimensions matches the number of core features; and then using a preset encoding rule, which is used to convert the standardized core feature information into a fixed-dimensional feature vector. The core content of this rule is to assign corresponding numerical ranges to different types of features based on their type, ensuring a one-to-one correspondence between each feature information and the vector dimension. The numerical range is uniformly set between 0 and 1, where continuous features are assigned corresponding values ​​based on the relative magnitude of the feature values, and discrete features are assigned corresponding values ​​based on the importance and relevance of the features. The coding rule's preset process involves combining the actual needs of cross-system audits of automotive parts, sorting out the types and manifestations of various core features, referring to industry-standard feature coding practices, determining the numerical allocation principles for different types of features, setting corresponding values ​​for discrete features such as common root cause features and standard clause sets based on their impact on audit rectification, and dividing intervals and assigning corresponding values ​​for continuous features such as location distribution and number of non-conformities based on the feature's value range. After multiple tests and verifications, it is ensured that the encoded feature vectors can accurately represent the core features of the cluster and can be used for subsequent similarity calculations, thus determining the preset coding rules. The standardized feature information is encoded into fixed-dimensional feature vectors, ensuring a one-to-one correspondence between each feature information and the vector dimension during the encoding process, and assigning corresponding values ​​to each feature information, with the value range set between 0 and 1 to accurately represent the core features of the cluster. The feature vectors of the current non-conformity cluster are obtained, and the number of dimensions of each feature vector is counted, transforming the unstructured cluster features into a computable and comparable vector form, laying the data foundation for subsequent similarity matching with historical rectification cases.

[0052] Step 405: Extract a historical rectification case library from the industry knowledge graph. This library contains feature vectors of historical non-conformities and corresponding root cause rectification plans. Specifically, this includes: extracting a historical rectification case library from the industry knowledge graph. This library is a structured collection of rectification cases for all non-conformities from past cross-system audits of automotive parts. It includes feature vectors of historical non-conformities, corresponding root cause rectification plans, rectification implementation processes, rectification effects, and verification results. The root cause rectification plans include: specific rectification measures, responsible parties, rectification cycles, required resources, and rectification acceptance standards. The extracted historical rectification case library is organized and standardized. The total number of historical rectification cases in the library is counted, and the cases are categorized and stored according to non-conformity type and root cause characteristics. The number of cases under each category is calculated, and the complete correspondence between the feature vectors in each case and the historical root cause rectification plans is verified. This ensures the integrity and standardization of the data in the case library, achieving structured collection and reuse of historical rectification data, and providing historical data references for generating rectification plans for current related non-conformity clusters.

[0053] Step 406 involves extracting the weight benchmark values ​​of each feature dimension from the industry knowledge graph as the first benchmark point, and simultaneously calculating the average distribution of each feature dimension from the historical rectification case database as the second benchmark point. Specifically, this includes: extracting the weight benchmark values ​​of each feature dimension from the industry knowledge graph as the first benchmark point. The feature dimensions correspond one-to-one with the dimensions of the feature vector. The weight benchmark values ​​are pre-set based on industry standards, audit priorities, and expert experience for cross-system audits of automotive parts, used to define the importance of each feature dimension; and calculating the sum of the weight benchmark values ​​for each feature dimension to ensure the sum of the weight benchmark values ​​is 1. Simultaneously, the average distribution of each feature dimension is calculated from the historical rectification case database. As a second benchmark, the statistical process involves collecting the values ​​of each dimension of the feature vectors of all historical non-compliance items in the case library, counting the total number of values ​​for each feature dimension, summing all the values ​​for each feature dimension to obtain the sum of values ​​for that dimension, dividing the sum of values ​​by the total number of values ​​to obtain the average value for each feature dimension, calculating the difference between each value and the average value, squaring the difference, summing the results, and dividing by the total number of values ​​to obtain the variance, determining the maximum and minimum values ​​for each feature dimension, and thus obtaining the distribution interval. This determines the distribution pattern of each feature dimension in historical cases. By extracting the dual benchmarks, the reference basis for subsequent feature weight optimization is determined, ensuring that the weight optimization conforms to industry standards and historical experience.

[0054] Step 407: With the goal of minimizing the cross-entropy of the first and second reference points, iteratively optimize the weight coefficients of each feature dimension to obtain the optimized feature weights. Specifically, this includes: iteratively optimizing the weight coefficients of each feature dimension with the goal of minimizing the cross-entropy of the first and second reference points; setting the initial weight coefficients according to the weight reference values ​​of each feature dimension, summing the initial weight coefficients to ensure the sum of the initial weight coefficients is 1; calculating the cross-entropy value between the first and second reference points under the current weight coefficients during each iteration, setting an upper limit of 100 iterations, and calculating the change in the cross-entropy value after each iteration; adjusting the weights of each feature dimension according to the magnitude of the cross-entropy value. For feature dimensions that significantly impact cross-entropy, appropriately increase the weight coefficient by 5% to 10% of the current weight coefficient. For feature dimensions that have a smaller impact on cross-entropy, appropriately decrease the weight coefficient by 5% to 10% of the current weight coefficient. After adjustment, recalculate the total weight coefficients to ensure the sum is still 1. Repeat the above iterative adjustment process until the cross-entropy value reaches the preset minimum value and tends to stabilize. The stabilization standard is that the change in cross-entropy value is less than 0.001 for three consecutive iterations. Stop the iteration to obtain the optimized feature weights. Statistically calculate the weight coefficients of each feature dimension after optimization to ensure that the sum of the weight coefficients is 1, and ensure that the weight allocation of each feature dimension conforms to the actual similarity matching requirements.

[0055] Step 408: Using the optimized feature weights, calculate the weighted similarity between the feature vector of the current non-compliant item cluster and the feature vector of historical non-compliant items. Select matching candidate rectification cases based on the weighted similarity. Specifically, this includes: using the optimized feature weights, calculating the weighted similarity between the feature vector of the current non-compliant item cluster and the feature vector of each historical non-compliant item in the historical rectification case library; during the calculation, multiplying the corresponding dimension values ​​of the current feature vector and the historical feature vector, and then multiplying by the optimized weight coefficient corresponding to that dimension, accumulating the calculation results of all dimensions to obtain the weighted similarity value. The weighted similarity value is set between 0 and 1. Perform the same calculation operation on each historical non-compliant item feature vector, count the weighted similarity value between the current non-compliant item cluster and all historical cases, set the weighted similarity threshold to 0.7, filter out historical rectification cases with a weighted similarity of 0.7 or higher, count the number of candidate rectification cases, and ensure that the number of candidate cases is not less than 3, as candidate rectification cases to match the current non-compliant item cluster, to achieve accurate matching of historical cases and improve the adaptability of candidate cases.

[0056] Step 409 involves extracting rectification measures, rectification cycles, and verification methods from the candidate rectification cases, and integrating them to generate a root cause rectification plan for the current cluster of related non-compliance items. Specifically, this includes: analyzing each selected candidate rectification case, extracting core rectification information from each case, including specific rectification measures, a reasonable rectification cycle, and standardized verification methods; calculating the average rectification cycle for each candidate case; adjusting and optimizing the extracted rectification information based on the core root cause characteristics, location, relevant standard clauses, and actual production situation of the current cluster of related non-compliance items; removing rectification content that does not match the current non-compliance item cluster; supplementing rectification details that fit the current scenario; determining the responsible party for rectification, rectification implementation steps, and acceptance standards; calculating the adjusted rectification cycle to ensure it aligns with the company's production schedule; integrating the adjusted and optimized rectification measures, rectification cycles, and verification methods; counting the number of core clauses in the integrated rectification plan to ensure its completeness; and forming a complete and standardized root cause rectification plan for the current cluster of related non-compliance items. This provides an executable and complete basis for resolving the root causes of related non-compliance items, improving rectification efficiency and standardization, and avoiding redundant rectification.

[0057] In this embodiment of the invention, core information on non-conformities during the implementation of the integrated on-site audit path is collected to achieve comprehensive and accurate data collection, providing a complete data foundation for subsequent source tracing and rectification plan generation. Non-conformity information is input into an industry knowledge graph to locate related entity sets, providing structured related data support. Non-conformities originating from the same source are clustered to form related clusters, improving the targeting and systematic nature of processing. Cluster features are extracted and encoded into feature vectors to achieve quantitative representation, laying the foundation for similarity matching. A historical rectification case library is extracted to achieve historical data collection and reuse, providing a reference basis. A dual-weight benchmark is determined to ensure that weight optimization aligns with industry standards and historical experience. Feature weights are iteratively optimized to improve the rationality of similarity calculation. Candidate rectification cases are screened through weighted similarity to improve case adaptability. Rectification-related information from the cases is integrated to generate root cause rectification plans, achieving plan standardization and improving rectification efficiency and standardization.

[0058] In a preferred embodiment of the present invention, step 500 includes: Step 501: Collect the root cause rectification plan and the path time, non-conformity distribution, and rectification effect data recorded during this audit process. Standardize this data to form structured full-volume data. Specifically, this includes: comprehensively collecting the root cause rectification plan and various data recorded in real-time during the implementation of this integrated on-site audit path, including path time data such as the actual dwell time at each audit node and the total time of the entire audit path; non-conformity distribution data such as the number of non-conformities, non-conformity type distribution, and non-conformity frequency in each sub-region unit; and rectification effect data such as the rectification completion status, rectification compliance rate, and post-rectification review results after the implementation of the root cause rectification plan. Standardize the collected data, unify the data expression format, statistical standards, and terminology, remove invalid data, supplement missing data, and perform structured transformation on unstructured data to form structured full-volume data containing all audit-related data. This achieves systematic collection and standardized organization of audit-related data, solving the problem of lack of unified standards and ineffective data collection in audits, and providing a unified and well-organized data foundation for subsequent knowledge graph updates and path reviews.

[0059] Step 502 involves adding the structured full data to the industry knowledge graph and updating the statistical attributes of relevant nodes to obtain an evolved knowledge graph. Specifically, this includes adding the structured full data to the automotive parts manufacturing industry knowledge graph one by one, according to the industry knowledge graph data storage specifications. For nodes in the industry knowledge graph related to this audit, including audit nodes, sub-region unit nodes, non-conformity nodes, and rectification plan nodes, the statistical attributes of each node are updated synchronously. The updates include: path time statistics, non-conformity frequency statistics, and rectification effect statistics for each node. This ensures that the attribute information of each node in the industry knowledge graph is consistent with the actual audit situation, completing the incremental update of the industry knowledge graph and obtaining an evolved knowledge graph. This expands the data dimensions and coverage of the industry knowledge graph, maintains synchronous adaptation between the industry knowledge graph and the actual audit scenario, and solves the problems of historical audit data not being effectively reused and knowledge not being accumulated.

[0060] Step 503: Extract the actual audit data of this audit as the data to be analyzed, and the historical benchmark data of each audit node and sub-region from the evolved knowledge graph. Specifically, this includes: extracting all actual audit data generated during this audit process from the evolved knowledge graph as the data to be analyzed, including complete data such as the path time, non-compliance distribution, rectification effect, and root cause rectification plan of this audit; simultaneously, extracting historical data accumulated from multiple audits of each audit node and each sub-region unit from the evolved knowledge graph, calculating the historical average dwell time and historical time fluctuation range of each audit node, and the historical average frequency of non-compliance and historical risk level of each sub-region unit, and organizing this historical data as historical benchmark data, realizing the accurate extraction and separation of actual data and historical data, determining the core objects of data comparison, providing clear and comparable data support for subsequent data comparison analysis, and ensuring the rationality of data comparison.

[0061] Step 504: Compare the data to be analyzed with the historical benchmark data to identify key constraint nodes where the actual time consumption exceeds the threshold and high-risk clusters where the frequency of non-conformities exceeds the threshold. Specifically, this includes: comparing the data to be analyzed with the historical benchmark data one by one; for each audit node, determining the historical benchmark data for each audit node. Specifically, the historical average dwell time for production audit nodes is 15 minutes, for quality inspection nodes it is 20 minutes, for equipment control nodes it is 18 minutes, and for environmental control nodes it is [missing data]. The average historical dwell time for occupational health nodes is 10 minutes, with a maximum of 12 minutes. Preset time thresholds are uniformly set at 1.5 times the historical average dwell time for the corresponding node. Specifically, the time threshold for production audit nodes is 22.5 minutes, for quality inspection nodes it is 30 minutes, for equipment control nodes it is 27 minutes, for environmental control nodes it is 18 minutes, and for occupational health nodes it is 15 minutes. The actual dwell time for each node is compared to the corresponding threshold. If the actual dwell time exceeds the threshold, the node is identified as a critical constraint node. For each sub-region, historical baseline data was determined. Specifically, the historical average frequency of non-conformities for the stamping workshop sub-region was 2 audits per period; for the welding workshop sub-region, it was 3 audits per period; for the assembly workshop sub-region, it was 1 audit per period; for the warehousing sub-region, it was 1 audit per period; and for the dust removal equipment perimeter sub-region, it was 4 audits per period. The preset non-conformity frequency threshold was uniformly set at twice the historical average frequency for the corresponding region; that is, the threshold for the stamping workshop sub-region was... Each audit is conducted 4 times; the threshold for the welding workshop sub-area is 6 times per audit; the threshold for the assembly workshop sub-area is 2 times per audit; the threshold for the storage sub-area is 2 times per audit; and the threshold for the dust removal equipment perimeter sub-area is 8 times per audit. The frequency of non-conformities in each sub-area unit is compared with the corresponding threshold. If the frequency of non-conformities exceeds the threshold, the area is identified as a high-risk cluster. Through threshold determination, differentiated analysis and key positioning of audit data are achieved, accurately capturing weak links in the audit process and providing a clear target direction for subsequent path optimization.

[0062] Step 505: Generate path optimization constraints based on key constraint nodes and high-risk cluster areas. These constraints include: adjusting the dwell time of key constraint nodes, adjusting the inspection priority of high-risk cluster areas, and merging rules for adjacent nodes. Specifically, this involves generating path optimization constraints based on key constraint nodes and high-risk cluster areas, combined with the actual needs of cross-system audits of automotive parts, production site layout, and audit efficiency requirements. The path optimization constraints specifically include three aspects: adjusting the dwell time of key constraint nodes by calculating the difference between the actual dwell time and the historical baseline dwell time, subtracting the historical average dwell time from the actual dwell time to obtain the dwell time difference; if the dwell time difference is positive... If the time difference is negative, it means the current review took longer than the benchmark. An adjustment factor is calculated, which is the ratio of the time difference to the historical average dwell time. The historical average dwell time is multiplied by (1 + adjustment factor × 0.6) to obtain the dwell time for subsequent reviews of this node. This compensates for the current review's insufficient time while avoiding excessive time extension that could affect overall efficiency. If the time difference is negative, it means the current review took less than the benchmark. The reduction ratio is calculated by dividing the absolute value of the time difference by the historical average dwell time. If the reduction ratio is less than 20%, the historical average dwell time remains unchanged. If the reduction ratio is more than 20%, the historical average dwell time is multiplied by (1 + reduction ratio × 0.4) to prevent incomplete reviews due to excessively short review times.

[0063] The inspection priority of high-risk clusters is adjusted. The frequency of non-compliance in each high-risk cluster exceeds the threshold. The excess value is obtained by subtracting the threshold from the actual frequency of non-compliance. The excess ratio is calculated by dividing the excess value by the threshold. High-risk clusters are sorted from high to low according to the excess ratio. The audit node corresponding to the high-risk cluster with the highest excess ratio is set as the highest priority. This process is repeated to move the audit nodes corresponding to all high-risk clusters before ordinary audit nodes, ensuring that key areas are fully audited.

[0064] The adjacent node merging rule calculates the distance between two adjacent review nodes. If the distance is no more than 5 meters and the correlation between the review content of the two nodes is above 80% (calculated by dividing the number of overlapping review clauses by the total number of review clauses), and both nodes are confirmed to have no key constraints (i.e., the actual dwell time does not exceed the corresponding threshold, and the frequency of non-compliance items in the corresponding area does not exceed the corresponding threshold), then a merging rule is established to merge the two adjacent nodes into one comprehensive review node. The dwell time after merging is the sum of the historical average dwell times of the two nodes multiplied by 0.9, reducing the number of review nodes and shortening the review path. Through the above specific calculation process, the path optimization constraints are refined and standardized, transforming data comparison results into actionable path optimization criteria, improving the pertinence and operability of path optimization.

[0065] Step 506 involves storing the path optimization constraints in an evolved knowledge graph as the basis for the next preliminary review path generation and optimization adjustment. Specifically, this includes: classifying and storing the path optimization constraints according to the storage format of the evolved knowledge graph, associating them with corresponding review nodes, sub-region unit nodes, and historical review data; establishing the relationship between the constraints and the review path generation and optimization adjustment; determining the path optimization constraints as the core basis for the next preliminary review path generation and optimization adjustment; achieving structured storage and reuse of path optimization constraints; constructing a closed-loop feedback mechanism for review data; and ensuring that subsequent review path optimization can rely on the latest review data to continuously accumulate review experience and data resources.

[0066] In this embodiment of the invention, various audit-related data are collected and standardized to form structured full-volume data. This achieves standardized collection of audit data, eliminates format differences, and provides a unified data foundation for subsequent knowledge graph updates and path reviews. The structured full-volume data is added to the industry knowledge graph, and node statistical attributes are updated to obtain an evolved knowledge graph. This enables incremental updates of the knowledge graph, expands data dimensions, maintains adaptability to actual audit scenarios, and improves data integrity and timeliness. The actual audit data and historical benchmark data are extracted from the evolved knowledge graph to achieve precise separation of the two types of data and determine the comparison objects. This provides clear support for subsequent data comparison; by comparing the data to be analyzed with historical benchmark data, key constraint nodes and high-risk clusters are identified through threshold determination, enabling differentiated analysis and key positioning of audit data, and providing direction for path optimization; based on the identification results, path optimization constraints are generated, determining dwell time, inspection priority, and node merging rules, refining standardized optimization constraints, and transforming data comparison results into actionable optimization basis; the optimization constraints are stored in an evolving knowledge graph as the basis for subsequent audit path optimization, enabling constraint reuse, building a data closed-loop feedback mechanism, and leveraging the continuous value of data.

[0067] like Figure 2 As shown, embodiments of the present invention also provide a cross-system integrated audit path optimization system based on industry knowledge graphs, including: The graph construction module is used to collect multi-source heterogeneous data from the target industry, extract knowledge and resolve conflicts, and build an industry knowledge graph that covers cross-system standard associations and production process flows. The document library module is used to perform vector embedding and semantic matching on the documents to be reviewed of the target company based on industry knowledge graphs. By identifying and removing duplicate and conflicting content, an integrated review document library is obtained. The path planning module is used to extract cross-system clause association paths and historical audit data from the industry knowledge graph based on the integrated audit document library to obtain the preliminary audit path; based on the preliminary audit path, key audit areas are delineated and spatially discretized to obtain audit feature data; based on the audit feature data, the path cost correction coefficient is calculated, and the dwell time and inspection order of the preliminary audit path are optimized and adjusted to obtain the integrated on-site audit path; The non-compliance processing module is used to input non-compliance information into the industry knowledge graph for correlation and source tracing based on the integrated on-site audit path, obtain related non-compliance clusters, and perform similarity matching based on the historical rectification case library to obtain the root cause rectification plan; The feedback optimization module is used to feed the root cause rectification plan and the path time, non-conformity distribution and rectification effect of this audit process as full data to the industry knowledge graph for incremental updates, resulting in an evolved knowledge graph. The evolved knowledge graph is used to review the integrated on-site audit path to obtain path optimization constraints, which are used for dynamic optimization of the audit path in the future.

[0068] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0069] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0070] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A cross-system integrated audit path optimization method based on industry knowledge graph, characterized in that, The method includes: Step 100: Collect multi-source heterogeneous data of the target industry. The multi-source heterogeneous data includes: collecting multi-source heterogeneous data of the target industry, which includes at least standard clause texts of multiple management systems, industry production specifications, real-time production data of enterprises, and historical audit case reports; wherein the target industry is the automotive parts manufacturing industry; and perform knowledge extraction and conflict resolution to construct an industry knowledge graph covering cross-system standard associations and production process flows. Step 200: Based on the industry knowledge graph, vector embedding and semantic matching are performed on the materials to be reviewed of the target enterprise. By identifying and removing duplicate and conflicting content, an integrated review document library is obtained. Step 300: Based on the integrated audit document library, extract cross-system clause association paths and historical audit data from the industry knowledge graph to obtain a preliminary audit path; based on the preliminary audit path, delineate key audit areas and perform spatial discretization processing to obtain audit feature data; calculate the path cost correction coefficient based on the audit feature data, and optimize and adjust the dwell time and inspection order of the preliminary audit path to obtain an integrated on-site audit path; including: based on the integrated audit document library, extract cross-system clause association path data from the industry knowledge graph, wherein the association path data includes: nodes and their corresponding audit clauses, production links, and association edges between nodes, and each association edge records the two nodes connected and their logical relationship type; Historical audit data is extracted from the industry knowledge graph. The historical audit data includes at least the historical average time spent on each audit node and the historical non-compliance rate. Based on the nodes and their relationships in the cross-system clause association path data, and combined with the historical average time consumption and non-compliance occurrence rate of each node, a comprehensive cost weight is assigned to the associated edge corresponding to each node. The comprehensive cost weight is calculated by weighting the historical average time consumption and non-compliance occurrence rate of the node. Starting from the audit start node and ending at the audit end node, traverse all reachable paths and calculate the sum of the comprehensive cost weights of each edge on each path as the total cost of that path. The path with the lowest total cost is selected as the initial review path, which consists of a sequence of review nodes arranged in order. Based on the key process nodes of the enterprise's production site corresponding to the audit nodes involved in the preliminary audit path, the location information of each node in the physical space is mapped. Based on the position information of each node in the physical space, convex hull calculation is performed, the set of vertices that constitute the boundary of the convex hull is selected and connected in counterclockwise order to form the smallest convex polygon containing all nodes, and the area enclosed by the smallest convex polygon is used as the spatial range of the key review area. Based on the spatial range of the key review area, the granularity parameters of the grid division are set; the key review area is spatially discretized into a grid according to the granularity parameters, and the spatial range is divided into multiple sub-region units; the spatial coordinate range and boundary information of each sub-region unit are recorded to obtain structured spatial discretized grid data. Based on structured spatial discretized grid data, historical audit data corresponding to each sub-region unit is extracted from the industry knowledge graph. The distribution density of historical non-compliance items, the concentration of audit points, and the pre-labeled risk level within each sub-region unit are statistically analyzed to obtain the audit feature data of each sub-region. Based on the historical audit data of each sub-region, the weights of each indicator are calculated, the weight coefficients corresponding to each indicator are determined, and the weight coefficients based on historical data are obtained. Based on the audit feature data of each sub-region, the indicators in the audit feature data are normalized respectively to obtain the dimensionless normalized value of each indicator. Based on the weighting coefficients calibrated from historical data, the dimensionless normalized values ​​of each indicator are weighted and summed to calculate the path cost correction coefficient corresponding to each sub-region unit. Based on the path cost correction coefficient corresponding to each sub-region unit and the sequence information of the audit nodes in the preliminary audit path, the sub-region unit to which each audit node belongs is determined, and the correspondence between the node and the sub-region is established. Based on the correspondence between nodes and sub-regions, the dwell time of corresponding nodes in the preliminary review path is dynamically adjusted according to the path cost correction coefficient of the sub-region to which each node belongs, so as to obtain the adjusted dwell time of each node; at the same time, according to the relationship between the path cost correction coefficients of the sub-regions to which adjacent nodes belong, the inspection order of adjacent nodes in the preliminary review path is locally rearranged to obtain the adjusted review node sequence. The adjusted audit node sequence and the corresponding adjusted dwell time are integrated to form an integrated on-site audit path; Step 400: Based on the integrated on-site audit path, input the non-conformity information into the industry knowledge graph for correlation and source tracing to obtain the cluster of related non-conformities, and perform similarity matching based on the historical rectification case library to obtain the root cause rectification plan; Step 500 involves using the root cause rectification plan and the path time, non-conformity distribution, and rectification effect of this audit process as full data, and feeding it back to the industry knowledge graph for incremental updates to obtain an evolved knowledge graph. The evolved knowledge graph is then used to review the integrated on-site audit path, obtaining path optimization constraints for subsequent dynamic optimization of the audit path. This includes: collecting data on the root cause rectification plan and the path time, non-conformity distribution, and rectification effect recorded during this audit process, performing standardization processing, and forming structured full data. The structured full data is added to the industry knowledge graph, and the statistical attributes of the relevant nodes are updated to obtain the evolved knowledge graph. The actual audit data for this audit, as well as the historical baseline data for each audit node and sub-region, are extracted from the evolved knowledge graph. By comparing the data to be analyzed with the historical benchmark data, key constraint nodes that exceed the threshold in actual time consumption and high-risk clusters of non-compliance frequency exceeding the threshold are identified. Path optimization constraints are generated based on key constraint nodes and high-risk cluster areas. These constraints include: adjusting the dwell time of key constraint nodes, adjusting the priority of inspection order in high-risk cluster areas, and merging rules for adjacent nodes. The path optimization constraints are stored in the evolutionary knowledge graph as the basis for the next preliminary review of path generation and optimization adjustments.

2. The cross-system integrated audit path optimization method based on industry knowledge graph as described in claim 1, characterized in that, Step 100 includes: Preprocessing of multi-source heterogeneous data includes removing irrelevant characters, correcting format errors, and standardizing technical terms to obtain cleaned data; Based on the cleaned data, named entity recognition technology is used to identify entities in the text, resulting in an entity set. Extract relations from entities in the entity set, determine the semantic associations between entities, and obtain entity relation triples; Add attribute annotations to the entities in the entity set to supplement the entity description information and obtain the entity containing attribute information; The entity relation triples are merged with entities containing attribute information, and conflict detection and resolution are performed. The conflict detection includes at least entity conflict, attribute conflict, relation conflict and semantic conflict, resulting in a resolved knowledge unit. The deconstructed knowledge units are conceptually standardized to unify different expressions of the same concept from different data sources, and a conceptual hierarchy is constructed to obtain standardized knowledge units. Based on standardized knowledge units, an industry knowledge graph covering cross-system standard associations and production process flows is constructed.

3. The cross-system integrated audit path optimization method based on industry knowledge graph as described in claim 2, characterized in that, Step 200 includes: Obtain the target company's pending review materials, and perform format conversion and content extraction on the pending review materials to obtain text data to be processed; Based on industry knowledge graphs, a vector embedding model is used to vectorize the text data to be processed, thereby obtaining the embedding vector of the data to be reviewed. Embedded vectors of cross-system standard knowledge are extracted from industry knowledge graphs and used as a reference benchmark for semantic matching; The similarity between the embedding vector of the data to be reviewed and the embedding vector of cross-system standard knowledge is calculated to obtain the matching degree value. The matching degree value is then converted into a matching label according to a preset threshold. The matching label includes at least the following: duplicate, conflict, and consistency. Based on the matching tags, the content marked as duplicates and conflicts is removed and integrated according to the preset merging rules and priority strategies to obtain the deredundant audit document content. The content of the de-redundant audit documents is structured and formatted in a unified manner to obtain an integrated audit document library.

4. The cross-system integrated audit path optimization method based on industry knowledge graph as described in claim 3, characterized in that, Step 400 includes: During the implementation of the integrated on-site audit path, information on non-conformities is collected and identified. This information includes at least a description of the non-conformity, its location, and the relevant standard clauses. The non-compliance information is input into the industry knowledge graph, and a related search is performed starting from the key entities in the non-compliance description text to obtain a set of related entities; Based on the set of associated entities, corresponding historical non-compliance records are extracted from the industry knowledge graph, and non-compliance items with common root characteristics are aggregated to form a cluster of associated non-compliance items. Extract the feature information of the cluster of items that do not conform to the correlation, and encode the feature information into a feature vector to obtain the feature vector of the current cluster of items that does not conform to the correlation.

5. The cross-system integrated audit path optimization method based on industry knowledge graph as described in claim 4, characterized in that, Step 400 further includes: A historical rectification case library is extracted from the industry knowledge graph. The historical rectification case library contains the feature vectors of historical non-compliance items and the corresponding historical root cause rectification solutions. The weight benchmark values ​​of each feature dimension are extracted from the industry knowledge graph as the first benchmark point, and the average distribution of each feature dimension is statistically analyzed from the historical rectification case library as the second benchmark point. With the goal of minimizing the cross-entropy between the first and second reference points, the weight coefficients of each feature dimension are iteratively optimized to obtain the optimized feature weights. Using the optimized feature weights, the weighted similarity between the feature vector of the current non-compliant item cluster and the feature vector of the historical non-compliant items is calculated, and matching candidate rectification cases are selected based on the weighted similarity. Rectification measures, rectification cycles, and verification methods are extracted from the candidate rectification cases and integrated to generate a root cause rectification plan for the current cluster of items with non-compliance in relevance.

6. A cross-system integrated audit path optimization system based on industry knowledge graphs, wherein the system implements the method as described in any one of claims 1 to 5, characterized in that, include: The graph construction module is used to collect multi-source heterogeneous data from the target industry, extract knowledge and resolve conflicts, and build an industry knowledge graph that covers cross-system standard associations and production process flows. The document library module is used to perform vector embedding and semantic matching on the documents to be reviewed of the target company based on industry knowledge graphs. By identifying and removing duplicate and conflicting content, an integrated review document library is obtained. The path planning module is used to extract cross-system clause association paths and historical audit data from the industry knowledge graph based on the integrated audit document library to obtain the preliminary audit path. Based on the preliminary review path, key review areas are delineated and spatial discretization is performed to obtain review feature data; The path cost correction coefficient is calculated based on the audit feature data, and the dwell time and inspection order of the preliminary audit path are optimized and adjusted to obtain the integrated on-site audit path. The non-compliance processing module is used to input non-compliance information into the industry knowledge graph for correlation and source tracing based on the integrated on-site audit path, obtain related non-compliance clusters, and perform similarity matching based on the historical rectification case library to obtain the root cause rectification plan; The feedback optimization module is used to feed the root cause rectification plan and the path time, non-conformity distribution and rectification effect of this audit process as full data to the industry knowledge graph for incremental updates, resulting in an evolved knowledge graph. The evolved knowledge graph is used to review the integrated on-site audit path to obtain path optimization constraints, which are used for dynamic optimization of the audit path in the future.

Citation Information

Patent Citations

  • Intelligent auditing method based on knowledge graph constraint

    CN121437188A

  • College teacher development path recommendation method and system based on mapping knowledge domain

    CN121481048A