A method, device, equipment, medium and product for constructing a data set for inspection and investigation

By constructing a rule-constrained knowledge graph and a hierarchical rule verification system, combined with a dynamic weight adjustment parameter set, compliance verification and logical correction are performed on the inspection and supervision dataset. This solves the problem of the lack of systematic and standardized governance in the construction of datasets in existing technologies, and realizes the generation of datasets with high logical completeness, strong compliance and traceability.

CN122490429APending Publication Date: 2026-07-31GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG POWER GRID CO LTD
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

The existing inspection and supervision dataset construction schemes lack systematic and standardized governance and multi-source information fusion and optimization capabilities, which cannot meet the business needs of high logic, strong compliance and traceability. They have problems such as incoherent logical reasoning chains, unreasonable explanation of rule application priorities, lagging compliance verification and limited multimodal semantic fusion effects.

Method used

By constructing a rule-constrained knowledge graph, a cross-modal vector index library, and a hierarchical rule verification system, combined with a dynamic weight adjustment parameter set, compliance verification and logical correction are performed on the samples to be verified, generating a target inspection and patrol dataset, and realizing semantic alignment and compliance verification of multimodal data.

Benefits of technology

It has achieved the construction of datasets with high logical completeness, strong compliance constraints and full-chain traceability, which solves the problems of lack of structured and formalized expression of rule knowledge and lagging compliance verification in existing technologies, and improves the standardization and uniformity of datasets and their ability to iterate and optimize autonomously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490429A_ABST
    Figure CN122490429A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, equipment, medium, and product for constructing an inspection and supervision dataset, relating to the field of large-scale inspection and supervision model training technology. First, multimodal raw business data corresponding to inspection and supervision operations are acquired, and a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified are constructed. Then, based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph, a hierarchical rule verification system and a dynamic weight adjustment parameter set are built. Subsequently, relying on the aforementioned knowledge graph, vector index library, verification system, and parameter set, the samples to be verified undergo full-process compliance verification and logical correction to obtain corrected samples. Finally, the corrected samples are imported into a preset training resource library for incremental iterative updates, and associated with corresponding multimodal supporting evidence to generate the target inspection and supervision dataset. This approach ensures both the logical completeness, compliance constraints, and end-to-end traceability of the dataset, while also meeting the requirements of large-scale inspection and supervision model training for high-quality compliant datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large-scale training technology for inspection and supervision models, and in particular to a method, apparatus, equipment, medium, and product for constructing inspection and supervision datasets. Background Technology

[0002] In recent years, with the rapid development of Large Language Models (LLM) and Knowledge-Augmented Generation (KAG) technologies, natural language processing capabilities such as intelligent question answering, automatic compliance verification, and automatic business report generation have been gradually applied in specialized business scenarios in vertical industries with strong compliance control requirements and high rigor. Inspection and supervision, as a specialized business scenario with clear management standards and strong compliance constraints, has three core business characteristics in its data processing: First, the data sources exhibit multimodal and heterogeneous characteristics. The information carriers not only include structured and unstructured texts such as specialized business management standards, work guidelines, business ledgers, and work records, but also involve tables, images, and optical character recognition (OCR). Optical Character Recognition (OCR) documents, audio, and video materials, among other multimodal formats, impose clear technical requirements on the system's cross-modal feature alignment and multimodal semantic embedding capabilities, necessitating unified modeling and parsing of all business data. Secondly, business outputs are subject to strong compliance and high standardization constraints. The outputs related to inspections and audits must not only ensure factual consistency and logical completeness but also strictly adhere to corresponding specific business management standards and control requirements. Therefore, the construction of datasets for model training requires the introduction of core mechanisms such as rule ontology modeling, constraint reasoning, and compliance verification to ensure the dual consistency of generated samples at both the semantic and business specification levels. Thirdly, the business process chain is long, fully covering multiple stages such as problem identification, data extraction, logical reasoning, specification matching, and conclusion generation, placing rigid business requirements on the interpretability and traceability of the dataset and corresponding training models.

[0003] Currently, dataset construction solutions for this type of specialized business still have many technical shortcomings, making it difficult to meet the high-precision and high-standard business usage requirements. Existing solutions generally rely excessively on manual annotation and fixed static rule screening, failing to integrate knowledge graphs to build a systematic business specification constraint framework. This easily leads to problems such as incoherent logical reasoning chains and unreasonable explanations of rule application priorities. At the same time, data compliance verification mostly adopts keyword matching or statistical screening methods, which cannot conduct refined logical verification of business specification clauses. Compliance verification has a significant lag, easily resulting in content output that contradicts business control requirements and internal logical contradictions. In addition, existing technologies lack a complete closed-loop optimization system. Biased sample content and corrected optimized data cannot be effectively fed back to supplement the training resource library. In the long run, this easily leads to low model iteration optimization efficiency and insufficient autonomous convergence ability, requiring a large amount of manual screening and elimination of training data. At the level of multimodal data processing, the development of existing unified modeling technology is still imperfect, the effects of cross-modal semantic fusion and feature alignment are limited, the application of multi-form data fusion is low, and the problem of missing data is prominent. It cannot support content traceability and compliance verification in complex business scenarios, ultimately making it difficult for the overall dataset to meet the core usage requirements of specific businesses for high confidence, strong traceability and standardized uniformity. Summary of the Invention

[0004] This invention provides a method, apparatus, equipment, medium, and product for constructing inspection and supervision datasets, which solves the technical problems of existing inspection and supervision dataset construction lacking systematic and standardized governance and multi-source information fusion and optimization capabilities, and failing to meet the requirements of high logic, strong compliance, and traceability in dataset construction.

[0005] The first aspect of this invention provides a method for constructing an inspection and patrol dataset, comprising: Obtain multimodal raw business data corresponding to inspection and patrol operations, and construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data; Based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph, a hierarchical rule verification system and a dynamic weight adjustment parameter set are built. Based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set, the sample to be verified is subjected to compliance verification and logical correction to obtain a corrected sample. The corrected samples are imported into a preset training resource library and incremental iterative updates are performed. The corresponding multimodal supporting evidence data in the multimodal original business data are associated and bound to generate the target inspection and patrol dataset.

[0006] Optionally, the step of acquiring multimodal raw business data corresponding to the inspection and patrol business, and constructing a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data, includes: Collect text, table, image, optical character recognition document, audio and video data corresponding to inspection and patrol business, and generate multimodal raw business data; The original multimodal business data is sequentially cleaned, classified into modes, extracted into business elements, and extracted into multimodal features to generate standardized multimodal business data. The inspection business elements, business association logic, and compliance constraint clause attributes are extracted from the standardized multimodal business data. A rule-constrained knowledge graph is constructed by combining triple structured orchestration with ontology modeling. The heterogeneous features in the standardized multimodal business data are mapped to a unified semantic space to obtain a unified semantic space feature vector, and a cross-modal vector index library is constructed based on the unified semantic space feature vector. Extract the valid inspection business segments from the standardized multimodal business data and organize and integrate them to generate a sample to be verified.

[0007] Optionally, the step of building a hierarchical rule verification system and a dynamic weight adjustment parameter set based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph includes: Based on the business flow connection form and business coverage of the inspection business association logic in the rule-constrained knowledge graph, a multi-level business execution hierarchy is defined. Based on the compliance constraint clause data in the rule-constrained knowledge graph, the compliance constraint clauses corresponding to each business execution level are matched, and combined with the applicable scenarios and control scope of the compliance constraint clauses, a hierarchical rule verification system is formed through hierarchical integration. Based on the complexity of the business process layout recorded in the aforementioned inspection business association logic, risk differentiation criteria are used to distinguish different businesses. Based on the constraint strength identifier built into the compliance constraint clause data, the control level of different compliance constraint clauses is defined; Based on the aforementioned business risk differentiation criteria and clause control levels, the verification order of each compliance constraint clause in the aforementioned compliance constraint clause data is uniformly arranged to determine the rule priority; Based on the rule priority, static weight parameters are matched for each of the compliance constraint clauses, cross-modal alignment thresholds are configured for the cross-modal content verification stage, risk classification thresholds are configured for the business risk identification stage, and all configured parameters and thresholds are summarized and integrated to generate a dynamic weight adjustment parameter set.

[0008] Optionally, the step of performing compliance verification and logical correction on the sample to be verified based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set to obtain a corrected sample includes: The compliance constraint clauses built into the rule-constrained knowledge graph are invoked, and the hierarchical control specifications of the hierarchical rule verification system are combined to compare the compliance content of each sample to be verified, generating a compliance verification list; The system retrieves multimodal business feature data from the cross-modal vector index library and performs multi-source content comparison by combining the cross-modal alignment threshold corresponding to the dynamic weight adjustment parameter set, thereby generating a cross-modal consistency verification list. Based on the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, the illegal content, business logic conflict content, and cross-modal deviation content within the sample to be verified are identified, and a list of issues to be corrected is generated. Based on the hierarchical control requirements of the hierarchical rule verification system, the logical conflict content in the list of issues to be corrected is resolved by rule conflict resolution, and the business logic deviation is simultaneously calibrated by the multi-level business execution hierarchy to generate logical calibration data. Based on the constraints of the cross-modal alignment threshold, the descriptions of cross-modal deviations in the list of issues to be corrected are standardized, the violations are rectified and adjusted, and compliant correction data is generated. Based on the logical calibration data and the compliance correction data, the sample content of the sample to be verified is uniformly standardized to generate a corrected sample.

[0009] Optionally, the step of identifying violations, business logic conflicts, and cross-modal deviations within the sample to be verified, and generating a list of issues to be corrected, based on the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, includes: The static weight parameters and risk classification thresholds in the dynamic weight adjustment parameter set are called, and the compliance judgment content recorded in the compliance verification checklist is combined to identify non-compliant segments in the sample to be verified item by item, and generate compliance violation labeling data. Based on the discrimination criteria of the risk classification threshold, and combined with the multi-source content differences recorded in the cross-modal consistency verification list, the multi-modal business content difference segments in the sample to be verified are screened item by item to generate cross-modal deviation identification data. Based on the business flow specifications and hierarchical rule verification system control logic of the multi-level business execution hierarchy in the rule-constrained knowledge graph, combined with the compliance and violation labeling data and the cross-modal deviation identification data, the business connection contradictions in the samples to be verified are investigated, and business logic conflict identification data is generated. The compliance and violation labeling data, the cross-modal deviation identification data, and the business logic conflict identification data are summarized, categorized and organized according to problem type, and the corresponding rectification basis for the labeled problems is unified to compile a list of problems to be corrected.

[0010] Optionally, the step of importing the corrected samples into a preset training resource library and performing incremental iterative updates, associating and binding the corresponding multimodal supporting evidence data in the multimodal original business data, and generating a target inspection and patrol dataset includes: The corrected samples are standardized and formatted according to the hierarchical management and control specifications of the hierarchical rule verification system, and then imported into the preset training resource library to generate a sample archive record table. The dynamic weight adjustment parameter set is invoked, and combined with the sample archive record table, to perform incremental iterative updates on the existing samples and newly added correction samples within the preset training resource library, generating an incremental update log for the resource library. Based on the rule-constrained knowledge graph, the original inspection business segments corresponding to the corrected samples within the multimodal original business data are matched to generate a supporting evidence matching relationship table. Based on the hierarchical rule verification system and the supporting evidence matching table, the corrected sample is bidirectionally associated and bound with the corresponding multimodal supporting evidence data in the multimodal original business data to generate a sample supporting evidence binding list. The sample supporting binding list is subjected to compliance verification, invalid binding content within the list is removed, and a binding verification report is generated; Based on the business association logic of the rule-constrained knowledge graph, the sample archive record table, the resource library incremental update log, the sample supporting binding list, and the binding verification report are classified and organized to generate the target inspection and patrol dataset.

[0011] A second aspect of the present invention provides an apparatus for constructing inspection and patrol datasets, comprising: The basic data construction module is used to acquire multimodal raw business data corresponding to inspection and patrol business, and to construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data. The rule parameter construction module is used to build a hierarchical rule verification system and a dynamic weight adjustment parameter set based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph. The sample compliance processing module is used to perform compliance verification and logical correction on the sample to be verified based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set, so as to obtain the corrected sample. The dataset integration and generation module is used to import the corrected samples into a preset training resource library and perform incremental iterative updates, and associate and bind the corresponding multimodal supporting evidence data in the multimodal original business data to generate the target inspection and patrol dataset.

[0012] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the patrol and inspection dataset construction method described above.

[0013] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the patrol and inspection dataset construction method as described above.

[0014] The fifth aspect of the present invention provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein, when the program instructions are executed by a computer, the computer performs the patrol and inspection dataset construction method as described above.

[0015] As can be seen from the above technical solutions, the present invention has the following advantages: This invention provides a scheme for constructing an inspection and supervision dataset. First, it acquires multimodal raw business data corresponding to inspection and supervision operations, completes data standardization processing and feature extraction, and simultaneously constructs a rule-constrained knowledge graph carrying business-related logic and compliance constraints, a cross-modal vector index library for semantic alignment of multi-source heterogeneous data, and samples to be verified for subsequent compliance checks, thus building a solid data foundation and institutional constraint base for the dataset. Then, based on the business-related logic and compliance constraints data accumulated in the rule-constrained knowledge graph, it divides multi-level business execution hierarchies and matches corresponding compliance constraints, establishing a hierarchical rule verification system adapted to business scenarios and control scope. Simultaneously, it determines rule priorities by combining business risk levels and clause control levels, matches static weights, alignment thresholds, and risk classification thresholds required for full-process verification, forming an adaptively adjustable dynamic weight adjustment parameter set, and establishing a refined compliance control framework covering the entire business process. Subsequently, supported by a rule-constrained knowledge graph, a cross-modal vector index library, a hierarchical rule verification system, and a dynamic weight adjustment parameter set, the system performs line-by-line compliance content comparison and multi-source content consistency verification on the samples to be verified, accurately identifying violations, business logic conflicts, and cross-modal deviations within the samples. Then, it specifically completes rule conflict resolution, business logic calibration, and cross-modal content standardization, achieving full-dimensional compliance verification and logical correction of the samples to be verified, resulting in standardized corrected samples. Finally, the corrected samples are imported into a pre-set training resource library according to hierarchical management standards, and incremental iterative updates are performed using the dynamic weight adjustment parameter set. Simultaneously, based on the rule-constrained knowledge graph, corresponding multi-modal supporting evidence data is matched to complete bidirectional association binding, ultimately generating a target inspection and audit dataset with high logical completeness, strong compliance constraints, and full-chain traceability. This technical solution, through rule ontology modeling and the construction of a hierarchical rule verification system, solves the problems of lack of structured and formalized expression of rule knowledge and the lack of systematic and standardized governance caused by the reliance on static keyword matching for compliance verification in existing technologies. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the steps of a method for constructing an inspection and patrol dataset according to Embodiment 1 of the present invention. Figure 2 This is a flowchart illustrating the steps of a method for constructing an inspection and patrol dataset according to Embodiment 2 of the present invention.

[0018] Figure 3 This is a system architecture layered diagram corresponding to the patrol and inspection dataset construction method provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of a dynamic weighting mechanism provided in Embodiment 2 of the present invention; Figure 5 This is a technical flow diagram corresponding to the method for constructing an inspection and patrol dataset provided in Embodiment 2 of the present invention; Figure 6 This is a structural block diagram of an inspection and patrol dataset construction device provided in Embodiment 3 of the present invention; Figure 7 This is a structural block diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below in conjunction with the known technical basis and the problems existing in the prior art.

[0020] In the field of dataset construction and model training for business scenarios with strong compliance constraints, the following well-known technologies are widely used: 1. Knowledge Graph Modeling and Rule Reasoning Knowledge graphs are structured networks used for storing and reasoning about domain-specific knowledge. Their general form can be represented as: =( E , , ),in, It is a general-purpose knowledge graph; E For entities, in compliance and regulatory scenarios, it usually refers to structured noun nodes extracted from regulatory documents, such as regulatory names, institutions, positions, business processes, etc. A set of relationships, i.e., directed edges connecting different entities, is used to express the business logic and constraint relationships between entities, such as "contains", "belongs to", "prohibits", "prerequisites", etc. A set of attributes or triplet facts refers to the intrinsic characteristic values ​​of an entity (such as promulgation time, priority, scope of application, etc.), or a set of underlying facts consisting of (head entity, relation, tail entity).

[0021] Ontology modeling enables the construction of hierarchical knowledge graph structures, which can then be combined with a rule engine for matching and reasoning, including common compliance determination functions. It can be represented as: ; in, It is a general-purpose knowledge graph; yThe target dataset sample to be verified covers all types of data content from inspection and supervision operations, including structured / unstructured text, multimodal parsing information, logical reasoning results, rule matching content, and conclusion text—all data carriers to be reviewed. This method can determine whether the generated results conform to the constraints of the knowledge graph, but it has significant limitations: rule weights are usually statically set, making it difficult to adapt to different business scenarios and the uncertainty of the output results.

[0022] 2. Language Model Quality Control and Feedback Optimization

[0023] Existing methods often employ manual annotation, data augmentation, and Reinforcement Learning with Human Feedback (RLHF) to control the quality of generated results. Within the RLHF framework, the general definition of the reward function is: ; in, This is the final value of the reward function; This serves as an indicator of the accuracy of the generated results; For compliance metrics of the generated results; These are the weighting coefficients for accuracy and compliance, respectively, used to balance the two optimization objectives. However, this type of method can only penalize or reward the generated results at the overall level, without integrating the confidence distribution of the model output with the rule priority into the compliance verification process, thus failing to achieve refined compliance constraints.

[0024] 3. Multimodal feature extraction and alignment

[0025] For business scenarios involving multi-source data such as text, images, tables, and audio, well-known technologies often employ methods such as optical character recognition, table parsing, image embedding, and speech recognition to map different modal inputs to a unified semantic space. The general feature mapping process can be represented as: ; in, To represent the original input data After feature extraction by the encoder, the feature vectors are mapped to a unified high-dimensional semantic space. mod This serves as a modality type identifier, referring to a set of different data modalities; To represent a specific mode mod This refers to a dedicated feature extraction function or encoder. In known technologies, it is typically implemented using pre-trained deep neural networks (such as BERT for text processing, ResNet or ViT for image processing, etc.). To indicate that it belongs to a specific mode mThe original input data. For example, when mod When it is text, it is a sequence of characters; when it is an image, it is a matrix of pixels; when it is audio, it is an acoustic signal, etc.

[0026] To achieve cross-modal feature alignment, a contrastive learning loss function is typically introduced, with the general form being: ; in, To compare the learning loss, the optimization objective is to bring semantically relevant cross-modal features closer together and to infer semantically irrelevant features. This is the feature vector of all sample sets (including positive and negative samples) within the current batch. The summation operation in the denominator is used for global normalization. To form a "positive sample pair" of feature vectors, for example, a scanned copy of a document (image modality) of the same business material and its corresponding OCR parsing result (text modality); Cosine similarity is a commonly used similarity metric function between feature vectors in cross-modal alignment tasks. The temperature coefficient is used to scale the similarity score, adjust the smoothness of the distribution that the model uses to distinguish between positive and negative samples, and control the penalty for difficult negative samples.

[0027] This type of method can achieve semantic fusion of multimodal data, but it lacks a cross-modal compliance verification mechanism at the institutional level and cannot meet the evidence chain consistency verification in scenarios with high compliance requirements.

[0028] 4. Closed-loop optimization and incremental training

[0029] Some solutions introduce incremental fine-tuning and low-rank adaptation (LoRA) techniques, achieving iterative model optimization by rewriting samples back into the training set. The general parameter update formula is: ; in, For the updated model parameters, i.e. the first... t The model weight state after +1 iteration; The model parameters for the current iteration, i.e., the nth iteration. t The model weight state at the next iteration; The learning rate controls the step size for each parameter update. The gradient operator represents taking the partial derivative of the loss function with respect to the direction in which the loss function decreases the most; This is the loss function, used to calculate the error between the model's current output and the expected output. The input data typically refers to raw text or multimodal features; The corrected sample labels are the expected output results of writing back the training set after logical verification and correction.

[0030] However, most existing correction mechanisms rely on manual review, lack sufficient automation, and are difficult to achieve efficient closed-loop optimization.

[0031] The aforementioned known technologies still have significant shortcomings when applied to the construction of datasets in scenarios with strong compliance constraints, such as inspections and audits, and are difficult to meet business needs: 1. Rule knowledge lacks structured and formalized expression. Current technologies do not transform institutional norms into structured ontology models; rule sets are stored only as informal text, making logical reasoning through triple representation impossible. Let the rule base be... ,in, The rules, being in natural language, lack formal definitions of entities, relationships, and attributes, making it impossible to effectively determine the compliance of the output results using the aforementioned compliance determination function during constraint reasoning. Empirical data shows that in scenarios with high compliance requirements, the accuracy of this type of rule conflict detection is only 0.58, and the proportion requiring manual intervention is as high as 45%, severely impacting the efficiency of dataset construction.

[0032] 2. The separation of generation and verification architecture leads to delays in compliance verification.

[0033] Existing technologies generally adopt a "generation-verification" separation architecture, resulting in a time-separation gap between data generation and compliance checks. The process can be represented as follows: ; in, This is the initial data source or input instruction set generated; For large language models or algorithm engines used to generate content, Refers to model parameters; The initial sample set generated for the model (the raw output sequence that has not been filtered for compliance). For rules based on static rule base or manually preset rules R The hysteresis check function; This is the set of compliant samples that have been verified and screened to be valid.

[0034] When sample size At that time, the cost of error correction increases exponentially O( The policy violation rate surged from 5% with a low sample size to 28%; meanwhile, the sample correction time followed a quadratic function relationship: ; in, Total time spent correcting for the sample; , which is the quadratic time complexity coefficient, representing the non-linear time consumption caused by the lack of graph-based ontology constraints, cross-sample rule comparison, and complex logical conflict investigation. is the linear time complexity coefficient, representing the linear time consumption when performing basic traversal or keyword filtering on a single sample. This time-consuming characteristic leads to uncontrolled efficiency in correction when the sample size is large, making it unsuitable for building large-scale compliant datasets.

[0035] 3. Multimodal data lacks unified semantic mapping and compliance verification.

[0036] In existing technologies, multimodal data such as text, images, and tables typically employ independent feature extraction processes, lacking a unified semantic mapping mechanism. For example, the embedding representation of text data is... (e.g., the output of the Sentence-BERT model), the feature representation of image data is as follows: (As shown in the ResNet-50 model output), the two did not pass through the cross-modal mapping function. Feature alignment resulted in feature space conflicts in approximately 41% of multimodal samples, causing the Evidence Chain Integrity Index (ECI) to drop to 0.59 (of which...). This measure, used to measure the consistency of cross-modal evidence, cannot meet the requirements of business scenarios for the integrity of the evidence chain.

[0037] To address the problems existing in the prior art, this invention provides a method, apparatus, device, medium, and product for constructing inspection and supervision datasets. These solutions address the technical issues that existing inspection and supervision dataset construction lacks systematic and standardized governance and multi-source information fusion and optimization capabilities, failing to meet the requirements for high logic, strong compliance, and traceability in dataset construction.

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] Please see Figure 1 , Figure 1 A flowchart illustrating the steps of a method for constructing an inspection and patrol dataset according to an embodiment of the present invention.

[0040] This invention provides a method for constructing an inspection and patrol dataset, comprising: Step 101: Obtain multimodal raw business data corresponding to the inspection and patrol business, and construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data.

[0041] In this embodiment of the invention, text, tables, images, optical character recognition documents, audio, and video data corresponding to inspection and patrol operations are collected and aggregated to generate multimodal raw business data. The multimodal raw business data undergoes data cleaning, modality classification, business element extraction, and multimodal feature extraction processing sequentially to obtain standardized multimodal business data. Inspection and patrol business elements, business association logic, and compliance constraint clause attributes are extracted from the standardized multimodal business data. A rule-constrained knowledge graph is constructed using a triplet structured arrangement combined with ontology modeling. Heterogeneous features in the standardized multimodal business data are mapped to a unified semantic space to obtain unified semantic space feature vectors. A cross-modal vector index library is built based on these unified semantic space feature vectors. Effective inspection and patrol business segments within the standardized multimodal business data are extracted, regularized, and integrated to finally generate samples to be verified.

[0042] Step 102: Based on the business relationship logic and compliance constraint clause data in the rule-constrained knowledge graph, build a hierarchical rule verification system and a dynamic weight adjustment parameter set.

[0043] In this embodiment of the invention, based on the business flow connection form and business coverage corresponding to the inspection business association logic in the rule-constrained knowledge graph, a multi-level business execution hierarchy is formed; the compliance constraint clause data recorded in the rule-constrained knowledge graph is called, and the compliance constraint clauses corresponding to each multi-level business execution hierarchy are matched. Combining the applicable scenarios and control scope of each compliance constraint clause, a hierarchical rule verification system is constructed by hierarchical integration; according to the complexity of the business process arrangement recorded in the inspection business association logic, risk differentiation standards for different businesses are formed; based on the constraint strength identifier built into the compliance constraint clause data, the control level corresponding to each compliance constraint clause is defined; according to the business risk differentiation standards and clause control levels, the verification order of all compliance constraint clauses is uniformly arranged to determine the rule priority; based on the determined rule priority, static weight parameters are matched for each compliance constraint clause, cross-modal alignment thresholds are configured for the cross-modal content verification stage, and risk classification thresholds are configured for the business risk judgment stage. All configured parameters and thresholds are uniformly summarized and integrated to generate a dynamic weight adjustment parameter set.

[0044] Step 103: Based on the rule-constrained knowledge graph, cross-modal vector index library, hierarchical rule verification system and dynamic weight adjustment parameter set, perform compliance verification and logical correction on the sample to be verified to obtain the corrected sample.

[0045] In this embodiment of the invention, the compliance constraint clauses built into the rule-constrained knowledge graph are invoked, and the hierarchical control specifications of the hierarchical rule verification system are combined to compare the compliance content of each sample to be verified, generating a compliance verification list; the cross-modal vector index library is retrieved to search for multimodal business feature data, and multi-source content comparison is performed in combination with the cross-modal alignment threshold corresponding to the dynamic weight adjustment parameter set, generating a cross-modal consistency verification list; combining the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, non-compliant segments of the sample to be verified are identified, differences in multimodal business content are screened, and contradictions in business connections are investigated, according to... The system first generates compliance and violation labeling data, cross-modal deviation identification data, and business logic conflict identification data. After summarizing and classifying these multiple types of identification data, a list of issues to be corrected is compiled. Based on the hierarchical control requirements of the layered rule verification system, the logical conflicts within the list of issues to be corrected are resolved through rule conflict resolution. Simultaneously, business logic deviations are calibrated at multiple business execution levels to generate logical calibration data. According to the constraints of the cross-modal alignment threshold, the cross-modal deviation content in the list of issues to be corrected is standardized, and the violations are adjusted to generate compliance correction data. Finally, the logical calibration data and compliance correction data are integrated, and the overall content of the sample to be verified is uniformly standardized to obtain the corrected sample.

[0046] Step 104: Import the corrected samples into the preset training resource library and perform incremental iterative updates, associate and bind the corresponding multimodal supporting evidence data in the multimodal original business data, and generate the target inspection and patrol dataset.

[0047] In this embodiment of the invention, the corrected samples are standardized and formatted according to the hierarchical management and control specifications of the hierarchical rule verification system, and then uniformly imported into a preset training resource library to generate a sample archive record table. A dynamic weight adjustment parameter set is invoked, and incremental iterative updates are performed synchronously on the existing samples and newly added corrected samples within the preset training resource library, generating an incremental update log for the resource library. Based on the business association logic of the rule-constrained knowledge graph, the original inspection business segments corresponding to the corrected samples in the multimodal original business data are matched to generate a supporting evidence matching relationship table. Combining the hierarchical rule verification system and the supporting evidence matching relationship table, a two-way association binding is completed between the corrected samples and the corresponding multimodal supporting evidence data within the multimodal original business data, generating a sample supporting evidence binding list. Compliance verification is conducted on the sample supporting evidence binding list, invalid binding content is removed, and a binding verification report is generated. Based on the business association logic of the rule-constrained knowledge graph, various types of form records are uniformly classified, standardized, and integrated to finally generate the target inspection dataset.

[0048] Please see Figure 2 , Figure 2This is a flowchart illustrating the steps of a method for constructing an inspection and patrol dataset according to Embodiment 2 of the present invention.

[0049] This invention provides a method for constructing an inspection and patrol dataset, comprising: Step 201: Obtain multimodal raw business data corresponding to the inspection and supervision business, and construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data.

[0050] It is worth mentioning that, in order to improve sample coverage and semantic diversity, the system performs data augmentation by leveraging the generative capabilities of a large language model while maintaining rule constraints. Let the original samples be... The enhancement function is ,but: The enhancement operations include synonym substitution, semantic expansion, and context reconstruction to ensure the expanded sample set. This means that standardized multimodal business data maintains rule compliance while ensuring distribution consistency.

[0051] Further, step 201 may include the following sub-steps: S11. Collect text, table, image, optical character recognition document, audio and video data corresponding to the inspection and patrol business, and generate multimodal raw business data.

[0052] In this embodiment of the invention, the entire process of internal inspection and auditing of the enterprise is collected and archived with relevant business data, including business control ledgers, on-site business visit records, business verification drafts, original project vouchers, project initiation and filing materials, supply chain procurement and bidding archives, personnel management and salary assessment files, job performance self-inspection reports, archived materials of previous internal verification and rectification, industry operating standard provisions, written materials of business supervision and verification, and supporting documents for verification of business status clues, etc. Let the original input dataset corresponding to the multimodal original business data to be processed be: ; In the formula, This is a collection of multimodal raw business data. For the set of the first i A separate set of multimodal raw business data, n The total amount of data is obtained through multi-source collection, resulting in standardized multimodal raw business data, which provides the raw input basis for subsequent processing.

[0053] S12. Perform data cleaning, modality classification, business element extraction, and multimodal feature extraction on the original multimodal business data in sequence to generate standardized multimodal business data.

[0054] In this embodiment of the invention, preprocessing operations are performed sequentially on the multimodal raw business data: data cleaning is used to remove invalid characters, format marks, redundant information, and other non-business content; the data is classified into text, OCR image, table, audio, and video categories through modal classification; core entities of the inspection business are extracted through business element extraction, including systems, business positions, processes, and unmatched rule items; and for each modal data, a mapping function is embedded. Unified mapping to semantic space The complete form of the embedding mapping function is: ; The text data is used to generate text embedding feature vectors through the BERT-base model. Image data / OCR data are used to generate image feature vectors through the CLIP ViT-L / 14 model. Audio data is used to generate audio feature vectors through the Wav2Vec2.0 model. , further expressed as: ; Simultaneously, for individual modal data, projection is performed through the embedding function. Generate feature vectors: ; In the formula, , , , The original data are in the formats of text, image / OCR, table, and audio modalities, respectively. , , These are the feature vectors corresponding to each mode; This is the mapped unified semantic feature vector; It is a 1024-dimensional semantic space; For the first i Multimodal raw business data, For the first mod Projection of the embedding function corresponding to each modality For the first i Feature vectors of the data for d A unified semantic space is then established. Subsequently, a rule-constrained knowledge graph is utilized. Align entities and relationships: ,in, G It is a rule-constrained knowledge graph; E For a collection of entities; It is a set of constraint relations; It is a set of attribute features; Version timestamp; An executable rule set; Conflict resolution strategies; To finally match and bind the target entity of the knowledge graph; e ∈ E Candidate entities in the knowledge graph; The multimodal feature vector after preprocessing the input data; For similarity calculation functions (such as cosine similarity); A threshold for entity link similarity is used; alignment is only considered successful when the maximum similarity is greater than or equal to this threshold, filtering out low-confidence erroneous bindings and achieving structured binding between data and rule entities. Through the above embedding mapping, standardized representation and semantic alignment of multi-source heterogeneous data are achieved, generating standardized multimodal business data.

[0055] Furthermore, multimodal feature extraction can also be achieved using an evidence chain reconstruction architecture based on multimodal fusion: For multi-source heterogeneous data such as text, invoice images, audio recordings, and video surveillance, a cross-modal attention mechanism is used to jointly model information from different modalities. The joint modeling formula is as follows: ; in, This is the joint fusion feature encoding output after cross-modal attention computation; Q The query matrix is ​​generated by the linear mapping of dominant modality features. Its sequence length, d For feature dimensions; K The key matrix is ​​generated by the linear mapping of auxiliary modal features; V The actual feature content of the modality to be fused that is homologous to the key matrix; The dimension of the key vector. As a scaling factor, it prevents the softmax function from entering the gradient saturation region due to excessively large sum-to-dot product results, thus ensuring the stability of model training. This scheme can simultaneously utilize invoice OCR features, speech emotion vectors, and text embeddings to form a joint representation of multimodal features.

[0056] S13. Extract inspection business elements, business relationship logic, and compliance constraint clause attributes from standardized multimodal business data, and construct a rule-constrained knowledge graph through triple structured orchestration combined with ontology modeling.

[0057] In this embodiment of the invention, inspection business elements, business association logic, and compliance constraint clause attributes are extracted from standardized multimodal business data. Based on the OWL2 ontology language and RDF triple structure, the inspection policy rules are formalized into a six-tuple structure: In the formula, E A collection of entities (systems, business positions, processes); For a set of constraint relations (such as mustNot, mustHave, priority, subClassOf, partOf); For attribute feature set (timeliness, scope of application, weight coefficient); Version timestamp; An executable rule set; For conflict resolution strategies; arranged according to the logical combination of entity-relationship-entity triples, combined with ontology hierarchy. This process achieves hierarchical and formalized rule modeling to construct a rule-constrained knowledge graph, forming a knowledge graph of inspection and supervision policies and rules with institutional-level constraints. Joint encoding is performed on multi-source heterogeneous data such as text, images, tables, and relationship graphs to obtain joint feature codes for the multi-source heterogeneous data. This encoding is fundamental to ensuring the continuity of the reasoning chain. The encoding process is implemented through multimodal joint encoding and mapping functions. Inputs include feature codes extracted from text data, feature codes extracted from image data, feature codes extracted after parsing table data, and entity or topological structure feature codes extracted from relationship graph data. The mapping function can employ cross-modal attention mechanisms or contrastive learning alignment functions to map features from different modalities to a unified semantic space, achieving effective fusion of multimodal data.

[0058] Combination Figure 3 As shown, this invention employs a four-layer progressive system architecture to support the implementation of the above method. Each architecture layer and its built-in components correspond one-to-one with the steps in this embodiment: The knowledge graph ontology layer, i.e., the rule-constrained knowledge graph construction stage, includes an entity relationship definition module, a domain rule base component, and a semantic verification interface, responsible for the ontological modeling and structured accumulation of institutional clauses; The real-time compliance verification and logic correction layer, i.e., the subsequent compliance verification and logic correction stage, includes a conflict detection engine, a reasoning rule engine, and an entity link repair tool, responsible for rule conflict detection, logical reasoning verification, and entity repair and error correction; The confidence-rule weight dynamic adjustment layer, i.e., the dynamic weight adjustment and risk identification stage, includes a weight calculation module, a confidence assessment component, and a dynamic balance regulator, responsible for confidence assessment, weight calculation, and adaptive adjustment; The closed-loop feedback and multimodal fusion layer, i.e., the incremental iterative update and multimodal evidence binding stage, includes a feedback data collector, a text-image fusion interface, and a knowledge update trigger, responsible for data backflow collection, multimodal feature fusion, and knowledge system update. The four-layer architecture collaboratively completes the construction of the entire process dataset.

[0059] S14. Map the heterogeneous features in the standardized multimodal business data to a unified semantic space to obtain a unified semantic space feature vector, and build a cross-modal vector index library based on the unified semantic space feature vector.

[0060] In this embodiment of the invention, the feature vectors of each modal data generated in S12 are used... All feature vectors are associated with the original data identifiers and stored in batches in pre-built vector databases such as FAISS and Milvus. An HNSW graph index is constructed based on the cosine similarity algorithm to form a cross-modal vector index library. At the same time, the feature vectors are aligned with the entity set of the rule-constrained knowledge graph to establish the association between the feature vectors and the graph entities.

[0061] It is worth mentioning that the vector index library can be replaced with a retrieval architecture based on a vector database to adapt to large-scale knowledge base scenarios: Let the knowledge base document set be... By embedding functions Mapping documents to a vector space and constructing an HNSW (Hierarchical Navigable Small World) graph index results in a query complexity approximately O(n log n). Compared to traditional sequential retrieval methods This scheme significantly reduces computational complexity in large-scale retrieval of legal provisions and historical cases, ensuring the real-time availability of evidence required for the reasoning chain.

[0062] S15. Extract effective inspection business segments from standardized multimodal business data, organize and integrate them, and generate samples to be verified.

[0063] In this embodiment of the invention, invalid and redundant segments in standardized multimodal business data are removed, and valid inspection business segments for compliance verification and logic correction are extracted; combined with the feature vector set obtained through embedding mapping in step 201. (i.e., preprocessed results), through a large language model Generate candidate samples y The formula for generating the formula is: in, H This is a set of feature vectors obtained after standardizing multimodal business data through embedding mapping. For large language models, through receiving H Candidate samples are generated as input. y The candidate sample y This is the sample to be verified in subsequent compliance checks.

[0064] Furthermore, the generation of samples to be verified can also be achieved using a knowledge-driven reasoning architecture based on retrieval enhancement: Centered on Retrieval-Augmented Generation (RAG), a knowledge base and a legal case database for inspection and supervision are established to handle input queries. q First, through vectorized embedding functions Mapping text to a semantic space: Then, a Top-k similarity search was performed to obtain the evidence set. The data is then injected into the generative model to participate in the reasoning chain; this approach uses professional knowledge to support the generation process, improving sample compliance and evidence traceability.

[0065] It is worth mentioning that the reasoning process for generating the sample to be verified can be implemented using a dynamic reasoning replacement mechanism based on a process reward model, replacing the traditional static chained reasoning path: The reasoning task is modeled as a Markov Decision Process (MDP), and the state transition probabilities are defined. And introduce a process reward function: ; in, The logical consistency score is used to assess the logical consistency of the reasoning process. Indicates the fact matching degree, used to evaluate the degree of matching between the reasoning result and the facts in the knowledge base; This represents cross-modal consistency, used to evaluate the logical consistency of multimodal information; it is achieved through the policy gradient method. The model is dynamically optimized to enable the system to adaptively correct itself in complex tasks, thus solving the problem that traditional chain reasoning relies on a single path and suffers from cumulative path deviations.

[0066] It's worth noting that the large language model can be replaced with a lightweight sub-model based on knowledge distillation to adapt to scenarios with limited computing power: the original large model is used as the teacher model, and a lightweight student model is obtained through distillation. During the distillation process, temperature parameters are used to... Adjusting the smoothness of soft label distribution reduces inference latency by approximately 60% while maintaining student model accuracy, making it particularly suitable for real-time deployment of patrol and inspection edge terminals.

[0067] Step 202: Based on the business relationship logic and compliance constraint clause data in the rule-constrained knowledge graph, build a hierarchical rule verification system and a dynamic weight adjustment parameter set.

[0068] Furthermore, step 202 may include the following sub-steps: S21. Based on the business flow connection form and business coverage of the inspection business association logic in the rule-constrained knowledge graph, divide the business execution level into multiple levels.

[0069] In this embodiment of the invention, the rule-constrained knowledge graph is traversed, and the inspection business is divided into multiple business execution levels based on the business flow connection form and business coverage of the inspection business association logic. These levels include the overall inspection level, the special domain level, and the specific business link level. The business boundaries, flow logic, and control scope of each level are clearly defined, providing a basis for hierarchical division for the layered rule verification.

[0070] S22. Based on the compliance constraint clause data in the rule-constrained knowledge graph, match the compliance constraint clauses corresponding to each business execution level, and combine the applicable scenarios and control scope of the compliance constraint clauses to form a hierarchical rule verification system.

[0071] In this embodiment of the invention, compliance constraint clause data from the rule-constrained knowledge graph is retrieved, and the corresponding applicable compliance constraint clauses are accurately matched for each level of business execution. The clauses are then hierarchically integrated based on their applicable scenarios and control scope, and rule logic reasoning is implemented through the SWRL rule engine. The rule-constrained knowledge graph structure upon which this step relies is as follows: Based on this, the core decision unit of the hierarchical rule validation system—the rule constraint function—is defined as follows: ; in, y The sample to be verified is generated for subsequent steps; G For rule-constrained knowledge graphs, this function is used to evaluate samples generated by large language models. Does it trigger a constraint conflict in the knowledge graph? At that time, the logical consistency determiner is invoked to analyze internal inconsistencies in the sample, and the rule-driven correction operator is used. R Generate correction results To achieve real-time error correction; when When the condition is met, the sample is deemed to conform to the rule constraints and no correction is required.

[0072] This function provides real-time compliance verification capabilities for subsequent compliance verification processes, enabling hierarchical compliance control with "level correspondence and clause adaptation" to ensure that different business levels implement the corresponding compliance standards.

[0073] It is worth mentioning that, under the constraints of knowledge graphs, large language models Generate candidate sample sequences: And combine the chain reasoning template to develop the reasoning path. , A complete reasoning path unfolding from a large language model using a chain-based reasoning template. For (from the original) (replaced) represents the first inference path One intermediate reasoning step or sub-conclusion node (of which) j =1, 2, ..., ; This represents the total number of steps in this reasoning path. Ensure the generated results are logically interpretable and correspond to the evidence.

[0074] S23. Based on the complexity of the business process arrangement recorded in the inspection business association logic, distinguish the risk differentiation standards for different businesses.

[0075] In this embodiment of the invention, based on the complexity of the business process arrangement recorded in the rule-constrained knowledge graph, differentiated risk differentiation standards are defined for different inspection businesses: the higher the complexity of the business process and the stricter the control requirements, the more detailed the risk assessment dimensions; a higher static priority weight range for rules is set for businesses with high process complexity; a lower static priority weight range for rules is set for businesses with low process complexity; and the adjustment rules for dynamic weight factors are adjusted according to the business risk level. This risk differentiation standard provides a basis for the parameter configuration of the subsequent dynamic scoring function and provides a judgment benchmark for the subsequent dynamic scoring and logic correction stages.

[0076] S24. Determine the control level of different compliance constraint clauses according to the constraint strength identifier built into the compliance constraint clause data.

[0077] In this embodiment of the invention, compliance constraint clauses are divided into three control levels according to the constraint strength identifier built into the compliance constraint clause data: Level 1 core constraint (highest constraint strength); Level 2 regular constraint (medium constraint strength); and Level 3 process constraint (basic constraint strength). The constraint strength is positively correlated with the control level, which is directly mapped to the static priority parameter of the rules in the dynamic scoring function: Level 1 core constraint corresponds to the highest static priority weight, Level 2 regular constraint corresponds to medium weight, and Level 3 process constraint corresponds to basic weight, providing parameter basis for the subsequent configuration of the dynamic scoring function.

[0078] S25. Based on the business risk differentiation standards and clause control levels, uniformly arrange the verification order of each compliance constraint clause in the compliance constraint clause data to determine the rule priority.

[0079] In this embodiment of the invention, combining the business risk differentiation criteria defined in step S23 and the clause control levels defined in step S24, the verification order of all compliance constraint clauses is uniformly sorted: clauses with higher control levels and higher risk levels are verified first, thus determining the rule priority. The purpose of this priority sorting includes: resolving rule conflicts and ensuring that high-priority rules are executed first; and directly affecting the weighting of the static priority parameter of rules in the dynamic scoring function, with higher-priority rules having a higher weighting in the dynamic scoring.

[0080] S26. Based on rule priority, match static weight parameters for each compliance constraint clause, configure cross-modal alignment thresholds for cross-modal content verification, configure risk classification thresholds for business risk identification, and summarize and integrate all configured parameters and thresholds to generate a dynamic weight adjustment parameter set.

[0081] In this embodiment of the invention, based on the rule priority determined in step S25, the following parameter configuration is completed: 1. Static weight parameter configuration: Configure the static priority of rules in the dynamic scoring function for each compliance constraint clause. ; 2. Cross-modal alignment threshold configuration: Configure a cross-modal alignment threshold for the cross-modal content consistency verification stage, which is used for subsequent semantic similarity comparison of multimodal feature vectors; 3. Risk Classification Threshold Configuration: Configure risk classification thresholds for the business risk assessment process, including mandatory correction thresholds. and manual review threshold ; 4. Dynamic weight update configuration: Dynamic weight factor The update is performed using a Bayesian / Exponential Weighted Moving Average (EWMA) method, with the following formula: ; In the formula, To update the step size; This is the residual for the rule triggering frequency; this update mechanism can dynamically adjust the weights according to the rule triggering frequency, improving the model's responsiveness to high-risk rules. r For the first r Number of updates; r +1 is the first r +1 update count; the dynamic penalty / correction intensity is calculated through weighted summation, involving parameters including the total number of constraint rules triggered when generating content for logical reasoning, the corresponding preset weights, the confidence level of the model-generated content under the corresponding rules, and the indicator function; among them, the value of the preset weight of the rule is positively correlated with the priority of the rule in the knowledge graph; the lower the confidence level, the greater the correction penalty weight is assigned; the indicator function is used to determine whether the generated content conforms to the corresponding rule, and only the part that does not conform to the rule is accumulated with the penalty intensity, so as to realize the automatic correction of non-compliant or logically erroneous content; 5. Dynamic scoring function configuration: Configure the confidence-rule weight joint scoring function, which is defined as follows: ; in, This represents the joint score of the sample. The confidence level weighting coefficient; Output confidence scores for the model (based on token probability entropy) Normalization reflects the reliability of the model output. For a set of rules; For the first r The dynamic weighting factor of each rule (generated by the above update formula and dynamically adjusted according to the rule triggering frequency); For the first r The static priority weight of each rule (determined by the control level in step S24 and the priority in S25); this function integrates the model confidence and rule weights to achieve multi-dimensional risk assessment; 6. Modification Operator Configuration: Configure rules to drive the modification operator. It is used to force correction of high-risk samples and achieve real-time error correction.

[0082] Based on the parameters configured above, the scoring results will trigger different processing logic: when When this is detected, it is deemed a high-risk output and a forced correction is triggered: Call the rule-driven correction operator Generate corrected results; when When marked as medium-risk output, a soft penalty and manual review are applied; dynamic weights are updated using a Bayesian strategy. ; in, For the first t +1 iteration update r The static priority weight of each rule; For the first t The first iteration update r The static priority weight of each rule; For the first t The first iteration update j The static priority weights of the rules are used; the rule weights are adaptively converged over time through online updates.

[0083] The results of manual verification are recorded as labels. And it participates in model retraining as a supervisory signal: This mechanism ensures that critical outputs in highly sensitive scenarios are subject to human oversight to avoid systemic biases.

[0084] The static weight parameters, cross-modal alignment threshold, risk classification threshold, dynamic weight factor, scoring function parameters, and correction operator parameters mentioned above are all summarized to generate a dynamic weight adjustment parameter set, namely the confidence-rule weight dynamic coupling model, which provides parameter support for subsequent content recognition, risk assessment, and automatic correction.

[0085] Combination Figure 5 As shown, the complete process of dynamic weight adjustment and risk assessment is as follows: First, input the model confidence level and rule static weights; second, calculate the dynamic adjustment factor; then, apply the joint scoring function to evaluate the risk level of the sample; next, determine whether the scoring result triggers high / medium risk; perform forced correction on high-risk samples, and use soft penalty and manual review for medium-risk samples; then update the dynamic weights through a Bayesian update strategy; finally, output the adjusted weights and scoring results to provide accurate data support for sample correction and model iteration.

[0086] Step 203: Invoke the built-in compliance constraint clauses of the rule-constrained knowledge graph, and combine them with the hierarchical control specifications of the hierarchical rule verification system to compare the compliance content of each sample to be verified and generate a compliance verification list.

[0087] In this embodiment of the invention, the rule-constrained knowledge graph constructed in step 201 is invoked. Executable rule set Based on the hierarchical control specifications of the hierarchical rule verification system established in step 202, the core judgment unit—the rule verification function—is called. For the samples to be verified generated in step 201 (large language model based on preprocessed feature vector set) H Each generated candidate sample is compared against the compliance criteria: when When the sample is deemed to comply with the rule constraints and has no violations, it is determined that the sample conforms to the rule constraints and there are no violations. when If the sample is determined to have triggered a constraint conflict or violation, it needs to proceed to the subsequent logic correction stage. When multiple rules or pieces of evidence generate logical conflicts, the system uses a consistency check operator. Make a judgment: like If all constraints are satisfied, then if conflicts exist... This triggers the conflict resolution algorithm. C Conflict minimization can be achieved through priority ranking, weighted voting, or minimum cost correction rules: .

[0088] During the comparison process, the risk classification threshold of the dynamic weight adjustment parameter set in step 202 is combined to record the compliance judgment result, triggered compliance clauses, mismatch location and mismatch type of each sample. All comparison data are summarized to generate a compliance verification list, which provides a compliance judgment basis for subsequent problem identification and logic correction.

[0089] Furthermore, compliance content comparison can also be assisted by a knowledge-driven reasoning architecture based on retrieval enhancement: Based on the Retrieval Enhanced Generation (RAG) architecture, matching clauses and compliance cases are retrieved from the inspection and supervision knowledge base and the legal case base. The retrieval results are then semantically matched with the samples to be verified, replacing the pure rule keyword matching method and improving the accuracy and interpretability of compliance verification.

[0090] Step 204: Retrieve multimodal business feature data from the cross-modal vector index library, and perform multi-source content comparison by combining the cross-modal alignment threshold corresponding to the dynamic weight adjustment parameter set, and generate a cross-modal consistency verification list.

[0091] In this embodiment of the invention, the cross-modal vector index library constructed in step 201 is retrieved. Based on the feature vector of the sample to be verified (multimodal unified semantic features generated by the embedding mapping function Γ), multimodal business feature data with a semantic similarity not lower than the cross-modal alignment threshold configured in the dynamic weight adjustment parameter set in step 202 is retrieved. Cross-modal alignment is achieved through a bidirectional attention mechanism, and the attention calculation formula is as follows: ; In the formula, For text modal feature vectors, For image / OCR modal feature vectors, The feature dimension is used to record the differences, deviations, and alignment results of multiple sources such as text, OCR images, tables, and audio during the comparison process. All comparison data are summarized to generate a cross-modal consistency checklist, which solves the problems of insufficient cross-modal evidence alignment and evidence breakage, and improves the evidence chain integrity index (ECI) from 0.59 to 0.94.

[0092] Furthermore, cross-modal consistency verification can also be implemented using an evidence chain reconstruction architecture based on multimodal fusion: Constructing a multimodal consistency verification function: ; in, (This is the final output value of the multimodal consistency verification function), used to measure whether the evidence from different modalities in the same event is logically consistent and mutually corroborative. Sub-score for text modality consistency; Image modal consistency sub-score; For audio modal consistency sub-scores, The weights are for text, image, and audio modalities, respectively, and the sum of the three is always 1. The system can flexibly allocate weights according to the specific case type. This scheme ensures the integrity and robustness of the reasoning chain through cross-verification of multimodal evidence, and is especially suitable for cases that require cross-modal verification, such as "fake invoices" and "bid rigging".

[0093] Step 205: Based on the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, identify the illegal content, business logic conflict content, and cross-modal deviation content within the sample to be verified, and generate a list of issues to be corrected.

[0094] Furthermore, step 205 may include the following sub-steps: S31. Call the static weight parameters and risk classification thresholds in the dynamic weight adjustment parameter set, and combine them with the compliance judgment content recorded in the compliance verification checklist to identify non-compliant segments in the sample to be verified item by item, and generate compliance violation labeling data.

[0095] In this embodiment of the invention, the static weight parameters and risk classification thresholds in the dynamic weight adjustment parameter set of step 202 are called, and combined with the compliance judgment content of the compliance verification checklist in step 203, the non-compliant fragments in the sample to be verified are identified item by item through the confidence-rule weight joint scoring function, and the violation type, triggering clause, risk level and unmatched position are marked to generate compliance violation marking data, thereby completing the accurate identification of violation content.

[0096] S32. Based on the judgment criteria of risk classification threshold, and combined with the multi-source content differences recorded in the cross-modal consistency verification checklist, the multi-modal business content difference segments in the sample to be verified are screened item by item to generate cross-modal deviation identification data.

[0097] In this embodiment of the invention, based on the risk classification threshold discrimination standard of the dynamic weight adjustment parameter set in step 202, and combined with the multi-source content differences recorded in the cross-modal consistency verification checklist in step 204, the modal difference segments of text, OCR images, audio, and tables in the sample to be verified are screened item by item, the deviation type, deviation degree, and misalignment position are marked, cross-modal deviation identification data is generated, and cross-modal deviation content identification is completed.

[0098] S33. Based on the business flow specifications and hierarchical rule verification system control logic of the multi-level business execution hierarchy in the rule-constrained knowledge graph, combined with compliance and violation labeling data and cross-modal deviation identification data, the system identifies business connection contradictions in the samples to be verified and generates business logic conflict identification data.

[0099] In this embodiment of the invention, based on the business flow specifications of the multi-level business execution hierarchy in the rule-constrained knowledge graph and the control logic of the hierarchical rule verification system in step 202, combined with the compliance and violation labeling data generated in S31 and the cross-modal deviation identification data generated in S32, the business connection contradictions such as process sequence errors, business subject conflicts, and hierarchical logical contradictions in the sample to be verified are investigated, and business logic conflict identification data is generated to complete the identification of logical conflict content.

[0100] S34. Summarize compliance and non-compliance labeling data, cross-modal deviation identification data, and business logic conflict identification data, classify and organize them according to problem type, unify the rectification basis for the corresponding problems, and compile a list of problems to be corrected.

[0101] In this embodiment of the invention, the compliance and violation labeling data generated in S31, the cross-modal deviation identification data generated in S32, and the business logic conflict identification data generated in S33 are summarized and organized according to three categories of problems: violation content, business logic conflict, and cross-modal deviation. The rectification basis (including rule ID, business level, graph entity, and control specification) corresponding to each problem is uniformly labeled, and a list of problems to be corrected is compiled.

[0102] Step 206: Based on the hierarchical control requirements of the hierarchical rule verification system, resolve the logical conflicts in the list of issues to be corrected by resolving the rule conflicts, and simultaneously combine the multi-level business execution hierarchy to calibrate business logic deviations and generate logical calibration data.

[0103] In this embodiment of the invention, based on the hierarchical control requirements of the hierarchical rule verification system established in step 202, the conflict resolution strategy of the rule-constrained knowledge graph is invoked. C Introducing a conflict detection operator (in (This indicates the existence of a logical conflict, triggering a conflict resolution algorithm.) The algorithm resolves the logical conflicts in the list of issues to be corrected by using a rule-based conflict resolution method where higher-priority rules override lower-priority rules. Simultaneously, it combines the business flow specifications of the multi-level business execution hierarchy in the rule-constrained knowledge graph to calibrate logical deviations such as business process sequence, business subject, and hierarchical connection. The conflict resolution results and logical calibration content are then summarized to generate logical calibration data.

[0104] Step 207: Based on the constraints of the cross-modal alignment threshold, standardize the description of the cross-modal deviation content in the list of issues to be corrected, regulate and adjust the non-compliant content, and generate compliant correction data.

[0105] In this embodiment of the invention, according to the constraint requirements of the cross-modal alignment threshold corresponding to the dynamic weight adjustment parameter set in step 202, the cross-modal deviation content in the list of issues to be corrected is standardized, the multimodal expression format and semantic caliber are unified, and the consistency of expression among text, OCR images, tables, and audio is ensured; based on the compliance constraint clauses of the rule-constrained knowledge graph, the non-compliant content is rectified and adjusted, the rule-driven correction operator is called to execute the automatic correction logic, and the deviation standardization results and non-compliant content adjustment results are summarized to generate compliance correction data. The final output result is bound to the cross-modal evidence chain: ,in, For knowledge graph entities, For constraint relationships, This corresponds to the multimodal evidence fragments. This step addresses text, OCR documents, tables, images, and audio evidence using a unified embedding mapping function: Map features from different modalities to a unified semantic space. and with the knowledge graph entity set E Alignment enables cross-modal evidence fusion and compliance verification; the system binds a triplet-level evidence chain (entity, relation, evidence) to the output and records process logs to ensure that the results are traceable and verifiable throughout the entire chain.

[0106] It is worth mentioning that throughout the entire process of this invention, the system logs the generation results, rule triggering, correction steps, and manual interventions, forming a traceable chain of evidence: , A collection of structured tracing logs generated by the system; y The final output conclusion or text generated by the large language model for the current task; For the system to generate content y The status results obtained after performing compliance verification or consistency determination (such as compliance score, whether a corrective action has been triggered, etc.). A unique identifier (ID) is used to generate the core constraint rule of the knowledge graph that triggered or is based on this conclusion, in order to trace back the source of the rule; A unique identifier (ID) for the cross-modal evidence chain fragments supporting the conclusion is used to anchor the conclusion to the original input (such as text, image, audio recording); The precise system timestamp generated for this log entry ensures the timeliness and tamper-proof traceability of the evidence. This log can be used for post-event compliance checks and as metadata support during retraining, ensuring interpretability and compliance throughout the entire process.

[0107] Step 208: Based on the logical calibration data and compliance correction data, the sample content of the sample to be verified is uniformly standardized to generate a corrected sample.

[0108] In this embodiment of the invention, the logical calibration data generated in step 206 and the compliance correction data generated in step 207 are integrated to uniformly regulate all the content of the sample to be verified, eliminate illegal content, business logic conflicts, and cross-modal deviations, unify the sample format, semantic caliber and business logic, and generate a corrected sample that meets the requirements of high compliance, logical consistency and cross-modal integrity.

[0109] Step 209: Import the corrected samples into the preset training resource library and perform incremental iterative updates, associate and bind the corresponding multimodal supporting evidence data in the multimodal original business data, and generate the target inspection and patrol dataset.

[0110] In this embodiment of the invention, based on the closed-loop feedback incremental training framework, the corrected samples are imported into the preset training resource library and incremental iterative updates are performed. At the same time, the rule-constrained knowledge graph and the cross-modal vector index library are combined to associate and bind the corresponding multimodal supporting evidence data in the original multimodal business data, thereby generating a target inspection and patrol dataset with full-link traceability.

[0111] Furthermore, step 209 may include the following sub-steps: S41. The corrected samples are standardized and formatted according to the hierarchical control specifications of the stratified verification system, and then imported into the preset training resource library to generate a sample archive record table.

[0112] In this embodiment of the invention, the corrected samples generated in step 208 are verified according to the hierarchical control specifications of the hierarchical rule verification system in step 202. The sample format is unified, the business level and compliance level are marked, and the samples are imported into the preset training resource library. The sample number, generation time, compliance status and corresponding business level are recorded, and a sample archiving record table is generated to realize the standardized archiving management of corrected samples.

[0113] S42. Call the dynamic weight adjustment parameter set, and in conjunction with the sample archive record table, perform incremental iterative updates on the existing samples and newly added corrected samples in the preset training resource library, and generate an incremental update log for the resource library.

[0114] In this embodiment of the invention, the corrected sample With trigger log Stored in the training data pool D That is, a pre-defined training resource library drives the incremental fine-tuning process: based on a closed-loop feedback incremental training framework, a LoRA low-rank adapter is used to perform incremental fine-tuning, and the incremental update formula is: In the formula, These are the updated model parameters; These are the model parameters before the update. The learning rate; For rule compliance regular expressions; This represents the regularization term weight coefficient. This step employs a LoRA low-rank adapter to reduce computational overhead while ensuring compliance, achieving a closed-loop optimization of "generation-verification-correction-retraining". Simultaneously, it calls the dynamic weight adjustment parameter set from step 202, and, combined with the sample archive record table generated in S41, performs incremental iterative updates on the existing samples and newly added corrected samples in the preset training resource library, updating the dynamic weight factors, recording the update content, update time, and update basis, and generating an incremental update log for the resource library, thus achieving synchronous optimization of the dataset and the model.

[0115] Furthermore, incremental iterative updates can also be implemented using a parameter-efficient fine-tuning architecture based on low-rank decomposition: in this scheme, the Transformer backbone network remains frozen, and only the attention layer... With feedforward layer A low-rank adapter (LoRA) is introduced. Let the pre-trained weight matrix be... Through decomposition, we obtain: ; in, The updated weight matrix actually used in the inference phase after introducing the low-rank adapter; This is the original frozen pre-trained weight matrix for the large language model; The amount of weight update learned is used for incremental fine-tuning of the model. A and B These are two low-rank trainable matrices used to approximate high-dimensional update quantities; d and k The original pre-trained weight matrix Input and output dimensions; r Let be the rank of the low-rank adapter, satisfying ;rank r It can be flexibly configured according to different tasks (policy interpretation, risk detection, rectification tracking), and can reduce the number of training parameters by more than 90%. This solution has the advantages of rapid adaptation and low computing power consumption, and is suitable for real-time deployment at power grid terminal nodes and edge devices.

[0116] It is worth mentioning that incremental iterative updates can also employ a knowledge distillation mechanism to optimize the lightweight model. The distillation loss function is: ; in, CE This represents the cross-entropy loss, used to align the hard labels of the teacher and student models; KL This represents the Kullback-Leibler divergence, used to align the soft label distributions of the teacher and student models; The output distribution of the teacher model; The output distribution of the student model; This loss function enables the transfer of knowledge from large models to lightweight models, ensuring the compliance and accuracy of the lightweight models.

[0117] S43. Based on rule-constrained knowledge graphs, match the original inspection business fragments corresponding to the corrected samples within the multimodal original business data to generate a supporting evidence matching relationship table.

[0118] In this embodiment of the invention, the entity linking logic is based on the rule-constrained knowledge graph: , To finally match and bind the target entity of the knowledge graph; Candidate entities in the knowledge graph; E It is the set of all entities in the knowledge graph; The first part is the result of preprocessing the input data. i A multimodal feature vector (or semantic code); Similarity calculation functions (such as cosine similarity) are used to measure multimodal feature vectors. Semantic distance between the knowledge graph candidate entity e and the unified semantic space; This is the similarity threshold for entity links (alignment). Alignment is considered successful only when the maximum similarity is greater than or equal to this threshold, thus filtering out incorrect bindings with low confidence. Indicates in set E Find the similarity function sim The entity parameter that yields the maximum value. Combining the semantic retrieval results based on feature vectors in the cross-modal vector index library, the original inspection business fragments (including text, OCR images, audio, and tabular original evidence) corresponding to the corrected samples in the multimodal original business data are matched to clarify the correspondence between the samples and the original fragments, and a supporting evidence matching table is generated.

[0119] S44. Based on the hierarchical rule verification system and the supporting evidence matching relationship table, the corrected sample is bidirectionally associated and bound with the corresponding multimodal supporting evidence data in the original multimodal business data to generate a sample supporting evidence binding list.

[0120] In this embodiment of the invention, the control logic of the hierarchical rule verification system in step 202 and the supporting evidence matching relationship table generated in S43 are combined to bidirectionally associate and bind the corrected sample with the multimodal supporting evidence data in the original multimodal business data, establish a one-to-one correspondence between "sample-evidence", bind the triplet-level evidence chain (entity, relation, evidence) (where entity is a knowledge graph entity, relation is the business relationship between entities, and evidence is multimodal supporting evidence), and generate a sample supporting evidence binding list.

[0121] Furthermore, the binding of samples and evidence can also be achieved using an evidence chain reconstruction architecture based on multimodal fusion: the multimodal joint fusion features and consistency verification results are incorporated into the evidence chain binding logic to form an evidence chain structure with multiple sources of text, images, and audio for mutual verification, ensuring the complete closure of the evidence chain and improving the success rate and traceability of complex cases.

[0122] S45. Conduct compliance verification on the sample supporting binding list, remove invalid binding content from the list, and generate a binding verification report.

[0123] In this embodiment of the invention, the binding accuracy and evidence completeness of the sample supporting binding list generated in S44 are verified, invalid binding content with incorrect correspondence, missing evidence, or no business significance is removed, the verification results, reasons for invalid binding, and correction content are recorded, and a binding verification report is generated.

[0124] S46. According to the business association logic of the rule-constrained knowledge graph, classify and organize the sample archive record table, resource library incremental update log, sample evidence binding list and binding verification report to generate the target inspection and patrol dataset.

[0125] In this embodiment of the invention, according to the business association logic of the rule-constrained knowledge graph, the sample archive record table generated in S41, the resource library incremental update log generated in S42, the sample supporting binding list generated in S44, and the binding verification report generated in S45 are classified and organized according to business level and business type; a quadruple structure with a triplet-level traceability mechanism is adopted. In the formula, c This is the conclusion. EID For the chain of evidence ID, RID For the rule trigger ID, This represents the model version number; the logs are stored using a Merkle tree hash chain to ensure data immutability, ultimately generating a target inspection and patrol dataset with full-chain traceability and high confidence. The Merkle tree-based log hash chain storage ensures immutability, and evidence tracing takes less than 200ms.

[0126] Combination Figure 5 As shown, the overall technical process of this invention forms a standardized iterative pipeline, which consists of rule ontology modeling → real-time verification and correction → joint scoring-driven dynamic adjustment → incremental training closed loop → multimodal evidence fusion. Each sub-process completely matches the steps of this embodiment: the rule ontology modeling stage performs entity attribute definition and relation rule construction; the real-time verification and correction stage performs conflict detection, logical reasoning, and entity repair; the joint scoring-driven dynamic adjustment stage performs confidence calculation and dynamic weight allocation; the incremental training closed loop stage performs feedback data collection and model parameter update; and the multimodal evidence fusion stage performs text parsing, image feature extraction, and knowledge association. The overall process supports iterative training and correction, and the corrected data can be fed back to the rule ontology modeling stage to achieve long-term optimization of the dataset and model.

[0127] In the field of dataset construction for compliance verification scenarios, existing technologies still have significant limitations, making it difficult to support the construction of datasets with high credibility and strong traceability. On the one hand, the sample generation and compliance verification stages often adopt a generation-verification separation architecture, and its process can be represented as follows: ,in, This is the initial data source or input instruction set generated; For large language models or algorithm engines used to generate content, Refers to model parameters; The initial sample set generated for the model (the raw output sequence that has not been filtered for compliance). For rules based on static rule base or manually preset rules R The hysteresis check function; This is the set of valid and compliant samples that have been verified and screened; when the sample size... At that time, the cost of error correction increases exponentially with the increase of sample size, and the correction time meets the requirements. ,in,, Total time spent correcting for the sample; , which is the quadratic time complexity coefficient, representing the non-linear time consumption caused by the lack of graph-based ontology constraints, cross-sample rule comparison, and complex logical conflict investigation. is the time complexity coefficient for a single term, representing the linear time consumption when performing a basic traversal or keyword filtering on a single sample. On the other hand, multi-source heterogeneous data such as text, images, and tables typically employ independent feature extraction processes, lacking a unified semantic space mapping mechanism, and text embedding... Image features Ineffective alignment resulted in feature space conflicts in 41% of multimodal samples, affecting the chain of evidence integrity index. The total evidence for aligned evidence dropped to 0.59, making it difficult to support the construction of a complete and traceable chain of evidence.

[0128] To address the aforementioned issues, this invention proposes a four-layer technical framework, systematically constructing a solution covering the entire dataset construction process: First, a knowledge graph ontology constraint layer, constructing a rule-based knowledge graph to provide institutional-level constraints for compliance verification; second, a real-time compliance verification and logic correction layer, introducing rule verification functions. The system employs four main technical approaches: dynamic conflict identification and real-time error correction; a confidence-rule weight dynamic adjustment layer that calculates risk levels using a joint scoring function and achieves adaptive weight convergence using a Bayesian update formula; and a closed-loop feedback and multimodal fusion layer that optimizes the model through incremental training formulas and maps multi-source data to a unified semantic space using cross-modal embedding functions. The overall approach follows a technical path of "rule ontology modeling → real-time verification and correction → joint scoring-driven dynamic adjustment → incremental training closed loop → multimodal evidence fusion."

[0129] Through the above-mentioned technological innovations, this invention has achieved significant performance improvements: compliance rate increased by 39%, logical consistency improved by 46%, cross-modal evidence integrity rate increased by 54%, and manual intervention cost of datasets reduced by 73%. It effectively solves the core problems in existing technologies such as lagging rule conflict judgment, cross-modal evidence fragmentation, and low sample utilization efficiency, and can provide high-confidence, traceable, and standardized dataset construction capabilities for compliance verification scenarios.

[0130] Specifically, this embodiment uses the multimodal data processing of 237 engineering projects in a provincial power system as a scenario to verify the dataset construction method of the present invention: First, the collected multi-source heterogeneous data is preprocessed, covering 23,000 pages of PDF bidding documents, 1.5 million CSV fund flow records, 3,200 JPG contract scans, and 120 hours of bid evaluation audio recordings. The PaddleOCR engine (with a recognition accuracy of 98.7%) is used to convert the image files into structured text, and key entities such as the name of the bidding unit and the time of deposit payment are extracted using regular expressions (entity recognition F1-score=0.93). Simultaneously, a knowledge graph of "bidding unit-project-funds" containing 12,000 nodes (personnel / enterprises / projects) and 35,000 relationships is constructed. The relevant entity relationship triples are modeled using OWL2DL ontology, the rules are exported in SWRL format, the initial static weight is 0.8, and the dynamic weight is determined by the EWMA algorithm ( =0.3) Iterative updates; Based on the constructed rule ontology library, the Few-shot rule reasoning framework is called to break down the dataset construction task into a four-stage process of "entity extraction → rule matching → conflict resolution → evidence fusion", and five historical non-compliant cases are loaded as guiding examples. The LoRA low-rank adapter (rank) is used to update the dataset. r =16, =16) Fine-tune the Qwen2.5-7B model, using the HNSW index (M=32, =200) Construct an evidence retrieval library, perform Top-k=5 nearest neighbor retrieval to integrate relevant regulations and similar cases into the reasoning chain, and generate an interpretable sequence of intermediate steps; during the process, data quality is monitored in real time through a joint scoring function, and when abnormal data features are detected, the knowledge graph completion mechanism is automatically triggered to supplement related data, and finally outputs a standardized structure containing correction suggestions, evidence chain ID and model version number, which increases the automatic resolution rate of rule conflicts to 91% and the cross-modal evidence chain integrity index ECI from 0.62 to 0.95.

[0131] This invention employs a heterogeneous architecture of "server-side training + edge-side inference." The training environment is configured with a single-node NVIDIA A100 80GB GPU (CUDA 12.2), and distributed training is implemented through PyTorch Lightning. The AdamW optimizer (learning rate 3e-4, weight decay 0.02) iterates for 15k steps, using a LoRA low-rank adapter (rank r=16). =16, dropout=0.1) freezes 98.2% of the backbone model parameters; the inference terminal is equipped with an Intel i7-13700H CPU and 32GB DDR5 memory, and the model size is compressed to 2.8GB through GPTQ 4-bit quantization (per-channel quantization, zero-point optimization), and inference latency is optimized (p95<750ms) by combining TensorRT operator fusion technology; key technical parameters strictly follow industry standards, the knowledge graph adopts Neo4j graph database (supports 1 billion relations on a single machine), and the vector retrieval engine is configured with M=32, =200 runtime parameters, enabling EWMA update strategy for the dynamic weighted module ( =0.3) and dynamic optimization of scoring threshold ( =0.65); verified by third-party testing, the rule triggering accuracy was 0.97 and the evidence chain integrity ECI was 0.95 in this case, which is 23 times more efficient than traditional manual annotation, the multimodal evidence utilization rate reached 98.3%, and the dataset annotation cost was reduced by 76%.

[0132] The verification results of this embodiment demonstrate the comprehensive performance advantages of the five major technological innovations of this invention, achieving a 43% improvement in compliance verification accuracy (from 0.68 to 0.97), a 59% improvement in evidence chain integrity (from 0.59 to 0.94), and a 4.2-fold increase in training efficiency. This effectively solves problems in existing technologies such as the inability to automatically determine rule conflicts, cross-modal evidence breakage, and low sample utilization efficiency, meeting the intelligent requirements of high reliability, high efficiency, and traceability in power engineering project compliance verification scenarios. In terms of rule modeling, this embodiment uses the OWL2DL ontology language to formalize project policy clauses into a six-tuple knowledge graph. Logical reasoning is achieved through the SWRL rule engine, replacing the traditional text matching method. This improves the accuracy of automatic rule conflict determination from 0.68 to 0.97, reduces the non-compliance rate from 28% to 9%, and increases rule derivation efficiency by 8 times, solving the problem of resolving priority conflicts in policy clauses at different levels. In terms of dynamic decision-making, a confidence-rule weight joint scoring function and an EWMA dynamic weight update strategy are used, non-linearly coupled with the model confidence, replacing the static threshold. The value-based approach improves the high-risk output recognition rate from 0.75 to 0.92, reduces error correction costs by 67%, and dynamically optimizes the scoring threshold to 0.65 using ROC curves, achieving self-calibration of risk decisions. Regarding the training framework, an iterative pipeline of "corrected sample write-back - LoRA fine-tuning" is constructed, employing a low-rank adapter with rank 16 to achieve efficient parameter updates, replacing the full fine-tuning mode. Sample reuse rate increases from 0.3 to 0.89, convergence speed is accelerated by 4.2 times, training time is reduced from 120 hours to 28.6 hours, and manual annotation costs are reduced by 73%, meeting the needs of rapid iteration.

[0133] In terms of multimodal fusion, this embodiment constructs a unified semantic space mapping through contrastive learning and achieves multimodal alignment by combining a bidirectional attention mechanism. It integrates OCR invoice information, bank transaction text, and project-related voice data into a unified knowledge network. The Evidence Chain Integrity Index (ECI) increases from 0.59 to 0.94, the cross-modal conflict rate decreases by 82%, and the processing efficiency of complex projects improves by 37%. In terms of process link traceability, a four-tuple output structure containing key information, evidence chain identifier, rule identifier, and result identifier is designed. Merkle tree hash chain is used to store process operation logs. Edge deployment is achieved by combining INT8 quantization and TensorRT optimization, which reduces the evidence traceability time from 3000ms to 800ms, controls the inference latency within 800ms, and achieves 100% evidence citation accuracy, fully meeting the requirements of compliance verification scenarios for process traceability.

[0134] Please see Figure 6 , Figure 6 This is a structural block diagram of an inspection and patrol dataset construction device provided in an embodiment of the present invention.

[0135] This invention provides a device for constructing inspection and patrol datasets, comprising: The basic data construction module 601 is used to acquire multimodal raw business data corresponding to inspection and patrol business, and to construct a rule-constrained knowledge graph, a cross-modal vector index library and samples to be verified based on the multimodal raw business data; The rule parameter construction module 602 is used to build a hierarchical rule verification system and a dynamic weight adjustment parameter set based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph. The sample compliance processing module 603 is used to perform compliance verification and logical correction on the sample to be verified based on the rule-constrained knowledge graph, cross-modal vector index library, hierarchical rule verification system and dynamic weight adjustment parameter set, so as to obtain the corrected sample. The dataset integration and generation module 604 is used to import the corrected samples into the preset training resource library and perform incremental iterative updates, associate and bind the corresponding multimodal supporting evidence data in the multimodal original business data, and generate the target inspection and patrol dataset.

[0136] Since the above is a device corresponding to a method for constructing an inspection and patrol dataset, and its implementation principle is consistent with a method for constructing an inspection and patrol dataset, for the sake of convenience and brevity, those skilled in the art can clearly understand that the specific working process of the device and module described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here.

[0137] Please see Figure 7 , Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention.

[0138] An electronic device according to an embodiment of the present invention includes: a memory 701 and a processor 702. The memory 701 stores a computer program. When the computer program is executed by the processor 702, the processor 702 executes the patrol and inspection dataset construction method as described in any of the above embodiments.

[0139] Memory 701 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 701 has storage space 703 for program code 713 for performing any of the method steps described above. For example, storage space 703 for program code may include various program codes 713 for implementing the various steps in the methods described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When this code is run by a computing device, it causes the computing device to perform the various steps in the inspection and patrol dataset construction method described above.

[0140] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the patrol and inspection dataset construction method as described in any of the above embodiments.

[0141] This invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer performs the patrol and inspection dataset construction method as described in any of the above embodiments.

[0142] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0143] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0144] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0145] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0146] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0147] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing an inspection and patrol dataset, characterized in that, include: Obtain multimodal raw business data corresponding to inspection and patrol operations, and construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data; Based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph, a hierarchical rule verification system and a dynamic weight adjustment parameter set are built. Based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set, the sample to be verified is subjected to compliance verification and logical correction to obtain a corrected sample. The corrected samples are imported into a preset training resource library and incremental iterative updates are performed. The corresponding multimodal supporting evidence data in the multimodal original business data are associated and bound to generate the target inspection and patrol dataset.

2. The method for constructing inspection and patrol datasets according to claim 1, characterized in that, The steps of acquiring multimodal raw business data corresponding to inspection and patrol operations, and constructing a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data include: Collect text, table, image, optical character recognition document, audio and video data corresponding to inspection and patrol business, and generate multimodal raw business data; The original multimodal business data is sequentially cleaned, classified into modes, extracted into business elements, and extracted into multimodal features to generate standardized multimodal business data. The inspection business elements, business association logic, and compliance constraint clause attributes are extracted from the standardized multimodal business data. A rule-constrained knowledge graph is constructed by combining triple structured orchestration with ontology modeling. The heterogeneous features in the standardized multimodal business data are mapped to a unified semantic space to obtain a unified semantic space feature vector, and a cross-modal vector index library is constructed based on the unified semantic space feature vector. Extract the valid inspection business segments from the standardized multimodal business data and organize and integrate them to generate a sample to be verified.

3. The method for constructing inspection and patrol datasets according to claim 1, characterized in that, The steps of building a hierarchical rule verification system and a dynamic weight adjustment parameter set based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph include: Based on the business flow connection form and business coverage of the inspection business association logic in the rule-constrained knowledge graph, a multi-level business execution hierarchy is defined. Based on the compliance constraint clause data in the rule-constrained knowledge graph, the compliance constraint clauses corresponding to each business execution level are matched, and combined with the applicable scenarios and control scope of the compliance constraint clauses, a hierarchical rule verification system is formed through hierarchical integration. Based on the complexity of the business process layout recorded in the aforementioned inspection business association logic, risk differentiation criteria are used to distinguish different businesses. Based on the constraint strength identifier built into the compliance constraint clause data, the control level of different compliance constraint clauses is defined; Based on the aforementioned business risk differentiation criteria and clause control levels, the verification order of each compliance constraint clause in the aforementioned compliance constraint clause data is uniformly arranged to determine the rule priority; Based on the rule priority, static weight parameters are matched for each of the compliance constraint clauses, cross-modal alignment thresholds are configured for the cross-modal content verification stage, risk classification thresholds are configured for the business risk identification stage, and all configured parameters and thresholds are summarized and integrated to generate a dynamic weight adjustment parameter set.

4. The method for constructing inspection and patrol datasets according to claim 1, characterized in that, The step of performing compliance verification and logical correction on the sample to be verified based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set to obtain a corrected sample includes: The compliance constraint clauses built into the rule-constrained knowledge graph are invoked, and the hierarchical control specifications of the hierarchical rule verification system are combined to compare the compliance content of each sample to be verified, generating a compliance verification list; The system retrieves multimodal business feature data from the cross-modal vector index library and performs multi-source content comparison by combining the cross-modal alignment threshold corresponding to the dynamic weight adjustment parameter set, thereby generating a cross-modal consistency verification list. Based on the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, the illegal content, business logic conflict content, and cross-modal deviation content within the sample to be verified are identified, and a list of issues to be corrected is generated. Based on the hierarchical control requirements of the hierarchical rule verification system, the logical conflict content in the list of issues to be corrected is resolved by rule conflict resolution, and the business logic deviation is simultaneously calibrated by the multi-level business execution hierarchy to generate logical calibration data. Based on the constraints of the cross-modal alignment threshold, the descriptions of cross-modal deviations in the list of issues to be corrected are standardized, the violations are rectified and adjusted, and compliant correction data is generated. Based on the logical calibration data and the compliance correction data, the sample content of the sample to be verified is uniformly standardized to generate a corrected sample.

5. The method for constructing inspection and patrol datasets according to claim 4, characterized in that, The step of identifying violations, business logic conflicts, and cross-modal deviations within the sample to be verified, and generating a list of issues to be corrected, based on the dynamic weight adjustment parameter set, the multi-level business execution hierarchy recorded in the rule-constrained knowledge graph, the compliance verification list, and the cross-modal consistency verification list, includes: The static weight parameters and risk classification thresholds in the dynamic weight adjustment parameter set are called, and the compliance judgment content recorded in the compliance verification checklist is combined to identify non-compliant segments in the sample to be verified item by item, and generate compliance violation labeling data. Based on the discrimination criteria of the risk classification threshold, and combined with the multi-source content differences recorded in the cross-modal consistency verification list, the multi-modal business content difference segments in the sample to be verified are screened item by item to generate cross-modal deviation identification data. Based on the business flow specifications and hierarchical rule verification system control logic of the multi-level business execution hierarchy in the rule-constrained knowledge graph, combined with the compliance and violation labeling data and the cross-modal deviation identification data, the business connection contradictions in the samples to be verified are investigated, and business logic conflict identification data is generated. The compliance and violation labeling data, the cross-modal deviation identification data, and the business logic conflict identification data are summarized, categorized and organized according to problem type, and the corresponding rectification basis for the labeled problems is unified to compile a list of problems to be corrected.

6. The method for constructing inspection and patrol datasets according to claim 1, characterized in that, The steps of importing the corrected samples into a preset training resource library and performing incremental iterative updates, associating and binding the corresponding multimodal supporting evidence data in the multimodal original business data, and generating the target inspection and patrol dataset include: The corrected samples are standardized and formatted according to the hierarchical management and control specifications of the hierarchical rule verification system, and then imported into the preset training resource library to generate a sample archive record table. The dynamic weight adjustment parameter set is invoked, and combined with the sample archive record table, to perform incremental iterative updates on the existing samples and newly added correction samples within the preset training resource library, generating an incremental update log for the resource library. Based on the rule-constrained knowledge graph, the original inspection business segments corresponding to the corrected samples within the multimodal original business data are matched to generate a supporting evidence matching relationship table. Based on the hierarchical rule verification system and the supporting evidence matching relationship table, the corrected sample is bidirectionally associated and bound with the corresponding multimodal supporting evidence data in the multimodal original business data to generate a sample supporting evidence binding list. The sample supporting binding list is subjected to compliance verification, invalid binding content within the list is removed, and a binding verification report is generated; Based on the business association logic of the rule-constrained knowledge graph, the sample archive record table, the resource library incremental update log, the sample supporting binding list, and the binding verification report are classified and organized to generate the target inspection and patrol dataset.

7. A device for constructing inspection and patrol datasets, characterized in that, include: The basic data construction module is used to acquire multimodal raw business data corresponding to inspection and patrol business, and to construct a rule-constrained knowledge graph, a cross-modal vector index library, and samples to be verified based on the multimodal raw business data. The rule parameter construction module is used to build a hierarchical rule verification system and a dynamic weight adjustment parameter set based on the business association logic and compliance constraint clause data in the rule-constrained knowledge graph. The sample compliance processing module is used to perform compliance verification and logical correction on the sample to be verified based on the rule-constrained knowledge graph, the cross-modal vector index library, the hierarchical rule verification system, and the dynamic weight adjustment parameter set, so as to obtain the corrected sample. The dataset integration and generation module is used to import the corrected samples into a preset training resource library and perform incremental iterative updates, and associate and bind the corresponding multimodal supporting evidence data in the multimodal original business data to generate the target inspection and patrol dataset.

8. An electronic device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the patrol and inspection dataset construction method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the patrol and inspection dataset construction method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer performs the patrol and inspection dataset construction method as described in any one of claims 1-6.