Data auditing method and device, electronic equipment and nonvolatile storage medium
By building a multi-modal knowledge base and auditing model, multi-level verification of cross-modal data is solved, and the problem of insufficient semantic association and logical relationship processing capabilities in cross-modal data audit is realized, and comprehensive accurate verification and high-quality management of cross-modal data is realized.
Patent Information
- Application Number
- CN202510586190.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art is difficult to capture the semantic associations and complex logical relationships of cross-modal data in cross-modal data audits, resulting in poor content quality inspection results.
By building a multimodal knowledge base and auditing model, combining verification rule sets and semantic reasoning capabilities, multi-level verification of cross-modal data is performed, exception reports are generated and correction strategies are provided.
It realizes comprehensive and accurate verification of cross-modal data, improves the scientificity and consistency of content quality inspection, and ensures high-quality processing and management of data.
Smart Images

Figure CN120541537A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data management technology, and more specifically, to a data auditing method, device, electronic device, and non-volatile storage medium. Background Art
[0002] Cross-modal data typically consists of different modalities (e.g., text, images, video, speech, etc.) and carries metadata, which can provide rich contextual information and multimodal relationships. With the development of data fusion and integration technologies, the amount of data with cross-modal characteristics has exploded.
[0003] In the field of data management, verifying the content quality of cross-modal data is crucial. However, related technologies, when conducting data audits, primarily focus on single-modal data or perform simple cross-modal data processing. This makes it difficult to capture semantic associations within cross-modal data and lacks the ability to process complex logical relationships, resulting in poor results in cross-modal data content quality verification.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a data auditing method, device, electronic device and non-volatile storage medium to at least solve the technical problem that the related art is not effective in verifying the content quality of cross-modal data.
[0006] According to one aspect of an embodiment of the present application, a data auditing method is provided, comprising: obtaining cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; determining a verification rule set corresponding to a multimodal knowledge base, and performing a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base at least includes professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set is used to indicate constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data contains anomalies that do not comply with the rules in the verification rule set; using an audit large model, performing a second verification process on the cross-modal data to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning ability of the large model; determining an exception report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the exception report is used at least to characterize abnormal data in the cross-modal data.
[0007] Optionally, after obtaining the cross-modal data to be audited, the method further includes: using a feature separation network to separate the features of different modalities in the cross-modal data to obtain modal features, wherein the modal features include at least one of the following: texture features, color features, semantic features, grammatical features, spectral features, and phoneme features; performing structured extraction on the modal features to obtain key metadata information corresponding to the modal features, wherein the key metadata information is used to characterize the semantic information and format attribute information of the data content; constructing a multimodal data cube corresponding to the cross-modal data by normalizing the key metadata information, wherein the normalization process is used to unify the data format and dimension of the key metadata information.
[0008] Optionally, determining the verification rule set corresponding to the multimodal knowledge base includes: determining the entity objects contained in the knowledge information in the multimodal knowledge base, and determining the association relationships between the entity objects; determining the ontology constraints corresponding to the entity objects and / or association relationships by analyzing the knowledge information, wherein the ontology constraints are used to characterize the reasonable constraints that the entity objects and / or association relationships should comply with; based on the ontology constraints, constructing the verification rule set corresponding to the multimodal knowledge base, wherein the verification rule set specifies the constraints that should be met when different association relationships exist between different entity objects.
[0009] Optionally, the first verification result is used to indicate abnormal data that does not meet the constraints corresponding to the verification rule set, and the second verification result is used to indicate abnormal data with logical contradictions detected by the audit model; based on the first verification result and the second verification result, determining the abnormal report corresponding to the cross-modal data includes: integrating the abnormal data corresponding to the first verification result and the second verification result, wherein the abnormal data includes at least one of the following: data inconsistent with facts or common sense, data with inconsistent descriptions of different modal data, data that does not match related data, and data missing key information; determining the abnormality type and abnormality cause corresponding to the abnormal data, wherein the abnormality type includes at least one of the following: semantic abnormality, structural abnormality, attribute abnormality, modal feature conflict abnormality, contextual abnormality, knowledge base conflict abnormality, and the abnormality cause includes at least one of the following: collection error, entry error, labeling error, reasoning defect; using a Bayesian network model, combined with the abnormality type and abnormality cause, to analyze the abnormal data, obtain the confidence corresponding to the abnormal data, and determine the abnormality degree corresponding to the abnormal data based on the confidence; generating an abnormality report based on the abnormal data, and the abnormality type, abnormality cause and abnormality degree corresponding to the abnormal data.
[0010] Optionally, the method also includes: determining historical abnormal data in historical abnormal data records whose similarity with abnormal data in abnormal reports exceeds a preset similarity threshold as target historical abnormal data; obtaining historical correction strategies corresponding to the target historical abnormal data; using an audit large model, combined with a multimodal knowledge base, to analyze the abnormal reports and historical correction strategies to obtain the target correction strategies corresponding to the abnormal data.
[0011] Optionally, the method also includes: sending the exception report and the target correction strategy corresponding to the exception data indicated in the exception report to the front-end interactive interface of the terminal device of the target object for display; in response to a first instruction triggered in the front-end interactive interface, correcting the exception data according to the target correction strategy, wherein the first instruction indicates that the target object agrees with the exception data and the target correction strategy indicated in the exception report; in response to a second instruction triggered in the front-end interactive interface, sending the exception report and the target correction strategy to the front-end interactive interfaces of the terminal devices of multiple other target objects for display, wherein the second instruction indicates that the target object does not agree with the exception data and the target correction strategy indicated in the exception report; obtaining the audit results returned by the terminal devices of multiple other target objects, and determining whether the exception data and the target correction strategy indicated in the exception report are correct based on the audit results.
[0012] Optionally, the method also includes: when the abnormal data indicated by the abnormal report is corrected, obtaining log data of the entire life cycle of analysis and processing of the abnormal data, wherein the log data includes at least one of the following: the original content of the abnormal data, the target correction strategy given by the model, the operating instructions and opinions of the target object, the audit results of other target objects, and the final correction results; and updating the log data as new knowledge information to the multimodal knowledge base.
[0013] According to another aspect of an embodiment of the present application, a data auditing device is further provided, comprising: a data acquisition module for acquiring cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; a first verification module for determining a verification rule set corresponding to a multimodal knowledge base, and performing a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base at least includes professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set is used to indicate constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data contains anomalies that do not comply with the rules in the verification rule set; a second verification module for performing a second verification process on the cross-modal data using an audit large model to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning capability of the large model; and an anomaly analysis module for determining an anomaly report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the anomaly report is used at least to characterize abnormal data in the cross-modal data.
[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the data audit method is executed when the program is run.
[0015] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided. The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data auditing method by running the computer program.
[0016] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, which implements the steps of the data auditing method when executed by a processor.
[0017] In an embodiment of the present application, cross-modal data to be audited is obtained, wherein the cross-modal data includes data of at least two different modalities; a verification rule set corresponding to a multimodal knowledge base is determined, and based on the verification rule set, a first verification process is performed on the cross-modal data to obtain a first verification result, wherein the multimodal knowledge base at least includes professional knowledge and specifications in the corresponding field of the cross-modal data to be audited, the verification rule set is used to indicate the constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data has anomalies that do not comply with the rules in the verification rule set; a second verification process is performed on the cross-modal data using an audit macro model to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning ability of the macro model; based on the first verification result and the second verification result, an exception report corresponding to the cross-modal data is determined, wherein the exception report is used to at least characterize the abnormal data in the cross-modal data. By combining the multimodal knowledge base and the audit macro model, the purpose of achieving comprehensive and accurate verification of the cross-modal data is achieved, thereby solving the technical problem that the related art has poor content quality inspection effect on cross-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 This is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a data audit method provided in an embodiment of the present application;
[0020] Figure 2 This is a schematic diagram of a data audit method flow according to an embodiment of the present application;
[0021] Figure 3 This is a schematic diagram of a method flow for cross-modal data content logic intelligent auditing based on artificial intelligence and knowledge base according to an embodiment of the present application;
[0022] Figure 4 It is a structural diagram of a data auditing device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] To facilitate those skilled in the art to better understand the embodiments of the present application, some technical terms or nouns involved in the embodiments of the present application are explained as follows:
[0026] Cross-modal data refers to heterogeneous data that includes multiple forms (such as text, images, audio, and video). In this patent, these different modal data are collected and processed comprehensively to perform content logic audits.
[0027] Large model: An artificial intelligence model with large-scale parameters, strong learning and generalization capabilities, capable of deep semantic understanding, reasoning, and generation of multi-modal data such as text, images, and videos.
[0028] Knowledge base: A structured or semi-structured data set containing knowledge in a specific field or multiple fields, providing background knowledge for logical verification (such as a geographic feature library, a medical diagnostic standard library).
[0029] Standardized data cube: A highly standardized data set constructed by unifying the dimensions and measurement definitions of multi-source data, using standardized processes to clean and align the collected cross-modal data, and storing it in a multi-dimensional structure.
[0030] Double-blind review mechanism: an anonymous review method in which the reviewee and the reviewer are unaware of each other's identities, aiming to ensure the fairness and objectivity of the review process.
[0031] Back-to-back review: refers to multiple experts independently reviewing the same object without knowing each other's opinions, and ultimately reaching a more objective and fair conclusion by comparing or integrating the opinions of all parties.
[0032] Logical auditing: Through semantic reasoning and rule verification, it detects logical contradictions in cross-modal data content (such as inconsistencies between geographic images and labels, conflicts between medical images and patient information), and ensures the internal consistency of the data.
[0033] The complexity of cross-modal data poses even greater challenges to data quality control and auditing. On the one hand, cross-modal data contains a vast amount of knowledge and information, and its semantics are rich, making data quality control often difficult to rely solely on a single modality. On the other hand, cross-modal data often involves cross-modal association of multimodal knowledge, cross-modal reasoning of multimodal semantics, and cross-modal fusion of multimodal knowledge.
[0034] Data auditing methods in related technologies primarily focus on single-modal data or simple cross-modal data processing. Specifically, while each method has its own distinct limitations in terms of single-modal data verification, text data auditing often relies on keyword matching and grammar checking. Image data auditing utilizes feature extraction and matching algorithms, similar to image search engines that compare feature vectors for classification retrieval. However, these methods only superficially verify single-modal data and are unable to address complex cross-modal associations. Simple cross-modal data processing attempts, such as adding text tags to images in multimedia databases and then using text search to find images, and manual or rule-based matching in the medical field to link images with medical record text, lack deep semantic understanding and are unable to parse the underlying logic between geographic image scenes and labeled text, or between medical image features and medical record symptoms. Some cross-modal retrieval technologies measure similarity through feature space mapping, but their accuracy is limited by semantic gaps and insufficient fine-grained feature mining.
[0035] In summary, the data auditing methods of related technologies mostly rely on manual or simple rules, have poor adaptability and scalability, lack deep semantic understanding, have difficulty capturing semantic associations across modal data, and lack the ability to handle complex logical relationships, making it impossible to effectively audit cross-modal data.
[0036] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.
[0037] According to an embodiment of the present application, a method embodiment of data auditing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following is a hardware block diagram of a computer terminal (or electronic device) for implementing a data audit method. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or electronic device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data auditing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned data auditing method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0041] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0042] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or electronic device).
[0043] In the above operating environment, the embodiment of the present application provides a data audit method. Figure 2 This is a schematic diagram of a data audit method flow according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0044] Step S202: obtaining cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities;
[0045] Step S204: Determine a verification rule set corresponding to the multimodal knowledge base, and perform a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result. The multimodal knowledge base contains at least professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, and the verification rule set indicates the constraints that the cross-modal data should comply with. The first verification process is used to verify whether the cross-modal data contains any anomalies that do not comply with the rules in the verification rule set.
[0046] Step S206: Using the audit macro model, perform a second verification process on the cross-modal data to obtain a second verification result. The second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning capability of the macro model.
[0047] Step S208 : determining an exception report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the exception report is at least used to characterize abnormal data in the cross-modal data.
[0048] Through the above steps, by combining the multimodal knowledge base and the audit model, the goal of comprehensive and accurate verification of cross-modal data is achieved, thereby solving the technical problem that related technologies are not effective in content quality verification of cross-modal data.
[0049] The following further introduces the data audit method in steps S202 to S208 of the embodiment of the present application.
[0050] Figure 3 This is a schematic diagram of a method flow for cross-modal data content logic intelligent audit based on artificial intelligence and knowledge base according to an embodiment of the present application, such as Figure 3 As shown, by jointly processing cross-modal data, building a multi-domain knowledge base and designing a rule set, and using a combination of rule-driven and large-scale model reasoning to verify the logical rationality of data at multiple levels, the system automatically flags anomalies and generates graded correction suggestions based on the test results. A human-machine collaborative review mechanism ensures the accuracy of data corrections and returns the correction results to the system. Furthermore, the system iteratively optimizes the dynamic knowledge base to achieve high-quality data processing, precise verification, and effective management, thereby improving the scientificity, rationality, and consistency of data standards. A detailed description is provided below.
[0051] First, we collect and process cross-modal data, separate its modal features and clean up noise, extract and process metadata in a structured manner, and construct a standardized data cube. The details are as follows.
[0052] In some embodiments of the present application, after obtaining the cross-modal data to be audited, the method further includes the following steps: using a feature separation network to separate the features of different modalities in the cross-modal data to obtain modal features, wherein the modal features include at least one of the following: texture features, color features, semantic features, grammatical features, spectral features, and phoneme features; performing structured extraction on the modal features to obtain key metadata information corresponding to the modal features, wherein the key metadata information is used to characterize the semantic information and format attribute information of the data content; constructing a multimodal data cube corresponding to the cross-modal data by normalizing the key metadata information, wherein the normalization process is used to unify the data format and dimension of the key metadata information.
[0053] Specifically, it uses a variety of data acquisition devices and interfaces (for example, API / SDK / IoT protocols, etc.) to collect cross-modal data from different sources, including text, images, audio, video, etc., supports the fusion access of structured and unstructured data, and has a built-in data quality pre-inspection mechanism to automatically filter invalid data such as format errors and insufficient resolution to ensure data diversity and comprehensiveness; it uses a feature separation network (such as CNN+Transformer architecture) to separate the modal features of cross-modal data, such as texture / color features in images, semantic / grammatical features in text, and spectrum / phoneme features in audio, to achieve preliminary data classification and make the data structure clearer.
[0054] At the same time, data noise and ambiguous expressions in cross-modal data are eliminated to construct a standardized multimodal data cube; a metadata extractor is built based on the pre-trained language model (BERT / GPT) to intelligently extract key metadata information such as time (UTC standard timestamp), space (WGS-84 geographic coordinates), and equipment (camera model / sensor accuracy), and store the metadata in a structured manner.
[0055] Then, based on the verification rule set, the cross-modal data is subjected to the first verification processing (hard verification). In this embodiment, a multi-domain multimodal knowledge base can be constructed first, and the large model technology can be used to intelligently parse the knowledge base content, extract entity relationships and constraints, and form an explainable reasoning rule set (i.e., verification rule set), as follows.
[0056] In some embodiments of the present application, determining a verification rule set corresponding to a multimodal knowledge base includes the following steps: determining the entity objects contained in the knowledge information in the multimodal knowledge base, and determining the association relationships between the entity objects; determining the ontology constraints corresponding to the entity objects and / or association relationships by analyzing the knowledge information, wherein the ontology constraints are used to characterize the reasonable constraints that the entity objects and / or association relationships should comply with; based on the ontology constraints, constructing a verification rule set corresponding to the multimodal knowledge base, wherein the verification rule set specifies the constraints that should be complied with when different association relationships exist between different entity objects.
[0057] Specifically, it integrates knowledge data from different sources, including but not limited to the experiential knowledge of domain experts, publicly available academic research results, industry reports, and authoritative databases; integrates data from multiple modalities, including text, images, audio, and video, and meticulously classifies and organizes the constructed knowledge to form a clear knowledge structure, ultimately building a multi-domain multimodal knowledge base; based on a large model, it deeply understands the content of the knowledge base and extracts key information, identifies entities, concepts, events in the text, and the semantic relationships between them, extracts implicit entity relationships (association relationships) in the knowledge base, identifies ontological constraints in the knowledge base, that is, the rules and restrictions followed by entities and their associations, and generates logical chains. Based on the mined entity relationships and ontological constraints, it automatically generates a set of validation rules to provide clear guidance on data validation rules.
[0058] Entity relationships refer to the associations between different entities. For example, in a medical knowledge base, the relationship between "disease," "symptoms," and "treatment methods"; in a geographic knowledge base, the relationship between "city," "country," and "geographic location." Ontological constraints are rules that restrict entities and their relationships, specifying the legal ways in which entities can be associated. For example, in the news domain, we can analyze the relationships between entities such as people, time, place, and the course of events in a news event, as well as the appropriate constraints on these entities in news reports.
[0059] This in-depth analysis enables a better understanding of the knowledge base, providing an accurate foundation for rule generation. Interpretability is a key characteristic of rule generation, meaning that the generated rules clearly explain their basis and judgment logic. For example, the rule "If the geographical location of an image is marked as a desert area, then the image should not contain a large amount of tropical rainforest vegetation" is based on the knowledge base's knowledge of the geographical characteristics of deserts and tropical rainforests, making it easy to understand and execute.
[0060] After determining the verification rule set corresponding to the multimodal knowledge base, the first verification process (hard verification) can be performed on the cross-modal data based on the verification rule set to obtain a first verification result; in addition, the audit model can be used to perform a second verification process (soft verification) on the cross-modal data to obtain a second verification result; based on the first verification result and the second verification result, the abnormality report corresponding to the cross-modal data can be determined. In the embodiment of the present application, an abnormality label classification system can be constructed to verify the logical rationality of the cross-modal data content through multi-level testing, quantitatively evaluate the abnormalities and sort them, and display the abnormal results. The details are as follows.
[0061] In some embodiments of the present application, the first verification result is used to indicate abnormal data that does not meet the constraints corresponding to the verification rule set, and the second verification result is used to indicate abnormal data with logical contradictions detected by the audit model; based on the first verification result and the second verification result, determining the abnormal report corresponding to the cross-modal data includes the following steps: integrating the abnormal data corresponding to the first verification result and the second verification result, wherein the abnormal data includes at least one of the following: data that is inconsistent with facts or common sense, data with inconsistent descriptions of different modal data, data that does not match related data, and data that lacks key information; determining the abnormality type and abnormality cause corresponding to the abnormal data, wherein the abnormality type includes at least one of the following: semantic abnormality, structural abnormality, attribute abnormality, modal feature conflict abnormality, context abnormality, knowledge base conflict abnormality, and the abnormality cause includes at least one of the following: collection error, input error, labeling error, and reasoning defect; using a Bayesian network model, combined with the abnormality type and abnormality cause, the abnormal data is analyzed to obtain the confidence corresponding to the abnormal data, and based on the confidence, the degree of abnormality corresponding to the abnormal data is determined; generating an abnormality report based on the abnormal data, and the abnormality type, abnormality cause and abnormality degree corresponding to the abnormal data.
[0062] Specifically, an abnormal label classification system is constructed, covering the degree of abnormality, abnormality type and abnormality cause, etc. According to the verification rule set and the knowledge base content, the input cross-modal data is checked for hard content logical constraints to determine whether there are logical contradictions, omissions or abnormalities between the data, and logically intercept cross-modal data that violates common sense, and quickly locate the specific location of the abnormality and the data items involved; and, the fine-tuned multimodal large model is used as an inference engine to perform soft complex reasoning and analysis detection on cross-modal data, considering the contextual information, semantic relationships and potential logical connections of the data, identifying hidden errors and logical contradictions, and detecting unreasonable content. Based on the hard and soft verification results (i.e. the first verification result and the second verification result), the abnormality confidence is calculated through the Bayesian network for hierarchical ranking, triggering the corresponding processing mechanism, displaying the abnormality ranking, degree, type, cause and details, and forming a visual abnormality report.
[0063] For example, the labels of abnormality levels in the embodiments of this application include but are not limited to:
[0064] 1) Hard contradictions refer to situations where data clearly conflict with facts, common sense, or established rules. There is a direct and definite logical contradiction between two or more modal data, and it is absolutely impossible for them to be true at the same time.
[0065] 2) Probability contradiction means that the data is likely to be inconsistent with other relevant data or expected situations and needs further verification;
[0066] 3) Potential risks, that is, there may be potential problems or uncertainties in the data, which require attention and in-depth analysis, such as some data missing key information but not affecting the overall logic;
[0067] In the embodiment of this application, the abnormal type labels include but are not limited to:
[0068] 1) Semantic exceptions
[0069] Content inconsistency: semantic mismatch between text and images, videos, and audio. For example, a press release mentions “protests” but the accompanying image shows “celebrations.”
[0070] Semantic contradiction: direct conflict at the semantic level between texts, images, etc. For example: a press release mentions "typhoon landfall", but the time and weather records show "sunny".
[0071] Named entity errors: errors in the labeling of entities such as names of people, places, and time. For example, the geographical label is "Beijing", but the place mentioned in the text is "Shanghai".
[0072] 2) Structural exceptions
[0073] Format mismatch: There is a conflict between the formats of different modal data. For example, an audio file contains multiple languages but is not labeled.
[0074] Inconsistent encoding: Cross-modal data encoding formats conflict, for example: text encoding is UTF-8, but image encoding is ISO-8859-1.
[0075] Resolution / size conflict: There are differences in quality or specifications between cross-modal data. For example, the video file resolution is 1080P, but the accompanying subtitle file is 480P.
[0076] 3) Time and space anomalies
[0077] Time conflict: There are logical inconsistencies between modal data timestamps. For example, a press release was published in 2024, but the image shows 2023.
[0078] Geographic conflict: The geographic location information does not match the modal data content. For example, the image label is "Chengdu", but the image content is "Shanghai".
[0079] Spatial conflict: Different scenes appear in the video or audio without indicating the transition or switch. For example, the first half of the news video is about "Beijing", and the second half shows "Shanghai" scenes without explanation.
[0080] 4) Modal feature conflict
[0081] Inconsistent gender characteristics: The characteristics of the person in the image or video do not match the labeled information, for example, the person is labeled as male, but the image shows a female.
[0082] Inconsistent age features: The labeled age conflicts with the actual modality features, such as the subject being labeled as a child but the image shows an adult.
[0083] Brand / Logo Conflict: The brand or logo in the image conflicts with the text description, for example, it is labeled as "Brand A", but the image contains the logo of "Brand B".
[0084] 5) Contextual anomaly
[0085] Logical consistency conflict: The data content is logically inconsistent, such as the text mentions "increase" but the chart shows "decrease".
[0086] Context mismatch: There is a mismatch in the context between modal data, such as a press release describing a “disaster scene” but an image showing a “happy scene.”
[0087] Lack of coherence: Discontinuity in time or space across modal data, such as a video where the first segment describes an earthquake and the second segment describes a hurricane without annotating the change.
[0088] 6) Knowledge base conflict
[0089] Common sense errors: Cross-modal data does not conform to common sense or knowledge base rules, such as being labeled "the sun rises in the east and sets in the west", but the sun is in the north in the image.
[0090] Professionalism conflict: Cross-modal data conflicts with the expertise of the domain knowledge base, such as a medical image showing a “fracture” but the text describing it as “soft tissue injury”.
[0091] Conflict of historical facts: The modal data does not match the historical records, such as describing the “Olympic Games held in 1999”, but it should actually be “2000”.
[0092] The labels of abnormal reasons in the embodiment of this application include but are not limited to:
[0093] 1) Collection error: Due to the limitations of collection equipment or means, the data content does not match the label.
[0094] 2) Input error: manual input error.
[0095] 3) Labeling errors: errors that occur during data labeling or processing, such as labeling "Beijing" but the content is "Shanghai".
[0096] 4) Reasoning defects: defects or biases in AI model reasoning, such as the model reasoning "apple" but the actual result is "orange".
[0097] Taking medical imaging data as an example, in terms of severity, the complete inconsistency between the image and the diagnosis result can be classified as a hard contradiction; in terms of type, if the patient's gender marked on the image does not match the actual gender, it is considered gender characteristic inconsistency; in terms of cause, it is an input error.
[0098] Based on the logic rules of the validation rule set and the knowledge base content, the input cross-modal data is checked for hard content logical constraints to determine whether there are logical contradictions, omissions, or anomalies between the data. Cross-modal data that violates common sense is logically intercepted. For example, logical interception is performed on image and text combinations that violate medical common sense. For example, when a male patient is associated with uterine images, this type of data that clearly violates common sense is intercepted in real time:
[0099] For example, in the PACS system of a tertiary hospital, the electronic medical record of patient A shows:
[0100] Gender: Male;
[0101] Age: 45 years old
[0102] Chief complaint: lower abdominal pain;
[0103] Examination items: pelvic CT scan;
[0104] The cross-modal data collected synchronously by the system include:
[0105] Text data: the diagnosis description "prostatic hyperplasia to be investigated" in the electronic medical record;
[0106] Image data: pelvic CT DICOM file (including images and metadata);
[0107] Feature extraction: The image analysis unit detects the uterine outline using ResNet-50 (98% confidence), and the patient's gender label is "male" based on DICOM metadata analysis.
[0108] Hard test results:
[0109] Abnormality type: anatomical contradiction (hard contradiction);
[0110] Handling method: Mark and intercept.
[0111] Fine-tune the multimodal large model as an inference engine to perform soft complex reasoning and analysis on cross-modal data to detect unreasonable content.
[0112] Based on the results of hard and soft verification, the anomaly confidence is calculated through the Bayesian network for hierarchical sorting, triggering the corresponding processing mechanism, displaying the anomaly sorting, degree, type, cause and details, and forming a visual report. First, the prior probability of each type of anomaly is determined based on historical data and domain knowledge. For the input data, the conditional probability is calculated by analyzing the degree of match with the relevant information in the knowledge base. Then, the posterior probability, that is, the anomaly confidence, is calculated using the Bayesian formula. For example, for the deviation of a certain feature in the data from the standard value in the knowledge base, the conditional probability is determined based on the size of the deviation and the probability of anomalies caused by similar deviations in history, and then the anomaly confidence is obtained. Based on hard and soft verification, the anomaly confidence is evaluated through the Bayesian network to classify and sort the data. The relationship between confidence and anomaly degree is as follows: level one (hard contradiction, confidence ≥ 95%), level two (probability contradiction, 80% ≤ confidence < 95%), level three (potential risk, 60% ≤ confidence < 80%);
[0113] Furthermore, after obtaining the abnormal report, the abnormal data can be automatically analyzed based on the large model to generate a correction strategy, as follows.
[0114] In some embodiments of the present application, the method also includes the following steps: determining historical abnormal data in historical abnormal data records whose similarity with abnormal data in abnormal reports exceeds a preset similarity threshold as target historical abnormal data; obtaining historical correction strategies corresponding to the target historical abnormal data; using an audit large model, combined with a multimodal knowledge base, to analyze the abnormal reports and historical correction strategies to obtain the target correction strategies corresponding to the abnormal data.
[0115] Specifically, a grading system of correction suggestion labels can be constructed based on the type and severity of data errors. Leveraging a large model, detailed correction rationale can be generated for each abnormal data point. Furthermore, by combining historical correction records, data context, and a domain template library, actionable correction suggestions can be generated, providing multiple possible correction options, a confidence level for each option, and a rationale for the correction.
[0116] For example, based on the type and severity of the data error, a graded suggested label can be output:
[0117] Level 1 recommendation: Direct correction (e.g., coordinate correction: "Hefei → Chengdu", confidence level 92%);
[0118] Level 2 recommendation: additional information (e.g., “a chromosome report from the patient is required to verify gender”);
[0119] Level 3 recommendation: manual review (e.g., “features in the image are fuzzy and require manual confirmation”);
[0120] At the same time, it is possible to combine historical correction records, contextual information of the data and the domain template library to generate actionable correction suggestions, provide multiple possible correction options, give the confidence level of each option, and provide a basis for correction. The domain template library is a set of structured correction rules predefined for specific industries (such as medical and geography). By constraining the practicality and compliance of AI suggestions, it solves the problem of over-generalization of large model outputs. Its core includes a three-layer structure of standardized parameters, operating instructions and verification basis: Taking medical imaging as an example, for the problem of "abnormal window width", the template library not only limits the adjustment range of CT values (such as window width 300-2000HU), but also provides DICOM tool operation instructions and NCCN guideline references; in geographic data scenarios, coordinate system deviation correction must be specified in accordance with the national standard GB / T20257. The conversion model (such as the seven-parameter method) must meet the accuracy requirements (RMSE ≤ 0.5 pixels). The template library integrates AI generation capabilities with industry standards through dynamic, linked knowledge base updates and automated verification mechanisms. This transforms correction suggestions from vague prompts ("adjust parameters") into executable instructions ("set to 1600 HU, in accordance with liver examination standards"), significantly improving the practical value and effectiveness of the cross-modal audit system. An example of correction options: Based on historical correction records and context, it provides multiple options for geo-annotation errors: "Recommend changing to Chengdu (92% confidence) or Chongqing (85% confidence)" and provides the rationale for the correction (e.g., "The image contains architectural features of western Sichuan").
[0121] After the exception report and target correction strategy are generated, they can be sent to the target object's terminal device for manual preliminary review. If the manual preliminary review opinion is inconsistent with the automatically generated content, a double-blind expert review can be further carried out, as follows.
[0122] In some embodiments of the present application, the method also includes the following steps: sending the exception report and the target correction strategy corresponding to the exception data indicated in the exception report to the front-end interactive interface of the terminal device of the target object for display; in response to a first instruction triggered in the front-end interactive interface, correcting the exception data according to the target correction strategy, wherein the first instruction indicates that the target object recognizes the exception data and the target correction strategy indicated in the exception report; in response to a second instruction triggered in the front-end interactive interface, sending the exception report and the target correction strategy to the front-end interactive interfaces of the terminal devices of multiple other target objects for display, wherein the second instruction indicates that the target object does not recognize the exception data and the target correction strategy indicated in the exception report; obtaining the audit results returned by the terminal devices of multiple other target objects, and determining whether the exception data and the target correction strategy indicated in the exception report are correct based on the audit results.
[0123] In some embodiments of the present application, the method also includes: when the abnormal data indicated by the abnormal report is corrected, obtaining log data of the entire life cycle of analysis and processing of the abnormal data, wherein the log data includes at least one of the following: the original content of the abnormal data, the target correction strategy given by the model, the operating instructions and opinions of the target object, the audit results of other target objects, and the final correction results; and updating the log data as new knowledge information to the multimodal knowledge base.
[0124] Specifically, for data identified as abnormal, manual reviewers are automatically notified to conduct individual checks and corrections. Reviewers, based on the system's graded correction recommendations and their own industry knowledge and experience, conduct individual checks and corrections on the data transferred to the manual review queue, confirming, supplementing, and modifying the system's recommendations. When a conflict arises between the model's opinions and those of the manual reviewers, a double-blind second review mechanism is initiated, with three experts randomly assigned to conduct back-to-back verifications to reach a final conclusion. The entire lifecycle of data modifications is recorded, and detailed log data is generated for each correction action. Log data can be organized by chronological order, data type, or other classification methods to facilitate user query and retrieval, and facilitate the tracing and review of data change processes.
[0125] For example, the types of structures that may be subject to preliminary review may include:
[0126] 1) Confirmation: If the recommendation is found to be reasonable and accurate, the reviewer or system will act on it;
[0127] 2) Supplementation: If any information is deemed incomplete, the reviewer will supplement or adjust it based on the actual situation. For example, during a news data audit, the system may suggest that the location name in a news article be modified. However, after further investigation and verification, the reviewer may find that not only the location name needs to be revised, but also the relevant background information of the event needs to be supplemented. In this case, the reviewer will correct the location name and complete the background information of the event.
[0128] 3) Modification: If the AI system determines that an image shows a certain disease, the reviewer will make adjustments based on the actual situation. In medical imaging diagnosis, if the AI system believes the image shows a certain disease, but the manual reviewer, based on their clinical experience, determines a different disease, the reviewer will modify the disease information. If this type of verification is deemed to conflict with the model's initial review, it will automatically enter a double-blind second review.
[0129] When the model's opinions conflict with those of the manual reviewers, a double-blind second review mechanism can be activated, randomly assigning three experts to conduct a back-to-back verification to reach a final conclusion. Once a conflict of opinion is detected, the conflict handling unit immediately initiates the double-blind review mechanism. When the system's AI recommendations conflict with the manual initial review, three anonymous experts in the relevant field are randomly assigned to conduct a back-to-back verification of the conflicting anomalous data. Experts only have access to necessary information related to the review and are shielded from other sensitive information. Without knowledge of the model's recommendations or other expert opinions, the experts independently analyze and judge the data, thereby maximizing the fairness and objectivity of the review and preventing interference from other factors.
[0130] If the three experts disagree, a majority vote will be used to determine the final conclusion. If the consensus is 2 to 1, the opinion with the most support will be the final conclusion. If the consensus is 1 to 1, the experts' arguments will be further analyzed and weighted based on their professional background and experience. A weighted score will be calculated, and the opinion with the highest score will be the final conclusion.
[0131] Furthermore, the entire lifecycle of data modifications can be recorded, generating a detailed log for each correction action, facilitating the tracing and review of data change processes. This recorded information includes the original content of the abnormal data, the system's recommendations, the manual reviewer's operational process and opinions, the opinions of the double-blind review experts (if any), the final correction results, and the person who performed the correction. When data audits or investigations are necessary, relevant personnel can query the correction log to understand the data's source, changes, and the basis for processing. For example, blockchain technology can be used to maintain tamper-proof records based on rules, meeting compliance audit requirements in the medical and news sectors.
[0132] In an embodiment of the present application, the knowledge base can also be manually updated and automatically updated incrementally based on the newly added knowledge and abnormal modification results, and the knowledge base, rule set, abnormal threshold and model can be dynamically optimized. Specifically, for newly added or changed businesses, knowledge, or when errors or incomplete content are found in the knowledge base, the knowledge base can be updated directly to ensure the timeliness and accuracy of the knowledge base content. For example, in the medical field, when a new type of disease is discovered or a treatment method is updated, manual maintenance personnel can accurately enter the relevant information into the knowledge base, including detailed content such as the symptoms of the disease, diagnostic criteria, and treatment plans, to ensure the timeliness and accuracy of the knowledge base.
[0133] After manual review confirms cross-modal data anomalies and completes corrections, the corrected data is automatically fed back into the knowledge base. Based on the final data logic corrections, the knowledge base content is dynamically expanded, allowing it to promptly reflect the latest data and cover more business scenarios and knowledge areas, forming a closed loop of "audit-feedback-evolution."
[0134] At the same time, the inference rule set is regularly evaluated and analyzed to determine the accuracy of the rules in identifying anomalies in the data. Based on the results of the rule evaluation and analysis, the rules are adjusted and improved accordingly. Based on new knowledge base content, the rule set is automatically optimized based on the large model. By analyzing the entity relationships and ontological constraints in the new data, the inference rules are adjusted and improved, improving the verification accuracy and adaptability of the rule set. False positives and false negatives can also be analyzed, and the model verification results and final data correction results are regularly statistically analyzed. Based on the false positives and false negatives, the rule thresholds and large model parameters are dynamically adjusted. After adjusting the thresholds, re-evaluation is performed until performance is optimized.
[0135] Through the online learning mechanism, the large model can continuously learn from new data and improve its reasoning capabilities. The model's continuous learning frequency is set to once a week. At a fixed time each week, the system automatically screens new data that meets the requirements and inputs it into the large model for learning. During the learning process, the model parameters are adjusted based on the characteristics of the new data and the feedback from the model. After learning is completed, the updated model is tested and evaluated to ensure its performance has been improved. If the model performance degrades, it rolls back to the previous model state, reanalyzes the new data, and adjusts the learning parameters. Through the online learning mechanism, the large model can acquire knowledge from new data in real time and continuously update and improve its model parameters. Through continuous learning, the reasoning and generalization capabilities of the large model are improved, enabling it to better adapt to the ever-changing data characteristics and business needs.
[0136] The embodiments of this application significantly enhance the intelligence and accuracy of cross-modal data content logic auditing. By combining large-scale model technology with a knowledge base, rule generation becomes more scientific. Multi-level validation accurately locates anomalies, reducing misjudgments. Large-scale model-generated correction suggestions provide clear action directions, reducing the difficulty of manual decision-making. Manual and double-blind review ensures reliable results. Dynamic updates and optimization ensure the system keeps pace with data changes, while online learning enhances model reasoning capabilities.
[0137] This application achieves comprehensive and accurate verification of cross-modal data by combining a multimodal knowledge base and an auditing large model. On the one hand, the professional knowledge and specifications of the multimodal knowledge base are used to ensure that the data conforms to the standards and logic within the field. On the other hand, with the help of the semantic reasoning ability of the large model, logical contradictions between data are effectively detected, thereby improving data quality. In addition, by analyzing historical abnormal data and correction strategies, the audit model and knowledge base are continuously optimized to form a closed-loop feedback mechanism, which significantly improves the efficiency of identifying and correcting abnormal data, reduces the cost of manual review, and enhances the intelligence and adaptability of the system. The application of this method can not only ensure the accuracy and consistency of the data, but also promote the reliability of data-driven decision-making. It has important practical value for scenarios such as big data analysis and artificial intelligence system training.
[0138] According to an embodiment of the present application, an embodiment of a data auditing device is also provided. Figure 4 This is a structural diagram of a data audit device provided according to an embodiment of the present application. Figure 4 As shown, the device includes:
[0139] A data acquisition module 40 is configured to acquire cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities;
[0140] A first verification module 42 is configured to determine a verification rule set corresponding to the multimodal knowledge base and, based on the verification rule set, perform a first verification process on the cross-modal data to obtain a first verification result. The multimodal knowledge base contains at least professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set indicates constraints that the cross-modal data should comply with, and the first verification process verifies whether the cross-modal data contains any anomalies that do not comply with the rules in the verification rule set.
[0141] A second verification module 44 is configured to perform a second verification process on the cross-modal data using the audit macro model to obtain a second verification result. The second verification process is configured to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning capability of the macro model.
[0142] The anomaly analysis module 46 is configured to determine an anomaly report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the anomaly report is at least used to characterize the abnormal data in the cross-modal data.
[0143] Optionally, after obtaining the cross-modal data to be audited, the data auditing device is further used to: use a feature separation network to separate the features of different modalities in the cross-modal data to obtain modal features, wherein the modal features include at least one of the following: texture features, color features, semantic features, grammatical features, spectral features, and phoneme features; perform structured extraction on the modal features to obtain key metadata information corresponding to the modal features, wherein the key metadata information is used to characterize the semantic information and format attribute information of the data content; construct a multimodal data cube corresponding to the cross-modal data by normalizing the key metadata information, wherein the normalization process is used to unify the data format and dimension of the key metadata information.
[0144] Optionally, determining the verification rule set corresponding to the multimodal knowledge base includes: determining the entity objects contained in the knowledge information in the multimodal knowledge base, and determining the association relationships between the entity objects; determining the ontology constraints corresponding to the entity objects and / or association relationships by analyzing the knowledge information, wherein the ontology constraints are used to characterize the reasonable constraints that the entity objects and / or association relationships should comply with; based on the ontology constraints, constructing the verification rule set corresponding to the multimodal knowledge base, wherein the verification rule set specifies the constraints that should be met when different association relationships exist between different entity objects.
[0145] Optionally, the first verification result is used to indicate abnormal data that does not meet the constraints corresponding to the verification rule set, and the second verification result is used to indicate abnormal data with logical contradictions detected by the audit model; based on the first verification result and the second verification result, determining the abnormal report corresponding to the cross-modal data includes: integrating the abnormal data corresponding to the first verification result and the second verification result, wherein the abnormal data includes at least one of the following: data inconsistent with facts or common sense, data with inconsistent descriptions of different modal data, data that does not match related data, and data missing key information; determining the abnormality type and abnormality cause corresponding to the abnormal data, wherein the abnormality type includes at least one of the following: semantic abnormality, structural abnormality, attribute abnormality, modal feature conflict abnormality, contextual abnormality, knowledge base conflict abnormality, and the abnormality cause includes at least one of the following: collection error, entry error, labeling error, reasoning defect; using a Bayesian network model, combined with the abnormality type and abnormality cause, to analyze the abnormal data, obtain the confidence corresponding to the abnormal data, and determine the abnormality degree corresponding to the abnormal data based on the confidence; generating an abnormality report based on the abnormal data, and the abnormality type, abnormality cause and abnormality degree corresponding to the abnormal data.
[0146] Optionally, the data auditing device is also used to: determine historical abnormal data in the historical abnormal data record whose similarity with the abnormal data in the abnormal report exceeds a preset similarity threshold as target historical abnormal data; obtain the historical correction strategy corresponding to the target historical abnormal data; use the audit big model, combined with the multimodal knowledge base, to analyze the abnormal report and the historical correction strategy to obtain the target correction strategy corresponding to the abnormal data.
[0147] Optionally, the data auditing device is also used to: send the exception report and the target correction strategy corresponding to the abnormal data indicated in the exception report to the front-end interactive interface of the terminal device of the target object for display; in response to a first instruction triggered in the front-end interactive interface, correct the abnormal data according to the target correction strategy, wherein the first instruction indicates that the target object agrees with the abnormal data and the target correction strategy indicated in the exception report; in response to a second instruction triggered in the front-end interactive interface, send the exception report and the target correction strategy to the front-end interactive interfaces of the terminal devices of multiple other target objects for display, wherein the second instruction indicates that the target object does not agree with the abnormal data and the target correction strategy indicated in the exception report; obtain the audit results returned by the terminal devices of multiple other target objects, and determine whether the abnormal data and the target correction strategy indicated in the exception report are correct based on the audit results.
[0148] Optionally, the data auditing device is also used to: when the abnormal data indicated by the abnormal report is corrected, obtain log data of the entire life cycle of analysis and processing of the abnormal data, wherein the log data includes at least one of the following: the original content of the abnormal data, the target correction strategy given by the model, the operating instructions and opinions of the target object, the audit results of other target objects, and the final correction results; and update the log data as new knowledge information to the multimodal knowledge base.
[0149] It should be noted that the various modules in the above-mentioned data auditing device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0150] It should be noted that the data auditing device provided in this embodiment can be used to perform Figure 2 The data auditing method shown, therefore, the relevant explanations and descriptions of the above data auditing method are also applicable to the embodiments of this application and will not be repeated here.
[0151] An embodiment of the present application also provides a non-volatile storage medium, the non-volatile storage medium including a stored computer program, wherein a device where the non-volatile storage medium is located executes the following data auditing method by running the computer program: obtaining cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; determining a verification rule set corresponding to a multimodal knowledge base, and performing a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base at least includes professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set is used to indicate the constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data has anomalies that do not comply with the rules in the verification rule set; using an audit large model, performing a second verification process on the cross-modal data to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning ability of the large model; and determining an exception report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the exception report is used at least to characterize abnormal data in the cross-modal data.
[0152] An embodiment of the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the data audit method described in each embodiment of the present application: obtaining cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; determining a verification rule set corresponding to a multimodal knowledge base, and performing a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base at least includes professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set is used to indicate the constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data contains anomalies that do not comply with the rules in the verification rule set; using an audit large model, performing a second verification process on the cross-modal data to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning capability of the large model; and determining an exception report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the exception report is used at least to characterize abnormal data in the cross-modal data.
[0153] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0154] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0156] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0157] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0159] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data audit method, characterized in that: include: Acquiring cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; Determining a verification rule set corresponding to a multimodal knowledge base, and performing a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base contains at least professional knowledge and specifications in a field corresponding to the cross-modal data to be audited, the verification rule set is used to indicate constraints that the cross-modal data should comply with, and the first verification process is used to verify whether the cross-modal data has anomalies that do not comply with rules in the verification rule set; Using the auditing macro model, performing a second verification process on the cross-modal data to obtain a second verification result, wherein the second verification process is used to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning ability of the macro model; An exception report corresponding to the cross-modal data is determined based on the first verification result and the second verification result, wherein the exception report is at least used to characterize abnormal data in the cross-modal data.
2. The data auditing method according to claim 1, characterized in that: After obtaining the cross-modal data to be audited, the method further includes: Using a feature separation network, separating features of different modalities in the cross-modal data to obtain modal features, wherein the modal features include at least one of the following: texture features, color features, semantic features, grammatical features, spectral features, and phoneme features; Performing structured extraction on the modal features to obtain key metadata information corresponding to the modal features, wherein the key metadata information is used to characterize semantic information and format attribute information of data content; By normalizing the key metadata information, a multimodal data cube corresponding to the cross-modal data is constructed, wherein the normalization process is used to unify the data format and dimension of the key metadata information.
3. The data audit method according to claim 1, characterized in that: The verification rule set corresponding to the multimodal knowledge base includes: Determining entity objects included in the knowledge information in the multimodal knowledge base, and determining association relationships between the entity objects; Determining ontology constraints corresponding to the entity objects and / or the association relationships by analyzing the knowledge information, wherein the ontology constraints are used to represent reasonable constraints that the entity objects and / or the association relationships should comply with; Based on the ontology constraints, the verification rule set corresponding to the multimodal knowledge base is constructed, wherein the verification rule set specifies the constraint conditions that should be met when different association relationships exist between different entity objects.
4. The data auditing method according to claim 1, characterized in that: The first verification result is used to indicate abnormal data that does not meet the constraint conditions corresponding to the verification rule set, and the second verification result is used to indicate abnormal data with logical contradictions detected by the audit model; Determining, based on the first verification result and the second verification result, an abnormality report corresponding to the cross-modal data includes: Integrating the abnormal data corresponding to the first verification result and the second verification result, wherein the abnormal data includes at least one of the following: data inconsistent with facts or common sense, data with inconsistent descriptions of different modal data, data that does not match related data, and data missing key information; Determine the abnormality type and abnormality cause corresponding to the abnormal data, wherein the abnormality type includes at least one of the following: semantic abnormality, structural abnormality, attribute abnormality, modal feature conflict abnormality, context abnormality, and knowledge base conflict abnormality; the abnormality cause includes at least one of the following: acquisition error, input error, annotation error, and reasoning defect; Using a Bayesian network model, combined with the abnormality type and the abnormality cause, the abnormal data is analyzed to obtain a confidence level corresponding to the abnormal data, and based on the confidence level, the abnormality level corresponding to the abnormal data is determined; The exception report is generated based on the exception data, and the exception type, the exception cause, and the exception degree corresponding to the exception data.
5. The data audit method according to claim 4, characterized in that: The method further comprises: Determine historical abnormal data in the historical abnormal data record whose similarity with the abnormal data in the abnormal report exceeds a preset similarity threshold as target historical abnormal data; Obtaining a historical correction strategy corresponding to the target historical abnormal data; The audit big model is used in combination with the multimodal knowledge base to analyze the abnormal report and the historical correction strategy to obtain the target correction strategy corresponding to the abnormal data.
6. The data audit method according to claim 5, characterized in that: The method further comprises: Sending the exception report and the target correction strategy corresponding to the abnormal data indicated in the exception report to the front-end interactive interface of the terminal device of the target object for display; In response to a first instruction triggered in the front-end interactive interface, the abnormal data is corrected according to the target correction strategy, wherein the first instruction indicates that the target object agrees with the abnormal data indicated in the abnormal report and the target correction strategy; In response to a second instruction triggered in the front-end interactive interface, the abnormality report and the target correction strategy are respectively sent to the front-end interactive interfaces of terminal devices of multiple other target objects for display, wherein the second instruction indicates that the target object does not agree with the abnormal data and the target correction strategy indicated in the abnormality report; The audit results returned by the terminal devices of the plurality of other target objects are obtained, and based on the audit results, it is determined whether the abnormal data indicated in the abnormal report and the target correction strategy are correct.
7. The data audit method according to claim 6, characterized in that: The method further comprises: When the abnormal data indicated in the abnormal report has been corrected, obtaining log data of the entire life cycle of analyzing and processing the abnormal data, wherein the log data includes at least one of the following: the original content of the abnormal data, the target correction strategy given by the model, the operating instructions and opinions of the target object, the review results of other target objects, and the final correction result; The log data is used as new knowledge information to update the multimodal knowledge base.
8. A data auditing device, characterized in that: include: A data acquisition module, configured to acquire cross-modal data to be audited, wherein the cross-modal data includes data of at least two different modalities; a first verification module, configured to determine a verification rule set corresponding to a multimodal knowledge base, and perform a first verification process on the cross-modal data based on the verification rule set to obtain a first verification result, wherein the multimodal knowledge base contains at least professional knowledge and specifications in the field corresponding to the cross-modal data to be audited, the verification rule set indicates constraints that the cross-modal data should comply with, and the first verification process is configured to verify whether the cross-modal data contains any anomalies that do not comply with rules in the verification rule set; a second verification module, configured to perform a second verification process on the cross-modal data using the audit macro model to obtain a second verification result, wherein the second verification process is configured to detect whether there are logical contradictions in the cross-modal data based on the semantic reasoning capability of the macro model; An anomaly analysis module is configured to determine an anomaly report corresponding to the cross-modal data based on the first verification result and the second verification result, wherein the anomaly report is at least used to characterize abnormal data in the cross-modal data.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the data audit method according to any one of claims 1 to 7 is executed when the program is run.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data audit method according to any one of claims 1 to 7 by running the computer program.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data audit method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-mode network content security intelligent auditing system and method thereof
CN118312922A
Bidirectional reasoning evaluation method for fault risk of civil aircraft
CN118485152A
Intelligent voucher auditing method and system based on large language model
CN119027251A
Cited By
Power data dynamic verification method based on large model
CN121328527A
A large model-based power data dynamic verification method
CN121328527B