Data processing method and device based on artificial intelligence, computer equipment and medium
By utilizing multi-dimensional detection—including anomaly detection, data review, and correction agents—in intelligent underwriting and claims review within the financial insurance sector, the high misjudgment rate of existing models in complex scenarios has been resolved. This has enabled efficient and accurate correction of prediction results, thereby improving the automation and accuracy of business decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-08
AI Technical Summary
Existing intelligent underwriting and claims review models in the financial insurance field have a high misjudgment rate in complex business scenarios and lack real-time misjudgment identification and self-correction capabilities, resulting in additional compensation costs for insurance companies or damage to customer experience.
An AI-based data processing approach is employed, combining an anomaly detection agent, a data review agent, and a correction agent to perform multi-dimensional detection and automatically generate corrected target prediction results. This includes anomaly detection, knowledge retrieval, semantic comparison, and causal consistency testing, ensuring the accuracy and efficiency of the results.
It improves the efficiency of identifying and processing misjudgments in business model predictions, ensures the accuracy of generated correction results, reduces reliance on manual review, and enhances the precision and automation of financial and insurance business decisions.
Smart Images

Figure CN121998773A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to data processing methods, devices, computer equipment and storage media based on artificial intelligence. Background Technology
[0002] In intelligent underwriting and claims review scenarios within the financial insurance sector, while existing business models can generate predictive results based on input data, their accuracy verification and correction mechanisms still have significant shortcomings. Specifically, current technologies typically rely solely on manual review of model outputs on a case-by-case basis. This is not only inefficient and susceptible to subjective experience, but also fails to guarantee accuracy, hindering large-scale application. Furthermore, due to the lack of dynamic identification and self-correction capabilities for prediction errors, the model exhibits a high misjudgment rate in complex business scenarios (such as underwriting of non-standard individuals or claims with high fraud risk), leading to additional claims costs for insurance companies or negative customer experience. For instance, in health insurance claims review, existing models may misclassify pre-existing conditions that should be rejected as reimbursable items due to insufficient correlation between the insured's historical medical records and the current claim application. If this oversight is not detected during manual review, it will directly result in financial losses for the insurance company.
[0003] Therefore, there is an urgent need to develop an intelligent auditing system with real-time misjudgment identification and self-correction capabilities to improve the accuracy and automation of financial and insurance business decisions. Summary of the Invention
[0004] The purpose of this application is to propose a data processing method, apparatus, computer equipment, and storage medium based on artificial intelligence, so as to solve the technical problems of low efficiency and inaccuracy in the existing method of relying on manual review of model outputs.
[0005] Firstly, an artificial intelligence-based data processing method is provided, including: The system acquires the business data to be processed and performs inference processing on the business data based on a preset business model to obtain the corresponding prediction results. Obtain the context information corresponding to the business data; Based on the context information, a preset anomaly detection agent is used to perform anomaly detection processing on the prediction result to obtain the corresponding anomaly detection result. If the anomaly detection result indicates the presence of an anomaly, the business data will be processed for reasonableness detection based on a preset data review intelligence agent to obtain the corresponding review result. Based on the review results, a preset correction agent is used to perform result correction processing on the business data to obtain the corresponding target prediction results. The target prediction results are verified based on a preset verification strategy. If the target prediction result passes verification, the target prediction result is then output.
[0006] Secondly, an artificial intelligence-based data processing device is provided, comprising: The inference module is used to acquire the business data to be processed and to perform inference processing on the business data based on a preset business model to obtain the corresponding prediction results. The acquisition module is used to acquire context information corresponding to the business data; The first detection module is used to perform anomaly detection processing on the prediction result based on the context information using a preset anomaly detection agent, and obtain the corresponding anomaly detection result. The second detection module is used to perform reasonableness detection processing on the business data based on a preset data review intelligent agent if the anomaly detection result is that there is an anomaly, and obtain the corresponding review result. The correction module is used to correct the business data based on the review results using a preset correction agent to obtain the corresponding target prediction results. The verification module is used to verify the target prediction results based on a preset verification strategy. The output module is used to output the target prediction result if the target prediction result passes the verification.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based data processing method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based data processing method.
[0009] In the aforementioned solution implemented by the artificial intelligence-based data processing method, apparatus, computer equipment, and storage medium, the following steps are taken: First, business data to be processed is acquired, and the business data is inferred based on a preset business model to obtain a corresponding prediction result. Then, context information corresponding to the business data is acquired. Based on the context information, a preset anomaly detection agent is used to perform anomaly detection processing on the prediction result to obtain a corresponding anomaly detection result. If the anomaly detection result indicates an anomaly, a preset data review agent is used to perform rationality detection processing on the business data to obtain a corresponding review result. Subsequently, based on the review result, a preset correction agent is used to perform result correction processing on the business data to obtain a corresponding target prediction result. The target prediction result is then verified based on a preset verification strategy. If the target prediction result passes verification, the target prediction result is output. Unlike existing methods that rely on manual review of model outputs, this application, after obtaining prediction results by reasoning through business data based on the business model, intelligently and accurately performs multi-dimensional detection on the prediction results using a combination of anomaly detection agents, data review agents, and correction agents to identify model misjudgments and automatically generate corrected target prediction results. This effectively improves the processing efficiency of identifying misjudgments in the prediction results of the business model and ensures the accuracy of the generated corrected results. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the artificial intelligence-based data processing method according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of the artificial intelligence-based data processing apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0019] It should be noted that the data processing method based on artificial intelligence provided in the embodiments of this application is generally executed by a server / terminal device, and correspondingly, the data processing device based on artificial intelligence is generally set in the server / terminal device.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based data processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based data processing method provided in this application can be applied to any scenario requiring result analysis, and thus can be applied to products in these scenarios, such as result analysis scenarios in the financial insurance field. The AI-based data processing method includes the following steps: Step S201: Obtain the business data to be processed, and perform inference processing on the business data based on the preset business model to obtain the corresponding prediction result.
[0022] In this embodiment, the data processing method based on artificial intelligence runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the business data to be processed via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultrawideband) connections, and other currently known or future-developed wireless connection methods. The execution result of this application is specifically a data processing system, which can be simply referred to as the system. The aforementioned business model can be a model applied to tasks such as customer service quality inspection, intelligent underwriting, and claims review in the financial insurance field, such as a trained claims decision model or underwriting classification model. The aforementioned business data is the business data to be verified, which may include customer characteristics (user age, medical history), claims materials, underwriting applications, etc. Furthermore, by loading a trained business model, such as a claims decision model in PyTorch or TensorFlow format, the business data is converted into the tensor format required by the model. The business model performs inference calculations on the modified business data, generating raw outputs (such as logits vectors) through forward propagation, and then applying the Softmax function to transform them into probability distributions. Finally, it generates prediction results containing two types of outputs: prediction labels (e.g., "underwritten / rejected," "full compensation / partial compensation / rejection") and confidence vectors (the model's predicted probability distribution for each category). =[ , ,..., ],For example =[0.3,0.6,0.1] indicates that the model considers the probability of "partial compensation" to be 60%.
[0023] Step S202: Obtain the context information corresponding to the business data.
[0024] In this embodiment, the aforementioned contextual information may include at least timestamps, customer entity characteristics, and business scenarios. A timestamp refers to the time when the prediction is recorded. Customer entity characteristics include user age, gender, region, and occupation (in an insurance scenario, "insured's age = 65 years old"). Business scenarios may include claims scenarios, underwriting scenarios, and so on.
[0025] Step S203: Based on the context information, use a preset anomaly detection agent to perform anomaly detection processing on the prediction result to obtain the corresponding anomaly detection result.
[0026] In this embodiment, the specific implementation process of using a preset anomaly detection agent to perform anomaly detection processing on the prediction result based on the context information to obtain the corresponding anomaly detection result will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0027] Step S204: If the anomaly detection result indicates the presence of an anomaly, then the business data is processed for rationality detection based on a preset data review agent to obtain the corresponding review result.
[0028] In this embodiment, the specific implementation process of the above-mentioned reasonableness detection processing of the business data based on the preset data review intelligent agent to obtain the corresponding review result will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0029] Step S205: Based on the review results, use a preset correction agent to perform result correction processing on the business data to obtain the corresponding target prediction results.
[0030] In this embodiment, the specific implementation process of using a preset correction agent to perform result correction processing on the business data based on the review results to obtain the corresponding target prediction results will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0031] Step S206: Verify the target prediction result based on a preset verification strategy.
[0032] In this embodiment, the specific implementation process of verifying the target prediction result based on the preset verification strategy will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0033] Step S207: If the target prediction result passes the verification, the target prediction result is output.
[0034] In this embodiment, the obtained target prediction results can be sent to relevant personnel via email, text message, or interface display, thereby completing the output processing of the target prediction results.
[0035] This application first acquires the business data to be processed, and performs reasoning processing on the business data based on a preset business model to obtain the corresponding prediction result; then, it acquires the context information corresponding to the business data; and based on the context information, it uses a preset anomaly detection agent to perform anomaly detection processing on the prediction result to obtain the corresponding anomaly detection result; if the anomaly detection result indicates the presence of anomalies, it performs reasonableness detection processing on the business data based on a preset data review agent to obtain the corresponding review result; then, based on the review result, it uses a preset correction agent to perform result correction processing on the business data to obtain the corresponding target prediction result; subsequently, it performs verification processing on the target prediction result based on a preset verification strategy; if the target prediction result passes verification, it outputs the target prediction result. Unlike existing methods that rely on manual review of model outputs, this application, after obtaining prediction results by reasoning through business data based on the business model, intelligently and accurately performs multi-dimensional detection on the prediction results using a combination of anomaly detection agents, data review agents, and correction agents to identify model misjudgments and automatically generate corrected target prediction results. This effectively improves the processing efficiency of identifying misjudgments in the prediction results of the business model and ensures the accuracy of the generated corrected results.
[0036] In some alternative implementations, step S203 includes the following steps: Based on the anomaly detection agent, the confidence distribution entropy corresponding to the prediction result is calculated.
[0037] In this embodiment, the aforementioned anomaly detection agent, also known as the anomaly detection agent, is responsible for detecting potential errors or predictions with abnormal confidence levels from the model output of the business model. It is the first "perception checkpoint" in the cognitive link. The anomaly detection agent uses multi-signal fusion detection to comprehensively consider model uncertainty and input feature distribution drift, achieving "high recall and low false positive" detection of misjudged samples.
[0038] The prediction results mentioned above include predicted labels and confidence vectors. This can be achieved using the algorithm logic of confidence distribution entropy detection: To calculate the confidence distribution entropy corresponding to the prediction result, where, Let i be the confidence vector, and i be the label category. If the confidence distribution entropy is higher than the dynamic threshold... This indicates that the business model is uncertain about the prediction result, and the prediction result can be marked as high uncertainty. Furthermore, the process of setting the aforementioned dynamic threshold may include: statistically analyzing the entropy distribution from historical data (such as all prediction samples from the past 30 days), and taking the 95th percentile as the threshold. Example: If 95% of the samples in the historical data have an entropy value ≤ 0.7, then... =0.7.
[0039] Calculate the confidence margin index corresponding to the prediction result.
[0040] In this embodiment, it can be achieved through an algorithm. To calculate the confidence margin index, where... This refers to the difference between the two categories with the highest confidence levels. If... This indicates that the business model is hesitant among multiple categories. The setting methods include: setting an experience threshold. =0.2, if This is marked as "low discrimination". Example: =0.3>0.2 → Pass, but it is close to the threshold and needs to be combined with other indicators.
[0041] Based on a preset anomaly detection model, the business data and the context information are processed to perform anomaly detection calculations to obtain the corresponding anomaly values.
[0042] In this embodiment, the anomaly detection model described above can be an isolated forest model trained using historical input feature data (such as customer feature vectors) to obtain an output anomaly score. The model for the anomaly detection function. Specifically, the training and generation process of the anomaly detection model is based on the core idea of "quickly isolating anomalies through random partitioning". The steps include: First, randomly extracting sub-samples from historical input feature data (such as customer feature vectors, business scenarios, material descriptions, etc.) to construct multiple isolated trees (iTrees). Each tree recursively and randomly selects a feature and a random split value for that feature, continuously bisecting the data space until all samples are isolated to leaf nodes or the preset tree height limit is reached. Since anomalies are usually isolated with fewer partitions (shallower path depths) due to their sparse distribution or extreme feature values, while normal points require more partitions (deeper paths), the average path depth of the sample in all trees is calculated and normalized to an anomaly score / anomaly value in the range of 0-1 (the closer the score is to 1, the higher the probability of an anomaly). Finally, the model reduces the influence of randomness through ensemble learning of multiple trees (such as taking the average score) and outputs the anomaly score for each sample. The scores of normal samples are usually concentrated in the lower range (such as <0.5), while the scores of anomalies are significantly higher (such as >0.7).
[0043] Furthermore, the trained anomaly detection model can be used to perform anomaly detection calculations on the input business data and context information to output the corresponding anomaly values.
[0044] Based on a preset anomaly score calculation strategy, the confidence distribution entropy, the confidence margin index, and the anomaly value are calculated and processed to obtain the corresponding anomaly score.
[0045] In this embodiment, the formula corresponding to the above-mentioned abnormal score calculation strategy includes: .in, For abnormal scores, The confidence distribution entropy, The first preset weight corresponding to the confidence distribution entropy, The confidence margin index. The second preset weight corresponding to the confidence margin index, This is an abnormal value. This is the third preset weight corresponding to the abnormal values. Furthermore, the values of the first, second, and third preset weights can be set according to actual business needs. For example, the first, second, and third preset weights can be set to 0.4 (model uncertainty), 0.3 (discrimination), and 0.3 (input anomaly), respectively. Alternatively, they can be adjusted according to business priorities.
[0046] Furthermore, the corresponding anomaly score can be obtained by substituting the aforementioned confidence distribution entropy, confidence margin index, and anomaly values into the formula corresponding to the anomaly score calculation strategy.
[0047] Determine whether the abnormal score is greater than a preset abnormal threshold.
[0048] In this embodiment, the optimal split point can be determined using the ROC curve as the corresponding outlier threshold (e.g., anomaly threshold). =0.7).
[0049] If yes, generate a first anomaly detection result corresponding to the prediction result; otherwise, generate a second anomaly detection result corresponding to the prediction result.
[0050] In this embodiment, if the detected anomaly score is greater than a preset anomaly threshold, a first anomaly detection result corresponding to the predicted result is generated, and the business data is marked as a "suspected misjudgment sample," and an anomaly sample list with explanatory labels is output. For example: {Case ID: 202501, Reason for anomaly: "Low confidence + Feature drift", Anomaly score: 0.87}. Conversely, if the detected anomaly score is not greater than the preset anomaly threshold, a second anomaly detection result corresponding to the predicted result is generated, and no anomaly is found.
[0051] This application calculates the confidence distribution entropy corresponding to the prediction result based on the anomaly detection agent; calculates the confidence margin index corresponding to the prediction result; and performs anomaly detection calculation on the business data and the context information based on a preset anomaly detection model to obtain the corresponding anomaly value; then, it calculates the confidence distribution entropy, the confidence margin index, and the anomaly value based on a preset anomaly score calculation strategy to obtain the corresponding anomaly score; subsequently, it determines whether the anomaly score is greater than a preset anomaly threshold; if so, it generates a first anomaly detection result corresponding to the prediction result that contains an anomaly; otherwise, it generates a second anomaly detection result corresponding to the prediction result that does not contain an anomaly. Based on the above processing flow, this application, by using an anomaly detection agent and fusing model uncertainty (confidence distribution entropy, confidence margin index) and input anomaly (context deviation), can achieve more robust misjudgment detection processing than a single index, effectively improving the accuracy of the generated anomaly detection results.
[0052] In some optional implementations of this embodiment, step S204 includes the following steps: Based on the data review agent, the business data is retrieved using a preset knowledge base to obtain the corresponding knowledge retrieval results.
[0053] In this embodiment, the aforementioned data review agent, also known as the data review agent, performs the cognitive task of "understanding and comparison," conducts semantic review and knowledge comparison on abnormal samples, and verifies the rationality of the model's conclusions. The data review agent combines RAG retrieval, semantic consistency comparison, and counterfactual reasoning to achieve "interpretable cognitive comparison"; it can discover the root causes of model misjudgments (rule inconsistency, semantic drift, causal errors), rather than merely discovering "errors in results."
[0054] The aforementioned knowledge base may include an internal rule base and an external knowledge base. Key entities (such as "Hypertension Class II", "BMI", and "Claim Rejection") and keywords (such as "Payment Standard" and "Medical Impact") can be extracted from business data (e.g., customer characteristics: age = 45 years, BMI = 28; claims history: claim rejection; underwriting description: hypertension class II) and keywords (such as "payment standard", "medical impact"). Then, the keywords and entities are concatenated into a query vector (e.g., "Hypertension Class II Payment Standard, BMI, Medical Impact"), and the cosine similarity between the query vector and each document vector in the knowledge base is calculated. Subsequently, the documents are sorted according to their similarity scores, and the Top-K (e.g., Top-3) most relevant documents (e.g., the "Hypertension Payment Details" in the internal rule base and the external medical literature "BMI and Cardiovascular Disease Correlation") are used as the knowledge retrieval results. Furthermore, the sources of evidence can be recorded based on the knowledge retrieval results, such as medical clause #12 corresponding to Article 12 of the "Hypertension Payment Details," to facilitate subsequent information filling processing of the review results.
[0055] Extract factual data from the knowledge retrieval results.
[0056] In this embodiment, the aforementioned factual data may refer to the aforementioned knowledge retrieval results, i.e., the rules or facts in the Top-K documents.
[0057] Based on the preset first major language model, the prediction results are semantically compared with the factual data to obtain the corresponding semantic comparison results.
[0058] In this embodiment, the aforementioned first large language model can be a general large language model, such as GPT4. The implementation process of semantic comparison based on the first large language model includes: constructing a specified prompt word to request the first large language model to calculate the semantic consistency score. (Range 0-1) This parameter determines whether the model's output conclusion is consistent with the facts in the knowledge base. Example of a specified prompt: "Evaluate the logical consistency between the following two texts (prediction results and factual data), outputting a score of 0-1 (1 for complete consistency): Text 1: 'Model conclusion: Deny payment'; Text 2: 'Factual data in the knowledge base rules: Partial payment'." Example response from the largest language model: "Analysis: Model conclusion contradicts the rules, score 0.2." A pre-set comparison threshold (e.g., 0.5) is used. If the semantic consistency score is less than this threshold, it is considered "rule deviation" (model conclusion conflicts with the knowledge base), and a semantic comparison result with rule deviation content is generated. Otherwise, it is considered consistent.
[0059] The review report fields include: Review Type: "Rule Deviation" (if conflicting) or "Conformity" (if approved). Confidence Level: 0.92 (confidence level can be adjusted from 1- Normalization yields, for example, 0.92 = 1 - 0.08.
[0060] The prediction results are subjected to a causal consistency test to obtain the corresponding test results.
[0061] In this embodiment, the above-mentioned causal consistency test verifies the logical rationality of the model's conclusions through counterfactual reasoning and explains the specific reasons for rule deviations. The specific implementation process includes: 1) constructing counterfactual samples by generating new samples through feature perturbation (such as changing "high blood lipids = yes" to "no"); 2) comparing predicted changes: ,in, The original prediction results. This is a counterfactual sample. Mapped to numerical values (e.g., " For claim denial = 0, If the full compensation is 1", then we get... =-1.3) Medical common sense verification: check Does it conform to common medical causality (e.g., "abnormal BMI should affect compensation decisions")? If If the direction contradicts common sense (e.g., a decrease in payout after a lower BMI), an audit alert will be triggered. 4) Generate an explanation: Combining the knowledge base and counterfactual results, explain the specific reasons for the rule deviation (e.g., "The model did not consider the impact of abnormal BMI, leading to a claim rejection conclusion that conflicts with Article 12 of the 'Hypertension Payout Rules'"). Audit report field population: Explanation: "The model did not consider the impact of abnormal BMI".
[0062] Based on the knowledge retrieval results, the semantic comparison results, and the verification results, a report generation process is performed to obtain the corresponding target report.
[0063] In this embodiment, the generated knowledge retrieval results, semantic comparison results, and test results can be used to extract matching information according to a preset report structure and fill it into a preset review report template. The generated target report is then used as the corresponding review result. The report structure of the aforementioned review report template may include fields such as Case ID, Review Type, Evidence Source, Confidence Level, and Description. For example, the content of the generated target report, i.e., the review result (or review report), may include: {Case ID: 202501, Review Type: "Rule Deviation", Evidence Source: Medical Clause #12, Confidence Level: 0.92, Description: "The model did not consider the impact of abnormal BMI"}.
[0064] The target report is used as the review result.
[0065] This application utilizes a data review agent to perform knowledge retrieval on the business data using a preset knowledge base to obtain corresponding knowledge retrieval results. Factual data is then extracted from these results. Next, a semantic comparison is performed between the predicted results and the factual data based on a preset first language model to obtain corresponding semantic comparison results. Following this, a causal consistency test is performed on the predicted results to obtain corresponding test results. Subsequently, a report generation process is performed based on the knowledge retrieval results, the semantic comparison results, and the test results to obtain a corresponding target report. Finally, the target report is used as the review result. Based on this processing flow, this application, through the use of a data review agent and the triple verification of knowledge retrieval, semantic comparison, and causal consistency testing, can accurately pinpoint the root cause of model misjudgments, ensuring the accuracy of the generated review results and providing accurate structured evidence for subsequent result correction.
[0066] In some alternative implementations, step S205 includes the following steps: Based on the modified intelligent agent, historical similar cases corresponding to the business data are retrieved from a preset case library.
[0067] In this embodiment, the aforementioned corrective agent, also known as the labeling correction agent, simulates the cognitive correction behavior of a "human labeler," automatically generating new labels for abnormal samples that have been reviewed and judged. It can also ensure the credibility of the labels through self-adversarial verification. By combining multi-source consistency verification and adversarial discrimination, the risk of "false correction" caused by model self-labeling is avoided. Furthermore, through the dual constraints of confidence decay and semantic consistency, the quality of the returned samples is made close to the standards of human labeling.
[0068] Specifically, a corrective agent (or labeled corrective agent) can be used to calculate the semantic similarity between the original business data and each reference case in the aforementioned case library (e.g., through cosine similarity of embedded vectors). Then, a Top-K search is performed to select a specified number (e.g., 5) of cases with the highest similarity as the aforementioned historical similar cases. The aforementioned case library is a pre-built database used to store historical business cases, such as historical claims cases.
[0069] Based on a pre-defined second language model, the historical similar cases and the review results are analyzed and processed to generate multiple corresponding candidate results.
[0070] In this embodiment, the second large language model can be a general large language model, such as GPT4. A prompt word can be constructed to instruct the second large language model to generate a corresponding candidate result set (including multiple candidate results) based on the review results and historical similar cases. An example prompt word is: "Generate reasonable claims decision labels (full pay / partial pay / rejection) based on the following information: 1. The customer has high blood pressure, BMI 31; 2. Knowledge base rule: partial pay is required when BMI>30; 3. Historical cases: similar cases are mostly judged as partial pay." Accordingly, the candidate label set output by the second large language model (which may contain a probability distribution) is: Candidate Labels: {Full Pay: 0.1, Partial Pay: 0.8, Rejection: 0.1}. The candidate results can include the label results corresponding to the aforementioned historical similar cases.
[0071] Based on a preset multi-source consistency verification strategy, the similarity of each candidate result is calculated to obtain multiple corresponding similarity scores.
[0072] In this embodiment, the specific implementation process of performing similarity calculation on each candidate result based on the preset multi-source consistency verification strategy to obtain the corresponding multiple similarity scores will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0073] Select the target candidate with the highest similarity score from all candidate results.
[0074] In this embodiment, the similarity scores of each candidate result can be compared numerically to obtain the specific similarity score with the largest value. Then, the specific result corresponding to the specified similarity score can be extracted from all candidate results and used as the target candidate result.
[0075] The target candidate results are used as the target prediction results.
[0076] This application, based on the modified intelligent agent, queries historical similar cases corresponding to the business data from a preset case library; then, based on a preset second language model, it analyzes and processes the historical similar cases and the review results to generate multiple candidate results; subsequently, based on a preset multi-source consistency verification strategy, it performs similarity calculation on each candidate result to obtain multiple similarity scores; next, it selects the target candidate result with the highest similarity score from all candidate results; finally, it uses the target candidate result as the target prediction result. Based on the above processing flow, this application, by using the modified intelligent agent to query historical similar cases corresponding to the business data from a preset case library, and using the second language model to analyze and process the historical similar cases and review results to generate multiple candidate results, and then using the multi-source consistency verification strategy to perform similarity calculation on each candidate result, and then selecting the target candidate result with the highest similarity score from all candidate results as the final target prediction result, effectively ensures the reliability and accuracy of the obtained target prediction result.
[0077] In some optional implementations, the anomaly detection agent, data review agent, and annotation correction agent in this application constitute a "cognitive closed-loop" working mechanism. The three are responsible for different levels of cognitive tasks (perception → understanding → correction) and exchange knowledge and experience through a shared cognitive memory pool.
[0078] The collaboration and information flow among the three agents include: 1. Anomaly detection agent → marking potential problem samples; 2. Data review agent → retrieving rules and knowledge, verifying and generating explanatory reports; 3. Labeling and correction agent → generating, verifying, and outputting new labels; 4. All decision records are written to the cognitive memory pool: storing the reasons for each misjudgment, review logic, correction labels, and effects; subsequent agents can call upon this memory to achieve "experience sharing," such as automatically identifying similar problems, and can use it for subsequent strategy optimization. This cognitive memory pool is equivalent to the system's "experience brain," which can continuously evolve.
[0079] The system also includes a reinforcement tuning controller, which acts as the system's "reinforcement learning brain." This controller determines the model tuning strategy based on feedback signals from the cognitive layer. Specific implementations include: 1. Reward Mechanism Design. A collaborative optimization mechanism is implemented, with each agent sharing a reward signal. .
[0080] in: : Improvement in accuracy after retraining; Improved predictive consistency; The proportion of noise introduced by error correction. The enhancement and optimization controller uses this to calculate the overall reward value, which is used to evaluate the correction effect of each agent and dynamically adjust the sample sampling strategy.
[0081] 2. Policy Update. The Proximal Policy Optimization (PPO) algorithm is used to adjust the cooperation strategy among agents, reinforcing corrective paths that bring positive gains. If the reward is below a threshold, a "freeze mode" is entered, temporarily suspending automatic backflow to avoid false reinforcement.
[0082] In some optional implementations, the similarity calculation process for each candidate result based on a preset multi-source consistency verification strategy to obtain multiple corresponding similarity scores includes the following steps: Calculate the first similarity between a specified candidate result and a preset knowledge base rule; wherein the specified candidate result is any one of all the candidate results.
[0083] In this embodiment, the aforementioned knowledge base rules are a standardized set of rules pre-built based on the business domain knowledge system. Their core purpose is to provide clear business constraints for label correction. The generation of knowledge base rules typically relies on the following two approaches, combining business expert experience with historical data: 1. Manually coded explicit rules. Source: Rules manually written by business experts (such as those in law, insurance, and medical fields) based on industry standards, laws and regulations, and internal policies. Example: Insurance claims scenario: Article 12 of the "Claims Terms" clearly stipulates that "when hypertension is grade II and BMI ≥ 28, the reimbursement ratio is reduced to 70%". The rule can be expressed as: IF (Disease grade == "Hypertension Grade II" AND BMI ≥ 28) THEN Reimbursement label ∈ {"Partial Reimbursement"}. Medical diagnosis scenario: In the International Classification of Diseases (ICD) standard, "diabetes mellitus with nephropathy" corresponds to a specific diagnostic label. The characteristics of this generation method include: clear rules and strong interpretability, but limited coverage scenarios, requiring regular updates to adapt to policy changes. 2. Implicit rules mined from historical data. Source: By analyzing a large number of historical revision examples, frequently occurring association patterns between labels and features are automatically summarized. Method: Association rule mining (e.g., Apriori algorithm): Extracting strong association rules of "feature combination → label" from historical revision records. Example: In 80% of the "Hypertension Grade II + BMI=28" cases, the label is revised to "partial compensation," then the generated rule is: (Hypertension Grade II ∧ BMI=28) → Partial Compensation (Support = 80%, Confidence = 90%). Decision Tree / Random Forest: After training the model, feature importance and splitting rules are extracted and transformed into interpretable business rules. This generation method has the characteristics of wider scenario coverage, but requires manual review of the rule's rationality (e.g., to avoid erroneous inductions caused by data bias).
[0084] When calculating the similarity between a specified candidate result and the rules in the knowledge base, the system checks whether the candidate label matches the rules in the knowledge base. The higher the matching degree, the higher the score. The specific logic is as follows: 1) Rule matching: For each candidate result, check whether it meets the condition part of all relevant rules in the knowledge base (e.g., "Hypertension Grade II + BMI≥28"). If it does, then determine whether the candidate result is compliant based on the conclusion part of the rule (e.g., "Compensation label ∈ {"Partial Compensation"}). Similarity calculation: Complete match: If the candidate result completely matches the conclusion of a rule, then the similarity = 1. Partial match: If the candidate result meets the conclusion of multiple rules (e.g., simultaneously meeting "Partial Compensation" and "Supplementary materials required"), then take the average or weighted average. No match: If the candidate result conflicts with all rules, then the similarity = 0.
[0085] Calculate the second similarity between the specified candidate result and the explanatory vector of the business model.
[0086] In this embodiment, an explanatory vector corresponding to the prediction result output by the business model can be obtained. For example, if the original prediction result output by the business model is "claim rejection," but the explanatory vector implies "abnormal BMI may affect payout," then the similarity (such as cosine similarity) between the specified candidate result and the explanatory vector of the business model can be calculated to obtain the corresponding second similarity.
[0087] Calculate the third similarity between the specified candidate result and the label results of the historical similar cases.
[0088] In this embodiment, the semantic similarity (such as BERT embedding vector distance) between the specified candidate result and the label result of the historical similar case can be calculated by obtaining the label results of the aforementioned historical similar cases, and then used as the corresponding third similarity.
[0089] The first similarity, the second similarity, and the third similarity are weighted and calculated based on a preset comprehensive calculation formula to obtain the corresponding calculation results.
[0090] In this embodiment, the above comprehensive calculation formula includes: .in, The first similarity score, As the first weight corresponding to the knowledge base rules, For the second similarity, The second weight corresponds to the explanatory vector of the business model. The third similarity This is the third weight corresponding to the label results of similar historical cases. The values of the first, second, and third weights can be determined through historical data grid search; for example, the first, second, and third weights can be set to 0.4, 0.3, and 0.3 respectively. Alternatively, they can be adjusted according to business priorities, such as giving the highest weight to knowledge base rules.
[0091] Furthermore, the first similarity, second similarity, and third similarity can be substituted into the corresponding positions in the comprehensive calculation formula for weighted calculation, and the resulting calculation result can be used as the specified similarity score of the specified candidate result.
[0092] The calculation result is used as the specified similarity score for the specified candidate result.
[0093] This application calculates a first similarity between a specified candidate result and a preset knowledge base rule; wherein the specified candidate result is any one of all candidate results; simultaneously calculates a second similarity between the specified candidate result and the explanatory vector of the business model; and calculates a third similarity between the specified candidate result and the label results of historical similar cases; then, based on a preset comprehensive calculation formula, the first similarity, the second similarity, and the third similarity are weighted and calculated to obtain the corresponding calculation result; subsequently, the calculation result is used as the specified similarity score of the specified candidate result. Based on the above processing flow, this application, by using a comprehensive calculation formula, performs weighted calculation on the first similarity between the specified candidate result and the preset knowledge base rule, the second similarity between the specified candidate result and the explanatory vector of the business model, and the third similarity between the specified candidate result and the label results of the historical similar cases, and uses the obtained calculation result as the specified similarity score of the specified candidate result, thereby realizing multi-source and multi-dimensional similarity calculation processing of candidate results, ensuring the accuracy of the generated similarity score.
[0094] In some optional implementations of this embodiment, step S206 includes the following steps: Call the preset verification model.
[0095] In this embodiment, the verification model described above is a pre-trained model specifically designed to determine the consistency between "label and sample features". .
[0096] Based on the verification model, the business data and the target prediction results are processed to calculate the verification loss and obtain the corresponding loss value.
[0097] In this embodiment, business features can be extracted from business data, and the business features and the corrected target prediction results can be input into the above-mentioned verification model to calculate the cross-entropy loss: To obtain the corresponding loss value. As a business characteristic, The target prediction result.
[0098] Determine whether the loss value is less than a preset threshold.
[0099] In this embodiment, the value of the preset threshold is not specifically limited and can be set according to actual business needs, for example, it can be set to 0.2.
[0100] If yes, the target prediction result is deemed to have passed verification; otherwise, the target prediction result is deemed to have failed verification.
[0101] In this embodiment, if the calculated loss value is less than the preset threshold, the label (target prediction result) is considered to be consistent with the feature (business feature), and the target prediction result is deemed to have passed verification. However, if the calculated loss value is greater than or equal to the preset threshold, the label is considered to be consistent with the feature, the target prediction result is deemed to have failed verification, and the target prediction result will be returned to the data review agent for further review.
[0102] The system also features a function to reweight the confidence levels of the corrected target predictions to control correction risks and avoid overconfidence. The goal is to introduce dynamic confidence levels into the corrected labels, preventing over-correction due to review risks or data noise. The specific operation process includes: original confidence level (…). The initial revised label (target prediction result) may have a high confidence level (e.g., 0.95), but needs to be adjusted based on review risk. Confidence decay coefficient ( The risk factor is determined by the review process (e.g., when business data involves high-amount claims or ambiguous terms). (Higher). The formula is: Subsequently, the system will output high-confidence corrected label samples and confidence reports, such as {Case ID: 202501, New Label: “Partial Compensation”, Confidence: 0.91, Verification Passed: True}.
[0103] This application utilizes a pre-defined verification model. Based on this model, it calculates the verification loss for the business data and the target prediction result to obtain a corresponding loss value. Subsequently, it determines whether the loss value is less than a pre-defined threshold. If so, the target prediction result is deemed to have passed verification; otherwise, it is deemed to have failed verification. Based on this process, this application uses a verification model to calculate the verification loss for the business data and the target prediction result, obtaining a corresponding loss value. Then, by comparing the loss value with a pre-defined threshold, the application efficiently and accurately completes the verification of the target prediction result, ensuring the accuracy of the obtained verification result.
[0104] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps: Call the preset data flywheel execution module.
[0105] In this embodiment, the aforementioned data flywheel execution module (or data flywheel execution unit) is the core module for the system to achieve "continuous learning and self-evolution". It constructs a closed loop from sample feedback → model optimization → experience accumulation, that is, a closed-loop automated process from high-value sample feedback to retraining, which enables the model to continuously iterate and upgrade based on high-value feedback in real business scenarios.
[0106] The data flywheel execution module filters out target sample data that meets preset reflux conditions.
[0107] In this embodiment, the data flywheel execution module described above can be used to filter out truly valuable target sample data from a massive amount of corrected samples, thus avoiding "junk data" contaminating the model. The above-mentioned backflow conditions include: 1) >0.8: The corrected label confidence level must be higher than the threshold (e.g., 0.8) to ensure label reliability. For example, if a sample is corrected to "partial compensation" and the corrected confidence level is 0.91 (adjusted through multi-source validation and confidence decay), then it meets the condition. 2) >δ: The outlier score of a sample must exceed a preset threshold (δ) to ensure its business representativeness. For example, in claims review, a sample with "Hypertension Grade II + BMI=28" may be prioritized for retraining due to its significant difference from regular cases (high outlier score). Result: Only samples that simultaneously meet the above retraining conditions will enter the retraining pool as "fuel" for model optimization.
[0108] The business model is trained using lightweight fine-tuning based on the target sample data to obtain the corresponding target business model.
[0109] In this embodiment, lightweight fine-tuning (LoRA / Adapter) can be used to incrementally update and train the aforementioned business model using selected target sample data. This involves updating only some parameters of the model (e.g., LoRA injects new knowledge through low-rank decomposition), while maintaining the original model structure. This significantly reduces computational resource consumption and yields a finely tuned target business model. For example, in an insurance claims model, only the parameter layer related to "hypertension compensation rules" is fine-tuned, while other layers remain unchanged.
[0110] The target business model is subjected to comparative testing based on a preset testing strategy.
[0111] In this embodiment, the aforementioned testing strategy can specifically employ A / B testing. The effectiveness can be verified by performing A / B testing on the target business model using a validation set. Specifically, the new target business model (version B) and the old business model (version A) can be compared and tested on the validation set to evaluate key metrics (such as accuracy, F1 score, and business rule coverage). For example, if the "rule deviation rate" of version B decreases from 5% to 2% on the validation set, and the inference speed does not significantly decrease, then performance is considered improved. The validation set can be obtained by randomly selecting a specified proportion (e.g., 0.5) of data from the aforementioned target sample data.
[0112] If the target business model passes the comparison test, then the business model is updated based on the target business model.
[0113] In this embodiment, if the performance improvement of the new target business model is detected to exceed a threshold (e.g., an accuracy improvement of ≥2%), the target business model is determined to have passed the comparative test. The target business model is then used to replace the old business model and is deployed in the production environment to complete the model update process for the business model. Otherwise, the old business model is retained and the reasons are analyzed (e.g., whether the sample selection criteria need to be adjusted).
[0114] This application utilizes a pre-defined data flywheel execution module to filter target sample data that meets pre-defined feedback conditions. Then, it performs lightweight fine-tuning training on the business model based on the target sample data to obtain a corresponding target business model. Subsequently, it conducts comparative testing on the target business model based on a pre-defined testing strategy. If the target business model passes the comparative test, it updates the business model based on the target business model. Based on this processing flow, this application, through the use of the data flywheel execution module, can automatically and intelligently realize a closed-loop automated process from high-value sample feedback to model retraining. This allows the business model to continuously iterate and upgrade based on high-value feedback from real business scenarios, effectively improving the intelligence of model updates and enhancing the model's predictive performance.
[0115] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0116] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0117] Furthermore, this application integrates multi-agent cognitive reasoning, RAG knowledge review, and reinforcement learning optimization to construct a self-examining and self-correcting data flywheel system. In the financial insurance field, this technology not only significantly improves model reliability and interpretability but also drives the system from "passive learning" to "active self-evolution," demonstrating strong innovation and application value. Improvements and advantages: 1. Automatic cognitive correction mechanism: By simulating human review through multiple agents, it automates error identification, cause explanation and label correction, significantly reducing human intervention.
[0118] 2. High-quality data feedback: Combining confidence screening and adversarial verification ensures the purity of retraining data and avoids noise pollution.
[0119] 3. Continuous self-evolution of the model: By strengthening the tuning controller, the model automatically repairs itself before performance degrades, maintaining long-term stability.
[0120] 4. Cross-business portability: The framework is applicable to tasks such as underwriting, claims, customer service quality inspection, etc., and only the knowledge base and judgment rules need to be changed.
[0121] 5. Significantly reduce costs and increase efficiency: The system can complete the "self-correction-self-learning" closed loop without human intervention, which shortens the model iteration cycle of insurance institutions and reduces the error rate.
[0122] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0123] It should be emphasized that, in order to further ensure the privacy and security of the above target prediction results, the above target prediction results can also be stored in a blockchain node.
[0124] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0125] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0126] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0127] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0128] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data processing device based on artificial intelligence, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0129] like Figure 3 As shown, the artificial intelligence-based data processing device 300 described in this embodiment includes: an inference module 301, an acquisition module 302, a first detection module 303, a second detection module 304, a correction module 305, a verification module 306, and an output module 307. Wherein: The inference module 301 is used to acquire business data to be processed and to perform inference processing on the business data based on a preset business model to obtain the corresponding prediction result. The acquisition module 302 is used to acquire context information corresponding to the business data; The first detection module 303 is used to perform anomaly detection processing on the prediction result based on the context information using a preset anomaly detection agent, and obtain the corresponding anomaly detection result. The second detection module 304 is used to perform reasonableness detection processing on the business data based on a preset data review intelligent agent if the anomaly detection result is that there is an anomaly, and obtain the corresponding review result. The correction module 305 is used to perform result correction processing on the business data based on the review result using a preset correction agent to obtain the corresponding target prediction result; The verification module 306 is used to verify the target prediction result based on a preset verification strategy. The output module 307 is used to output the target prediction result if the target prediction result passes the verification.
[0130] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0131] In some optional implementations of this embodiment, the first detection module 303 includes: The first calculation submodule is used to calculate the confidence distribution entropy corresponding to the prediction result based on the anomaly detection agent. The second calculation submodule is used to calculate the confidence margin index corresponding to the prediction result; The third calculation submodule is used to perform anomaly detection calculation processing on the business data and the context information based on a preset anomaly detection model to obtain the corresponding anomaly value. The fourth calculation submodule is used to calculate and process the confidence distribution entropy, the confidence margin index and the abnormal value based on the preset anomaly score calculation strategy to obtain the corresponding anomaly score. The first judgment submodule is used to determine whether the abnormal score is greater than a preset abnormal threshold. The first generation submodule is used to generate a first anomaly detection result corresponding to the prediction result if the anomaly exists, and otherwise generate a second anomaly detection result corresponding to the prediction result if the anomaly does not exist.
[0132] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0133] In some optional implementations of this embodiment, the second detection module 304 includes: The retrieval submodule is used to perform knowledge retrieval on the business data based on the data review agent and use a preset knowledge base to obtain the corresponding knowledge retrieval results. An extraction submodule is used to extract factual data from the knowledge retrieval results; The comparison submodule is used to perform semantic comparison between the prediction results and the factual data based on a preset first language model, and obtain the corresponding semantic comparison results. The verification submodule is used to perform causal consistency verification on the prediction results and obtain the corresponding verification results. The second generation submodule is used to perform report generation processing based on the knowledge retrieval results, the semantic comparison results, and the verification results to obtain the corresponding target report; The first determining submodule is used to take the target report as the review result.
[0134] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0135] In some optional implementations of this embodiment, the correction module 305 includes: The query submodule is used to query historical similar cases corresponding to the business data from a preset case library based on the correction agent. The analysis submodule is used to analyze and process the historical similar cases and the review results based on the preset second language model, and generate multiple corresponding candidate results; The fifth calculation submodule is used to perform similarity calculation on each candidate result based on a preset multi-source consistency verification strategy to obtain multiple corresponding similarity scores; The filtering submodule is used to filter out the target candidate result with the highest similarity score from all candidate results; The second determining submodule is used to use the target candidate result as the target prediction result.
[0136] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0137] In some optional implementations of this embodiment, the fifth calculation submodule includes: The first calculation unit is used to calculate the first similarity between a specified candidate result and a preset knowledge base rule; wherein, the specified candidate result is any one of all the candidate results; The second calculation unit is used to calculate the second similarity between the specified candidate result and the explanatory vector of the business model; The third calculation unit is used to calculate the third similarity between the specified candidate result and the label result of the historical similar cases; The fourth calculation unit is used to perform weighted calculation processing on the first similarity, the second similarity and the third similarity based on a preset comprehensive calculation formula to obtain the corresponding calculation result; A determining unit is used to use the calculation result as a specified similarity score for the specified candidate result.
[0138] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here. In some optional implementations of this embodiment, the verification module 306 includes: Call the submodule to invoke the preset verification model; The sixth calculation submodule is used to perform verification loss calculation processing on the business data and the target prediction result based on the verification model to obtain the corresponding loss value; The second judgment submodule is used to determine whether the loss value is less than a preset threshold. The determination submodule is used to determine if the target prediction result passes verification if so, and otherwise determine if the target prediction result fails verification.
[0139] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0140] In some optional implementations of this embodiment, the artificial intelligence-based data processing device further includes: The calling module is used to invoke the preset data flywheel execution module; The filtering module is used to filter target sample data that meet preset reflux conditions based on the data flywheel execution module; The fine-tuning module is used to perform lightweight fine-tuning training on the business model based on the target sample data to obtain the corresponding target business model. The testing module is used to conduct comparative tests on the target business model based on a preset testing strategy. The update module is used to update the business model based on the target business model if the target business model passes the comparison test.
[0141] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here. To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0142] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0143] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0144] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0145] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based data processing method.
[0146] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0147] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based data processing method described above.
[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0149] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A data processing method based on artificial intelligence, characterized in that, Includes the following steps: The system acquires the business data to be processed and performs inference processing on the business data based on a preset business model to obtain the corresponding prediction results. Obtain the context information corresponding to the business data; Based on the context information, a preset anomaly detection agent is used to perform anomaly detection processing on the prediction result to obtain the corresponding anomaly detection result. If the anomaly detection result indicates the presence of an anomaly, the business data will be processed for reasonableness detection based on a preset data review intelligence agent to obtain the corresponding review result. Based on the review results, a preset correction agent is used to perform result correction processing on the business data to obtain the corresponding target prediction results. The target prediction results are verified based on a preset verification strategy. If the target prediction result passes verification, the target prediction result is then output.
2. The data processing method based on artificial intelligence according to claim 1, characterized in that, The step of using a preset anomaly detection agent to perform anomaly detection processing on the prediction result based on the context information to obtain the corresponding anomaly detection result specifically includes: Based on the anomaly detection agent, the confidence distribution entropy corresponding to the prediction result is calculated; Calculate the confidence margin index corresponding to the prediction result; Based on a preset anomaly detection model, anomaly detection calculations are performed on the business data and the context information to obtain the corresponding anomaly values. The confidence distribution entropy, the confidence margin index, and the abnormal value are calculated and processed based on the preset anomaly score calculation strategy to obtain the corresponding anomaly score. Determine whether the anomaly score is greater than a preset anomaly threshold; If yes, generate a first anomaly detection result corresponding to the prediction result; otherwise, generate a second anomaly detection result corresponding to the prediction result.
3. The data processing method based on artificial intelligence according to claim 1, characterized in that, The step of performing reasonableness detection processing on the business data based on a preset data review intelligent agent to obtain the corresponding review result specifically includes: Based on the data review agent, the business data is retrieved using a preset knowledge base to obtain the corresponding knowledge retrieval results. Extract factual data from the knowledge retrieval results; Based on the preset first language model, the prediction results are semantically compared with the factual data to obtain the corresponding semantic comparison results; The prediction results are subjected to a causal consistency test to obtain the corresponding test results; Based on the knowledge retrieval results, the semantic comparison results, and the verification results, a report generation process is performed to obtain the corresponding target report; The target report is used as the review result.
4. The data processing method based on artificial intelligence according to claim 1, characterized in that, The step of using a preset correction agent to correct the business data based on the review results to obtain the corresponding target prediction results specifically includes: Based on the corrective agent, historical similar cases corresponding to the business data are retrieved from a preset case library; Based on a pre-defined second language model, the historical similar cases and the review results are analyzed and processed to generate multiple corresponding candidate results. Based on a preset multi-source consistency verification strategy, the similarity of each candidate result is calculated to obtain multiple corresponding similarity scores. Select the target candidate with the highest similarity score from all candidate results; The candidate target results are used as the target prediction results.
5. The data processing method based on artificial intelligence according to claim 4, characterized in that, The step of calculating the similarity of each candidate result based on a preset multi-source consistency verification strategy to obtain multiple corresponding similarity scores specifically includes: Calculate the first similarity between a specified candidate result and a preset knowledge base rule; wherein, the specified candidate result is any one of all the candidate results; Calculate the second similarity between the specified candidate result and the explanatory vector of the business model; Calculate the third similarity between the specified candidate result and the label results of the historical similar cases; The first similarity, the second similarity, and the third similarity are weighted and calculated based on a preset comprehensive calculation formula to obtain the corresponding calculation results; The calculation result is used as the specified similarity score for the specified candidate result.
6. The data processing method based on artificial intelligence according to claim 1, characterized in that, The step of verifying the target prediction result based on a preset verification strategy specifically includes: Invoke the preset verification model; Based on the verification model, the business data and the target prediction results are subjected to verification loss calculation to obtain the corresponding loss value; Determine whether the loss value is less than a preset threshold; If yes, the target prediction result is deemed to have passed verification; otherwise, the target prediction result is deemed to have failed verification.
7. The data processing method based on artificial intelligence according to claim 1, characterized in that, After the step of outputting the target prediction result, the method further includes: Invoke the preset data flywheel execution module; The data flywheel execution module filters out target sample data that meets the preset reflux conditions. The business model is trained with lightweight fine-tuning based on the target sample data to obtain the corresponding target business model. The target business model is subjected to comparative testing based on a preset testing strategy; If the target business model passes the comparison test, then the business model is updated based on the target business model.
8. A data processing device based on artificial intelligence, characterized in that, include: The inference module is used to acquire the business data to be processed and to perform inference processing on the business data based on a preset business model to obtain the corresponding prediction results. The acquisition module is used to acquire context information corresponding to the business data; The first detection module is used to perform anomaly detection processing on the prediction result based on the context information using a preset anomaly detection agent, and obtain the corresponding anomaly detection result. The second detection module is used to perform reasonableness detection processing on the business data based on a preset data review intelligent agent if the anomaly detection result is that there is an anomaly, and obtain the corresponding review result. The correction module is used to correct the business data based on the review results using a preset correction agent to obtain the corresponding target prediction results. The verification module is used to verify the target prediction results based on a preset verification strategy. The output module is used to output the target prediction result if the target prediction result passes the verification.
9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data processing method based on artificial intelligence as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data processing method based on artificial intelligence as described in any one of claims 1 to 7.