Data auditing method and device, computer equipment and storage medium

By using automated data auditing methods and semantic recognition and rule matching technologies, the accuracy and efficiency issues of traditional data auditing have been resolved. This enables precise risk assessment of target users and multi-dimensional data fusion in financial business, adapting to the needs of multiple scenarios.

CN120852030APending Publication Date: 2025-10-28IND CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510751472.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional data audit methods make it difficult to achieve personalized assessments in financial services, resulting in inaccurate and inefficient audit results, increased credit risks and labor costs.

Method used

By acquiring basic data of target users, performing category matching and semantic recognition, generating a question set, transforming non-standard question and answer data into standardized question and answer data, and adjusting the word set using an exponentially weighted moving average algorithm, the question output order is optimized to achieve automated data review.

Benefits of technology

It achieves accurate and efficient data review, reduces subjective judgment errors, supports parallel processing of massive users, reduces time and labor costs, adapts to business changes, and covers multi-scenario applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852030A_ABST
    Figure CN120852030A_ABST
Patent Text Reader

Abstract

The invention relates to a data auditing method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring basic data of a target user; determining the category of the target user according to the basic data, and matching the category of the target user with a preset rule to obtain a matching result; generating a corresponding question set based on the matching result and the verbal skill set; obtaining corresponding to-be-audited data based on the problem set, and performing semantic recognition processing on the to-be-audited data to obtain alternative data; converting non-standardized question and answer data in the alternative data into standardized question and answer data to obtain target data; obtaining an audit result of the target user based on the target data; the auditing result comprises the default probability. According to the method, the time cost and the labor cost are reduced through automation of processes such as data acquisition, rule matching and semantic recognition. And moreover, multi-dimensional data fusion is realized, subjective judgment errors are reduced, and accurate judgment of the auditing result is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial data processing technology, and in particular to a data auditing method, apparatus, computer equipment, and storage medium. Background Technology

[0002] With the rapid development of computer technology and fintech, data verification, as a crucial link in risk control during financial transactions, directly impacts the sound operation of financial institutions through its scientific and efficient processes. For a long time, traditional data verification has relied on keyword matching and manual review. However, this model has revealed several limitations in practical application.

[0003] Regarding the accuracy of the review process, traditional methods struggle to provide personalized assessments based on each customer's unique financial situation, credit history, and borrowing needs. This significantly reduces the accuracy and reasonableness of the review results and may expose financial institutions to additional credit risks. In terms of efficiency, manual review requires substantial time and manpower, making it difficult to cope with the rapid growth of financial institutions' business, and highlighting the increasingly prominent problem of low review efficiency.

[0004] Therefore, how to conduct accurate and efficient data auditing of financial data is an urgent problem to be solved. Summary of the Invention

[0005] Therefore, it is necessary to provide a data auditing method, apparatus, computer equipment, and storage medium that can accurately assess risks in response to the above-mentioned technical problems.

[0006] Firstly, this application provides a data verification method, including:

[0007] Obtain basic data of the target user;

[0008] The category of the target user is determined based on the basic data, and the category of the target user is matched with preset rules to obtain the matching result;

[0009] Based on the matching results and the set of dialogue options, a corresponding set of questions is generated;

[0010] Based on the set of questions, obtain the corresponding data to be reviewed, and perform semantic recognition processing on the data to be reviewed to obtain candidate data;

[0011] The non-standardized question-and-answer data in the candidate data is transformed into standardized question-and-answer data to obtain the target data;

[0012] The review results for the target user are obtained based on the target data; the review results include the probability of default.

[0013] In one embodiment, the method further includes:

[0014] Based on the audit results, a validity index is determined for each question in the set of scripts; the validity index is used to characterize the degree of influence of the question on the audit results.

[0015] The weights for each question are adjusted based on the exponentially weighted moving average algorithm and the aforementioned effectiveness index.

[0016] The set of dialogues is updated based on the adjusted weights to obtain the updated set of dialogues.

[0017] In one embodiment, the method further includes:

[0018] During the review process of obtaining the review result based on the data to be reviewed, the current review status is determined; the current review status is used to characterize the current stage of the review.

[0019] Based on the current review status, the output order of each question in the script set is optimized to obtain an optimized script set.

[0020] In one embodiment, the data to be reviewed includes voice data and first text data; the step of obtaining the corresponding data to be reviewed based on the question set and performing semantic recognition processing on the data to be reviewed to obtain candidate data includes:

[0021] The speech data is semantically recognized and converted into second text data, and the first text data and the second text data are concatenated and / or weighted and fused to obtain candidate data;

[0022] The candidate data is obtained by performing semantic recognition processing on the candidate data.

[0023] In one embodiment, the step of performing semantic recognition processing on the candidate data to obtain the alternative data includes:

[0024] The candidate data is identified based on a named entity recognition model to extract target entities;

[0025] If the target entity includes a numerical entity, determine the numerical range corresponding to the numerical entity and / or convert the currency unit of the numerical entity to obtain the alternative data.

[0026] In one embodiment, the method further includes:

[0027] The first text data and the second text data are compared to obtain the anomaly prediction result;

[0028] If the anomaly prediction result indicates that the data to be reviewed is abnormal, an anomaly warning message is output; the anomaly warning message includes abnormal data items and warning prompts.

[0029] In one embodiment, the step of performing semantic recognition processing on the candidate data to obtain the alternative data includes:

[0030] The candidate data acquired each time is subjected to semantic recognition processing to obtain semantic recognition results;

[0031] Based on the semantic recognition results, intent classification and labeling are performed to obtain candidate data for labeled intent classification;

[0032] According to the intent classification, follow-up questions are matched from the preset set of follow-up questions, the follow-up questions are updated to the question set, and the candidate data is updated based on the data to be reviewed obtained from the follow-up questions to obtain the alternative data.

[0033] Secondly, this application also provides a data verification device, the device comprising:

[0034] The acquisition module is used to acquire basic data of the target user;

[0035] The matching module is used to determine the category of the target user based on the basic data, and match the category of the target user with preset rules to obtain a matching result;

[0036] The generation module is used to generate a corresponding set of questions based on the matching results and the set of dialogues;

[0037] The identification module is used to obtain the corresponding data to be reviewed based on the set of questions, and to perform semantic recognition processing on the data to be reviewed to obtain candidate data;

[0038] The conversion module is used to convert non-standardized question-and-answer data in the candidate data into standardized question-and-answer data to obtain the target data;

[0039] The review module is used to obtain the review result of the target user based on the target data; the review result includes the probability of default.

[0040] Thirdly, this application also provides a computer device, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in any embodiment of this application.

[0041] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any embodiment of this application.

[0042] The aforementioned data review method automates processes such as data collection, rule matching, and semantic recognition, replacing manual item-by-item review. This reduces processing time per target user, lowers both time and labor costs, and supports parallel processing of massive numbers of users. Furthermore, the preset rules can be flexibly expanded to quickly respond to business changes, thus covering multiple scenarios (such as loan, insurance, or leasing services). Moreover, by integrating basic user data (occupation, age, and credit reports, etc.) with data to be reviewed, multi-dimensional data fusion is achieved, enabling the construction of a comprehensive risk profile of the target user. Standardized quantification of non-standardized Q&A data reduces subjective judgment errors and achieves accurate review results. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating a data auditing method according to an exemplary embodiment;

[0045] Figure 2 This is a flowchart illustrating a data auditing method according to an exemplary embodiment;

[0046] Figure 3 This is a structural block diagram of a data verification device according to an exemplary embodiment;

[0047] Figure 4 This is an internal structural diagram of a computer device according to an exemplary embodiment. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0049] The terms "first," "second," and "third" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "at least one" is used to indicate one or more; "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0050] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0051] The data verification method provided in this application can be applied to computer devices, which can be mobile terminals or fixed terminals. For example, the terminal can include, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.

[0052] In one exemplary embodiment, such as Figure 1 As shown, a data auditing method is provided. Taking the application of this method to the computer device in the above embodiment as an example, the method includes the following steps S101 to S106. Wherein:

[0053] Step S101: Obtain the basic data of the target user.

[0054] In this embodiment, the basic data may include, but is not limited to, user age, user occupation, user gender, user education level, user income, user credit transaction records, public records, product requirements, and historical review data. Public records may include administrative and legal penalty records; product requirements may include, but are not limited to, at least one of loan amount, loan term, and loan interest rate; historical review data may include, but is not limited to, historical review questions, historical answer texts, and historical review results.

[0055] In some embodiments, obtaining the target user's basic data includes:

[0056] Obtain the initial basic data of the target user;

[0057] Based on the data type, the initial basic data is preprocessed to obtain the basic data.

[0058] In this embodiment, preprocessing refers to the process of collecting, cleaning, and transforming the raw data before analysis or modeling. Preprocessing may include, but is not limited to, at least one of data integration, data cleaning, data standardization, and data reduction.

[0059] Optionally, the data type is user age; the computer device can normalize the user age data to obtain user age data in the range [0,1].

[0060] Optionally, the data type is the user's occupation; the computer device can perform one-hot encoding on the user's occupation data, converting each occupation category data into a binary vector.

[0061] Optionally, the data type is loan amount; the computer equipment can scale the loan amount data proportionally to make the loan amount data fall within the same range as other types of data.

[0062] Optionally, the data type is the data interest rate; the computer device can normalize the data interest rate data to obtain the data interest rate data in the range [0,1].

[0063] Optionally, the data type is historical review questions and / or historical answer text; the computer device can use the bag-of-words model and / or the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm to convert the data into a numerical vector.

[0064] Optionally, the data type is historical review results; the computer device can mark the historical review result data as "pass" or "reject".

[0065] In some embodiments, the computer device can perform feature extraction on the preprocessed basic data to obtain at least one of numerical features, classification features, time series features, and text features; for example, the computer device uses word embedding technology to map each word in the basic data into a low-dimensional vector, and converts the low-dimensional vector into a fixed-length feature vector through methods such as average pooling or max pooling.

[0066] Step S102: Determine the category of the target user based on the basic data, and match the category of the target user with a preset rule to obtain a matching result.

[0067] In this application embodiment, the target user category may include, but is not limited to, at least one of high-risk users, low-risk users, micro and small enterprise users, medium-sized enterprise users, large enterprise users, and individual users.

[0068] In some embodiments, financial professionals / experts, based on industry experience and risk assessment standards, set different preset rules for different categories of users and business requirements. For example, if the target user category is high-risk users, the preset rule could be to add project risk assessment questions; or, if the target user category is high-limit applicants, the preset rule could be to add asset verification questions.

[0069] In one embodiment, a computer device can use a clustering algorithm to perform data analysis on the basic data of target users to determine the category of the target users; and match the target user category with preset rules to obtain a matching result (such as the preset matching rules). For example, if the target user categories are micro and small enterprise users and high-amount applicants, then the preset matching rules in the matching result may include adding questions about asset verification, adding questions about business expansion plans, and / or adding questions about the purpose of funds.

[0070] Step S103: Based on the matching results and the set of dialogues, generate a corresponding set of questions.

[0071] In this embodiment, the script set can be a standardized language framework (general script template) pre-defined by financial institutions for customer qualification review, risk control, and compliance communication. The script set may include, but is not limited to, at least one of the following: identity verification scripts, income / qualification review scripts, risk disclosure scripts, and supplementary information scripts.

[0072] In some embodiments, the computer device can generate a set of questions corresponding to the target user based on the matching results and the set of dialogue scripts. For example, if the target user categories are micro and small enterprise users and high-amount applicants, and the preset matching rules in the matching results can be to add questions about asset verification, business expansion plans, and / or the use of funds, then questions such as "Please provide business operating cash flow statements" and "Please explain the business expansion plan and the use of funds in detail" can be added to the set of dialogue scripts to generate the corresponding set of questions.

[0073] Step S104: Obtain the corresponding data to be reviewed based on the set of questions, and perform semantic recognition processing on the data to be reviewed to obtain candidate data.

[0074] In this embodiment of the application, the data to be reviewed can be voice / text data representing a set of questions answered by a target user, which is obtained by a computer device.

[0075] In one embodiment, a computer device can use speech recognition technology (such as a speech recognition model based on deep learning) to perform data preprocessing such as speech signal framing / windowing on the speech data in the data to be reviewed, extract speech features such as Mel-Frequency Cepstral Coefficients (MFCC) for semantic recognition processing, and convert the speech features into text features to obtain candidate data.

[0076] In some embodiments, such as Figure 2 As shown, the data to be reviewed includes voice data and first text data; step S104 includes:

[0077] Step S1041: The speech data is semantically recognized and converted into second text data, and the first text data and the second text data are concatenated and / or weighted and fused to obtain candidate data;

[0078] Step S1042: Perform semantic recognition processing on the candidate data to obtain the alternative data.

[0079] In this embodiment, the first text data refers to text data directly acquired by the computer device; for example, the first text data may be text data entered by the target user when filling out a business application form. The second text data refers to text data converted by the computer device based on voice data; for example, the second text data may be text data converted by the computer device from voice data of an acquired set of answer questions.

[0080] In one embodiment, a computer device may determine whether a target user's voice questions and answers are consistent with their text questions and answers based on first text data and second text data.

[0081] Step S105: Convert the non-standardized question-and-answer data in the candidate data into standardized question-and-answer data to obtain the target data.

[0082] In this application embodiment, non-standardized question-and-answer data indicates fuzzy question-and-answer data with unclear / unclear semantic representation.

[0083] In this embodiment of the application, standardized question-and-answer data indicates question-and-answer data with clear / explicit semantic representation.

[0084] In some embodiments, a computer device may construct a semantic understanding model based on a Transformer architecture; capture semantic features of text data corresponding to non-standard question-and-answer data using a multi-layer Transformer encoder in the semantic understanding model; and use a fully connected layer to map the semantic features into numerical representations or category labels based on preset mapping rules and output them.

[0085] For example, if the non-standard question-and-answer data is "income is okay", the computer device, based on mapping rules and combined with the contextual data of the non-standard question-and-answer data and the industry background of the target user, converts the non-standard question-and-answer data into a specific value / category corresponding to the income level. For example, if the target user's industry background is the financial industry, and the contextual data includes the target user's application amount and specific occupational information, then "income is moderate" can be converted into a specific value "20,000 to 50,000" corresponding to the income level.

[0086] Step S106: Obtain the audit result of the target user based on the target data; the audit result includes the probability of default.

[0087] In some embodiments, the computer device inputs target data into a prediction model to predict the probability of default. If the probability of default is greater than a preset threshold for non-compliance, the probability of default can indicate that the target user is more likely to default, and the review result can be marked as "not passed". If the probability of default is less than or equal to the threshold for non-compliance, the probability of default can indicate that the target user is less likely to default, and the review result can be determined by combining other influencing factors (such as historical overdue records, manual review opinions, etc.).

[0088] For example, the threshold for failure to meet the standard can be 0.2, 0.25, 0.3, etc.

[0089] The aforementioned data review method automates processes such as data collection, rule matching, and semantic recognition, replacing manual item-by-item review. This reduces processing time per target user, lowers both time and labor costs, and supports parallel processing of massive numbers of users. Furthermore, the preset rules can be flexibly expanded to quickly respond to business changes, thus covering multiple scenarios (such as loan, insurance, or leasing services). Moreover, by integrating basic user data (occupation, age, and credit reports, etc.) with data to be reviewed, multi-dimensional data fusion is achieved, enabling the construction of a comprehensive risk profile of the target user. Standardized quantification of non-standardized Q&A data reduces subjective judgment errors and achieves accurate review results.

[0090] In some embodiments, the method further includes:

[0091] Based on the audit results, a validity index is determined for each question in the set of scripts; the validity index is used to characterize the degree of influence of the question on the audit results.

[0092] The weights for each question are adjusted based on the exponentially weighted moving average algorithm and the aforementioned effectiveness index.

[0093] The set of dialogues is updated based on the adjusted weights to obtain the updated set of dialogues.

[0094] In this embodiment of the application, the effectiveness index can be determined based on factors such as the frequency of obtaining effective information from each question in the script set, the accuracy / completeness of the effective information, and the importance of the effective information.

[0095] For example, if the data to be reviewed for the first question is non-standardized question-and-answer data, and the answers are vague and difficult to provide a clear basis for judging the review results, then it can be determined that the accuracy and completeness of the effective information for the first question are relatively low, and the corresponding effectiveness index is low.

[0096] For example, if the target user's income level and income source are critical to the audit results, then the importance of the income level and income source issues can be determined to be high, and the corresponding effectiveness indicators are high.

[0097] In this embodiment of the application, the Exponential Weighted Moving Average (EWMA) algorithm can assign weights to each question in the speech set through an exponential decay mechanism.

[0098] For example, if the frequency of obtaining effective information from the k-th question in the dialogue set decreases, the computer device can quickly reduce the weight assigned to the k-th question according to the EWMA algorithm.

[0099] In this embodiment, the effectiveness index (contribution) of each question to the risk review is derived backward from the review results (such as default probability, pass rate, etc.), which can quantify the value of the questions. The weights of each question are adjusted based on the effectiveness index, thereby automatically identifying and eliminating inefficient questions. The priority of questions can be dynamically updated based on the weight assignments, for example, prioritizing high-weight questions. Furthermore, eliminating inefficient questions can shorten user response time, reduce user fatigue, and improve user experience.

[0100] In one embodiment, the method further includes:

[0101] During the review process of obtaining the review result based on the data to be reviewed, the current review status is determined; the current review status is used to characterize the current stage of the review.

[0102] Based on the current review status, the output order of each question in the script set is optimized to obtain an optimized script set.

[0103] In this embodiment, the current review status may include, but is not limited to, the questions already answered by the target user and the corresponding data pending review, the current review progress, and the feature values ​​corresponding to each customer feature / text feature. The current review status can be represented as a vector.

[0104] In some embodiments, the computer device can optimize the output order of questions in the dialogue set based on the current review status and in conjunction with a reward function mechanism. For example, the next question may include three candidate questions. The first candidate question corresponds to information of high importance and, based on the review status, can be determined to have a close connection with the previous question. Since the first candidate question and the previous question have a progressive relationship, the first candidate question can be given a positive reward based on the reward function mechanism. The second candidate question corresponds to information of high importance but has no inherent logical connection with the previous question, and the third candidate question corresponds to information of low importance. Therefore, the second and third candidate questions can be given negative rewards based on the reward function mechanism. Thus, the first candidate question can be determined as the next question, and so on, until the output order of each question is optimized to obtain the optimized dialogue set.

[0105] For example, if the previous question involved "recent unemployment", the next question can skip candidate questions related to the stability of the current job; or, if the previous question involved "being a freelancer", the next question can prioritize questions related to recent job performance.

[0106] In this embodiment of the application, by dynamically identifying the current review status in real time, the next question that is closely related to or of higher importance to the previous question can be selected based on the current review status. The priority of questions can be dynamically adjusted according to the stage goals, thereby improving the efficiency of obtaining key information, thus enhancing the review capability (accuracy of output questions) and the accuracy of review results, reducing invalid questions and answers to improve the user experience.

[0107] In one embodiment, the step of performing semantic recognition processing on the candidate data to obtain the alternative data includes:

[0108] The candidate data is identified based on a named entity recognition model to extract target entities;

[0109] If the target entity includes a numerical entity, determine the numerical range corresponding to the numerical entity and / or convert the currency unit of the numerical entity to obtain the alternative data.

[0110] In one embodiment, a computer device can input candidate data into a named recognition model, capture contextual information of the candidate data through a multi-layer Transformer encoder, and use a conditional random field (CRF) in the output layer to predict and label target entities and categories; for example, target entities can be "income", "liability" and "loan term", etc.

[0111] In this embodiment, the numerical entity refers to the target entity involving a specific numerical value. For example, target entities such as "income," "loan amount," and "liabilities" all involve specific numerical values.

[0112] In one embodiment, the candidate data is “last year’s bonus plus basic salary totaled about 250,000, but the monthly mortgage payment is 12,000”. The named entity recognition model can identify and label the income entity (250,000 annual income, including bonus) and the liability entity (12,000 / month, mortgage payment).

[0113] For example, if the numerical entity is "monthly income of US$3,000", the computer device can convert the currency unit "US$" of the numerical entity into "RMB" according to the current exchange rate and conversion rules, and determine the corresponding value "21,000 yuan".

[0114] For example, if the numerical entity is "initial monthly income", the computer device can determine the corresponding numerical range as "5000~10000 yuan".

[0115] In this embodiment of the application, when the target entity includes a numerical entity, real-time conversion between multiple currency units can be realized, and fuzzy expressions can be converted into precise ranges and labeled with confidence levels to realize multi-dimensional complex numerical relationship analysis, identify potential default risks, and provide data support for subsequent review.

[0116] In one embodiment, the method further includes:

[0117] The first text data and the second text data are compared to obtain the anomaly prediction result;

[0118] If the anomaly prediction result indicates that the data to be reviewed is abnormal, an anomaly warning message is output; the anomaly warning message includes abnormal data items and warning prompts.

[0119] In some embodiments, the computer device can train a data association model based on training sample data. The data association model can learn and identify the inherent relationships between various data items; for example, the reasonable ratio range between income and liabilities, the ratio between assets and loan amounts, etc. The computer device compares the first text data and the second text data acquired in real time using the data association model to determine whether there are any anomalies. If anomalies are found, for example, if the difference between the income data in the first text data and the income data in the second text data exceeds a difference threshold, or if the ratio between the income data in the first text data and the liability data in the second text data is not within a reasonable range, an early warning mechanism can be triggered, and an anomaly warning message can be output.

[0120] In this embodiment, the abnormal data item indicates the data where the abnormal situation occurs. For example, if the abnormal situation is that the difference between the income data contained in the first text data and the income data contained in the second text data exceeds a difference threshold, then the abnormal data item is the income data.

[0121] In some embodiments, the way to output abnormal warning information may include, but is not limited to, at least one of system message push, voice output, text output, graphical list output, email output, and SMS output.

[0122] In some embodiments, the output of abnormal warning information includes at least one of the following:

[0123] The first type: using a computer-based voice device to output abnormal warning information;

[0124] The second method is to output an abnormal warning message on the computer device's screen.

[0125] The third type: Early warning information for abnormal motor vibration output based on computer equipment;

[0126] The fourth method is to send the abnormality warning information to a third-party platform; the abnormality warning information is used for display by the third-party platform.

[0127] In this embodiment, by comparing the target user's declared data in real time, non-obvious risks such as data tampering and logical contradictions can be identified in advance; by marking and outputting abnormal data items, contradiction risk points can be accurately detected, making it easier for business personnel to provide timely feedback, realizing a transparent feedback mechanism, reducing customer objection rate and improving customer trust.

[0128] In one embodiment, the step of performing semantic recognition processing on the candidate data to obtain the alternative data includes:

[0129] The candidate data acquired each time is subjected to semantic recognition processing to obtain semantic recognition results;

[0130] Based on the semantic recognition results, intent classification and labeling are performed to obtain candidate data for labeled intent classification;

[0131] According to the intent classification, follow-up questions are matched from the preset set of follow-up questions, the follow-up questions are updated to the question set, and the candidate data is updated based on the data to be reviewed obtained from the follow-up questions to obtain the alternative data.

[0132] In some embodiments, the computer device trains an intent classification model based on training sample data labeled with intent classification. For example, intent classification may include, but is not limited to, querying loan amounts, displaying income status, and explaining repayment risks. During the training process, the model can be iteratively optimized using the cross-entropy loss function, and the model parameters of the intent classification model can be updated through the backpropagation algorithm to improve the accuracy of intent classification.

[0133] In some embodiments, the computer device uses an intent classification model to perform semantic recognition processing on candidate data, obtains semantic recognition results, and obtains intent classification of the candidate data based on the semantic recognition results; matches corresponding follow-up questions from the follow-up questioning script according to the intent classification, for example, when the intent classification indicates that the target user's income is unstable, the matched follow-up questions can be "the range of income fluctuations in the last six months" and / or "what are the main reasons for income instability"; or, when the intent classification indicates that the target user's work flow is unclear, the matched follow-up questions can be follow-up questions in another form of asking about work flow; updates the matched follow-up questions to the question set, obtains the data to be reviewed corresponding to the follow-up questions, and updates the candidate data to obtain alternative data.

[0134] In this embodiment, potential risk points of target users are accurately identified through intent classification. Further multi-level and multi-angle follow-up questions are used to deeply uncover hidden risk clues from target users. Compared to the traditional single-round question-and-answer model, this can further improve the accuracy of data review. Furthermore, the question-and-answer strategy is dynamically adjusted based on the candidate data of the target user's real-time responses, further refining effective communication, obtaining more detailed data, improving the completeness of key information, and dynamically adapting to the responses of different target users.

[0135] This application also provides an application scenario in which the above-described data auditing method is applied. Specifically, the data auditing method is applied in this scenario as follows:

[0136] In the dynamic dialogue template generation module, the computer equipment uses a clustering algorithm to classify the basic data of the target user and determine the category of the target user. The basic data may include the basic information of the target user, product requirements, and historical review data. The category of the target user is matched with preset rules to obtain the matching result. Based on the matching result and the dialogue set, a personalized set of questions corresponding to the category of the target user is generated.

[0137] In the intelligent question-answering module, the computer device can convert the voice data in the data to be reviewed (i.e., the target user's response data) corresponding to the acquired question set into second text data, and then concatenate and fuse the second text data with the first text data in the data to be reviewed to obtain candidate data; input the candidate data into a multimodal model for semantic recognition processing, extract target entities, and parse the corresponding numerical range of the target entities; and / or, compare the first text data and the second text data, and output an abnormal warning message when the semantic results represented by the first text data and the second text data are inconsistent; and / or, input the candidate data into an intent classification model to obtain the intent classification result corresponding to the candidate data; match follow-up questions based on the intent classification result, update the question set with successfully matched follow-up questions to obtain the data to be reviewed corresponding to the follow-up questions, and update the candidate data to obtain alternative data.

[0138] In the multi-model collaborative decision-making module, the computer equipment can convert non-standardized question-and-answer data from the candidate data into standardized question-and-answer data to obtain the target data; it can use a rule definition language (such as Drools rule language) to preset and set review rules and store them in the rule base; the review rules include a condition part and an action part; the target users are screened according to the review rules, and if the target data does not meet the review rules, the screening result can be determined as a failure; the target data corresponding to the screened target users can be input into the prediction model to predict the default probability of the target users and determine the final review result.

[0139] In the self-optimization feedback module, the computer equipment can determine the validity index of each question in the script set based on the review results; adjust the weight assignment of each question based on the exponentially weighted moving average algorithm and the validity index; update the script set according to the adjusted weight assignment to obtain the updated script set; and determine the current review status during the review process based on the review results obtained from the data to be reviewed; optimize the output order of each question in the script set based on the current review status to obtain the optimized script set.

[0140] The aforementioned data review method automates processes such as data collection, rule matching, and semantic recognition, replacing manual item-by-item review. This reduces processing time per target user, lowers both time and labor costs, and supports parallel processing of massive numbers of users. Furthermore, the preset rules can be flexibly expanded to quickly respond to business changes, thus covering multiple scenarios (such as loan, insurance, or leasing services). Moreover, by integrating basic user data (occupation, age, and credit reports, etc.) with data to be reviewed, multi-dimensional data fusion is achieved, enabling the construction of a comprehensive risk profile of the target user. Standardized quantification of non-standardized Q&A data reduces subjective judgment errors and achieves accurate review results.

[0141] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0142] Based on the same inventive concept, this application also provides a data auditing apparatus for implementing the data auditing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more data auditing apparatus embodiments provided below can be found in the limitations of the data auditing method described above, and will not be repeated here.

[0143] In one exemplary embodiment, such as Figure 3 As shown, a data auditing device is provided, comprising:

[0144] Module 10 is used to acquire basic data of the target user;

[0145] Matching module 20 is used to determine the category of the target user based on the basic data, and match the category of the target user with preset rules to obtain a matching result;

[0146] The generation module 30 is used to generate a corresponding set of questions based on the matching results and the set of dialogues;

[0147] The identification module 40 is used to obtain the corresponding data to be reviewed based on the question set, and to perform semantic recognition processing on the data to be reviewed to obtain candidate data;

[0148] The conversion module 50 is used to convert the non-standardized question and answer data in the candidate data into standardized question and answer data to obtain the target data;

[0149] The review module 60 is used to obtain the review result of the target user based on the target data; the review result includes the probability of default.

[0150] In one embodiment, the apparatus further includes:

[0151] The first determining module is used to determine the validity index of each question in the script set based on the audit results; the validity index is used to characterize the degree of influence of the question on the audit results.

[0152] The first adjustment module is used to adjust the weight assignment for each problem based on the exponentially weighted moving average algorithm and in combination with the effectiveness index.

[0153] The update module is used to update the script set according to the adjusted weight assignment, so as to obtain the updated script set.

[0154] In one embodiment, the apparatus further includes:

[0155] The second determining module is used to determine the current audit status during the audit process of obtaining the audit result based on the data to be audited; the current audit status is used to characterize the current stage of the audit.

[0156] The second adjustment module is used to optimize the output order of each question in the script set based on the current review status, so as to obtain an optimized script set.

[0157] In one embodiment, the data to be reviewed includes voice data and first text data; the recognition module 40 includes:

[0158] The conversion unit is used to perform semantic recognition on the speech data and convert it into second text data, and to concatenate and / or weightedly fuse the first text data and the second text data to obtain candidate data;

[0159] The identification unit is used to perform semantic recognition processing on the candidate data to obtain the alternative data.

[0160] In one embodiment, the identification unit is configured to perform the following steps:

[0161] The candidate data is identified based on a named entity recognition model to extract target entities;

[0162] If the target entity includes a numerical entity, determine the numerical range corresponding to the numerical entity and / or convert the currency unit of the numerical entity to obtain the alternative data.

[0163] In one embodiment, the apparatus further includes:

[0164] The comparison module is used to compare the first text data and the second text data to obtain an anomaly prediction result;

[0165] The output module is used to output an anomaly warning message when the anomaly prediction result indicates that the data to be reviewed is abnormal; the anomaly warning message includes an abnormal data item and a warning prompt.

[0166] In one embodiment, the identification unit is configured to perform the following steps:

[0167] The candidate data acquired each time is subjected to semantic recognition processing to obtain semantic recognition results;

[0168] Based on the semantic recognition results, intent classification and labeling are performed to obtain candidate data for labeled intent classification;

[0169] According to the intent classification, follow-up questions are matched from the preset set of follow-up questions, the follow-up questions are updated to the question set, and the candidate data is updated based on the data to be reviewed obtained from the follow-up questions to obtain the alternative data.

[0170] Each module in the aforementioned data verification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0171] In one embodiment, a computer device is provided, which can be any terminal, and its internal structure diagram can be as follows: Figure 4As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a sample analysis method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0172] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0173] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0174] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0175] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0176] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0177] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data auditing method, characterized in that, The method includes: Obtain basic data of the target user; The target user's category is determined based on the basic data, and the target user's category is matched with preset rules to obtain a matching result; Based on the matching results and the set of dialogue options, a corresponding set of questions is generated; Based on the set of questions, obtain the corresponding data to be reviewed, and perform semantic recognition processing on the data to be reviewed to obtain candidate data; The non-standardized question-and-answer data in the candidate data is transformed into standardized question-and-answer data to obtain the target data; The review results for the target user are obtained based on the target data; the review results include the probability of default.

2. The method according to claim 1, characterized in that, The method further includes: Based on the audit results, a validity index is determined for each question in the set of dialogue options; the validity index is used to characterize the degree of influence of the question on the audit results. The weights for each question are adjusted based on the exponentially weighted moving average algorithm and the aforementioned effectiveness index. The set of dialogues is updated based on the adjusted weights to obtain the updated set of dialogues.

3. The method according to claim 1, characterized in that, The method further includes: During the review process of obtaining the review result based on the data to be reviewed, the current review status is determined; the current review status is used to characterize the current stage of the review. Based on the current review status, the output order of each question in the script set is optimized to obtain an optimized script set.

4. The method according to claim 1, characterized in that, The data to be reviewed includes voice data and first text data; the process of obtaining the corresponding data to be reviewed based on the question set, and performing semantic recognition processing on the data to be reviewed to obtain candidate data, includes: The speech data is semantically recognized and converted into second text data, and the first text data and the second text data are concatenated and / or weighted and fused to obtain candidate data; The candidate data is obtained by performing semantic recognition processing on the candidate data.

5. The method according to claim 4, characterized in that, The step of performing semantic recognition processing on the candidate data to obtain the alternative data includes: The candidate data is identified based on a named entity recognition model to extract target entities; If the target entity includes a numerical entity, determine the numerical range corresponding to the numerical entity and / or convert the currency unit of the numerical entity to obtain the alternative data.

6. The method according to claim 4, characterized in that, The method further includes: The first text data and the second text data are compared to obtain the anomaly prediction result; If the anomaly prediction result indicates that the data to be reviewed is abnormal, an anomaly warning message is output; the anomaly warning message includes abnormal data items and warning prompts.

7. The method according to claim 4, characterized in that, The step of performing semantic recognition processing on the candidate data to obtain the alternative data includes: The candidate data acquired each time is subjected to semantic recognition processing to obtain semantic recognition results; Based on the semantic recognition results, intent classification and labeling are performed to obtain candidate data for labeled intent classification; According to the intent classification, follow-up questions are matched from the preset set of follow-up questions, the follow-up questions are updated to the question set, and the candidate data is updated based on the data to be reviewed obtained from the follow-up questions to obtain the alternative data.

8. A data verification device, characterized in that, The device includes: The acquisition module is used to acquire basic data of the target user; The matching module is used to determine the category of the target user based on the basic data, and match the category of the target user with preset rules to obtain a matching result; The generation module is used to generate a corresponding set of questions based on the matching results and the set of dialogues; The identification module is used to obtain the corresponding data to be reviewed based on the set of questions, and to perform semantic recognition processing on the data to be reviewed to obtain candidate data; The conversion module is used to convert non-standardized question-and-answer data in the candidate data into standardized question-and-answer data to obtain the target data; The review module is used to obtain the review result of the target user based on the target data; the review result includes the probability of default.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.