A data verification method, device, equipment, storage medium and product
Through the combination of internal and external joint mechanisms and preset classification models, the problem of fallacy identification in the verification of business data of financial institutions is solved, and accurate verification and risk identification of data reported by financial institutions is achieved.
Patent Information
- Application Number
- CN202510200032.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In financial audit scenarios, business data reported by financial institutions may have human operational errors or tampering, resulting in difficulty in data verification.
Through internal and external joint mechanisms, the business data to be verified and the set of internal and external prediction models of the target organization are obtained. The internal prediction model is based on the internal modeling of historical business data of the target organization, and the external prediction model is based on the external modeling of historical business data of the target organization and similar organizations. These models are used to predict the verification data, and the data verification results are determined through the preset classification model.
Accurate verification of business data reported by financial institutions can be realized, and possible fallacies in the data can be identified, thereby assisting regulatory agencies to discover related risks in a timely manner.
Smart Images

Figure CN119691528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data verification method, apparatus, device, storage medium and product. Background Art
[0002] In the financial audit scenario, there is a need to verify the business data reported by financial institutions. On the one hand, the business data reported by financial institutions may have errors such as misalignment and incorrect filling due to human operation errors; on the other hand, the reported business data may be tampered with by humans. Therefore, how to accurately verify the business data reported by financial institutions has become an urgent technical problem to be solved. Summary of the Invention
[0003] The present invention provides a data verification method, apparatus, device, storage medium and product to verify the business data reported by an institution through an internal and external joint mechanism and accurately identify possible errors in the business data.
[0004] According to one aspect of the present invention, a data verification method is provided. The method includes:
[0005] Obtaining the business data to be verified at a target time point of a target institution, as well as a target internal prediction model set and a target external prediction model set; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institutions as the target institution;
[0006] Based on the business data to be verified and the verified data of each same type of institution, respectively predicting the target time point according to the target internal prediction model set and the target external prediction model set to obtain corresponding internal modeling prediction results and external modeling prediction results;
[0007] Determining the data verification results corresponding to the business data to be verified, the internal modeling prediction results and the external modeling prediction results according to a preset classification model.
[0008] According to another aspect of the present invention, a data verification apparatus is provided. The apparatus includes:
[0009] An obtaining module, configured to obtain the business data to be verified at a target time point of a target institution, as well as a target internal prediction model set and a target external prediction model set; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institutions as the target institution;
[0010] A prediction module, configured to predict a target time point respectively according to a set of target internal prediction models and a set of target external prediction models based on the business data to be verified and the verified data of each similar institution, and respectively obtain corresponding internal modeling prediction results and external modeling prediction results;
[0011] A verification module, configured to determine data verification results corresponding to the business data to be verified, the internal modeling prediction results and the external modeling prediction results according to a preset classification model.
[0012] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data verification method according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the data verification method according to any embodiment of the present invention when executed.
[0017] According to another aspect of the present invention, there is provided a computer program product including a computer program which implements the data verification method according to any embodiment of the present invention when executed by a processor.
[0018] The technical solution of the embodiment of the present invention is to obtain the business data to be verified at the target time point of the target institution, as well as the target internal prediction model set and the target external prediction model set; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institutions as the target institution; based on the business data to be verified and the verification-passed data of each same type of institution, respectively predict the target time point according to the target internal prediction model set and the target external prediction model set to obtain the corresponding internal modeling prediction result and external modeling prediction result; determine the data verification result corresponding to the business data to be verified, the internal modeling prediction result and the external modeling prediction result according to the preset classification model. Through the internal modeling and external modeling methods, this technical solution can fully model the complex business mechanism of the institution, learn the operating characteristics of the institution itself, and then through the fusion modeling process of the preset classification model, it can accurately identify the possible fallacies in the business data, so as to assist the regulatory agency to discover relevant risks in a timely manner.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a data verification method provided in Embodiment 1 of the present invention;
[0022] Figure 2 is a flowchart of a data verification method provided in Embodiment 2 of the present invention;
[0023] Figure 3 is a schematic structural diagram of a data verification device provided in Embodiment 3 of the present invention;
[0024] Figure 4 is a schematic structural diagram of an electronic device for implementing the data verification method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] Embodiment 1
[0028] Figure 1 FIG. 10 is a flowchart of a data verification method provided in Embodiment 1 of the present invention. This embodiment is applicable to verifying the business data reported by an institution to identify spurious data. This method can be executed by a data verification device, which can be implemented in the form of hardware and / or software, and the data verification device can be configured in an electronic device. As Figure 1 shown, a data verification method provided in Embodiment 1 specifically includes the following steps:
[0029] S110. Obtain the business data to be verified at the target time point of the target institution, as well as the target internal prediction model set and the target external prediction model set; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institution to which the target institution belongs.
[0030] Among them, the target institution may refer to the institution to be verified for data, such as a certain supervised banking financial institution. The target time point may refer to the time point corresponding to the data reported by the target institution in the current period.
[0031] The business data to be verified can be understood as the business data reported by the target institution in the current period, which may include, but is not limited to, regulatory data standardization data (Examination and Analysis System Technology, EAST), off-site supervision reports (such as 1104 reports), etc. Further, the business data to be verified may contain several business fields (i.e., business dimensions), and each business field corresponds to a data item (i.e., field value). For example, the 1104 report may include business fields such as investment in interbank certificates of deposit and payable dividends.
[0032] The set of target internal prediction models contains several target internal prediction models. Each target internal prediction model is associated with a corresponding first business field and can be obtained by internal modeling based on the first historical business data of the target institution. Among them, the first historical business data may include the historical business data reported by the target institution for each historical time point.
[0033] The set of target external prediction models contains several target external prediction models. Each target external prediction model is associated with a corresponding second business field and can be obtained by external modeling based on the first historical business data of the target institution and the second historical business data of the institutions of the same type as the target institution. Among them, institutions of the same type refer to institutions belonging to the same category as the target institution; the second historical business data may include the historical business data reported by each institution of the same type for each historical time point.
[0034] In the embodiment of the present invention, during the supervision of the target institution, the business data to be verified reported by the target institution for the target time point can be obtained, and at the same time, the set of target internal prediction models and the set of target external prediction models obtained by internal modeling and external modeling respectively in advance can be obtained, so that the business data at the target time point can be predicted by using the set of target internal prediction models and the set of target external prediction models respectively.
[0035] It should be understood that by using internal modeling, those fallacy data that only modify a small amount of data and do not achieve the self-consistency of the internal operation logic of the institution can be found; by using external modeling, the possible fallacies in the data reported by the target institution can be found through the business data and operation modes of the institutions of the same type as the target institution.
[0036] S120. Based on the business data to be verified and the verified data of each institution of the same type, predict the target time point according to the set of target internal prediction models and the set of target external prediction models respectively, and obtain the corresponding internal modeling prediction result and external modeling prediction result.
[0037] Among them, the verified data can be understood as the business data that has passed verification by similar institutions for the target time point. That is, the verified data can be regarded as the real business data of similar institutions, which can be used as an important reference for discovering incorrect data in the data reported by the target institution.
[0038] The internal modeling prediction results can include the prediction results of each target internal prediction model in the target internal prediction model set for the target time point respectively. Similarly, the external modeling prediction results can include the prediction results of each target external prediction model in the target external prediction model set for the target time point respectively.
[0039] In the embodiment of the present invention, based on the business data to be verified of the target institution, each target internal prediction model in the target internal prediction model set can be used to predict the business data at the target time point respectively, and the first target prediction results output by the models are obtained respectively, and each first target prediction result is used as the internal modeling prediction result corresponding to the target institution at the target time point. Similarly, based on the business data to be verified of the target institution and the verified data of the similar institutions to which the target institution belongs, each target external prediction model in the target external prediction model set can be used to predict the business data at the target time point respectively, and the second target prediction results output by the models are obtained respectively, and each second target prediction result is used as the external modeling prediction result corresponding to the target institution at the target time point.
[0040] Further, in order to eliminate the magnitude difference between the data items of different business fields in the business data to be verified and the verified data, before using the business data to be verified of the target institution and the verified data of the similar institutions for prediction, a pre-configured standardization processing method can be called to standardize the business data to be verified and the verified data respectively.
[0041] S130. Determine the data verification results corresponding to the business data to be verified, the internal modeling prediction results, and the external modeling prediction results according to a preset classification model.
[0042] Among them, the preset classification model can refer to a model for realizing incorrect data identification, which can include but is not limited to: BERT (Bidirectional Encoder Representations from Transformers) model, CatBoost model, XGB (eXtreme Gradient Boosting) model, etc. The data verification results can be used to characterize whether there is incorrect data in the business data reported by the target institution in this period.
[0043] In an embodiment of the present invention, the obtained business data to be verified, the internal modeling prediction result, and the external modeling prediction result can be jointly input into a pre-trained preset classification model, and the model is used for fusion modeling processing to identify the possible fallacy data in the business data to be verified, so as to obtain the corresponding data verification result; wherein, the data verification result output by the model can be represented by 0 or 1, and 0 indicates the existence of fallacy, and 1 indicates the non-existence of fallacy.
[0044] It should be understood that the classification logic of the preset classification model in this embodiment is not specifically limited. In one embodiment, the classification logic of the preset classification model can be to separately analyze and process the data items of a single business field in the business data to be verified, and as long as the data item corresponding to one business field is identified as having a fallacy, the data verification result finally output by the model is that there is a fallacy. In another embodiment, the classification logic of the preset classification model can also be to perform an overall analysis and processing on the data items of all business fields in the business data to be verified. If the actual overall situation of the data items corresponding to all business fields deviates from the internal modeling prediction result and the external modeling prediction result by more than a preset deviation threshold, the data verification result finally output by the model is that there is a fallacy.
[0045] The technical solution of the embodiment of the present invention is to obtain the business data to be verified at the target time point of the target institution, as well as the target internal prediction model set and the target external prediction model set; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institutions as the target institution; based on the business data to be verified and the verified data of each same type of institution, the target time point is predicted respectively according to the target internal prediction model set and the target external prediction model set to obtain the corresponding internal modeling prediction result and external modeling prediction result; the data verification result corresponding to the business data to be verified, the internal modeling prediction result, and the external modeling prediction result is determined according to the preset classification model. Through the internal modeling and external modeling methods, this technical solution can fully model the complex business mechanism of the institution, learn the operating characteristics of the institution itself, and then through the fusion modeling processing of the preset classification model, it can accurately identify the possible fallacies in the business data, so as to assist the regulatory agency to discover relevant risks in a timely manner.
[0046] Further, on the basis of the above-mentioned embodiment of the invention, a data verification method provided in this embodiment further includes:
[0047] Obtain the historical business data of each institution corresponding to each historical time point, and perform standardization processing on the historical business data to obtain historical business samples;
[0048] Based on the business fields in the preset supervision data table, use historical business samples to cluster each institution to determine the target category corresponding to each institution.
[0049] Among them, the historical business data can refer to the business data reported by each institution in history. Standardization processing can be used to eliminate the magnitude differences between data items of different business fields. The standardization processing methods can include but are not limited to range standardization method, Z-score standardization method, etc.
[0050] The historical business sample represents the business data after standardization processing corresponding to a historical time point. The historical business sample can contain different business dimension information of a certain institution at a certain historical time point. For example, the historical business sample in a certain XX year and XX month can contain the following business data: {(business field 1, data item); (business field 2, data item); (business field 3, data item)}.
[0051] The preset supervision data table can refer to the supervision business reports reported by each institution, which can include but are not limited to 1104 reports, EAST reports, the central bank's large centralized reports, etc.; the preset supervision data table contains different business fields, and each business field represents a business type of a different dimension. And the business fields covered by each institution may vary, that is, not every institution has corresponding business for each business field. If an institution has no relevant business, the data item of the corresponding business field is empty.
[0052] The target category can be understood as the category corresponding to each institution after clustering the institutions. And each institution corresponds to a unique target category. At the same time, a group of institutions belonging to the same target category have the same business fields for the reported business data.
[0053] In the embodiment of the present invention, all institutions can be clustered based on the historical business samples of each supervised institution, so as to obtain the target category corresponding to each institution. Specifically, the historical business data reported by each supervised institution for each historical time point can be obtained from a data storage location such as a local or cloud server, and then the data items corresponding to each business field in the above historical business data are standardized, so as to obtain the historical business samples corresponding to each institution at different historical time points; then, based on the business fields in the preset supervision data table and combined with the business type of each institution itself, all institutions can be divided into several categories, so as to determine the target category corresponding to each institution.
[0054] Further, on the basis of the above-mentioned embodiment of the invention, using historical business samples to cluster each institution based on the business fields in the preset supervision data table to determine the target category corresponding to each institution includes:
[0055] Step 1: Coarsely classify each institution according to the presence or absence of the corresponding business fields of each institution;
[0056] Step 2: Determine whether the number of institutions included in each first category after coarse classification is greater than the preset institution number threshold;
[0057] Step 3: If so, call the preset clustering algorithm to cluster the historical business samples of the institutions within the first category, count the number of samples of the historical business samples of each institution belonging to each second category after clustering, and use the second category with the largest corresponding number of samples as the target category of the corresponding institution;
[0058] Step 4: If not, use the first category as the target category of the corresponding institution.
[0059] Among them, the coarse classification can be understood as a rough classification of all institutions based on the presence or absence of the corresponding business fields of each institution. The category corresponding to each institution after coarse classification is denoted as the first category; it can be understood that the number of the corresponding first categories after coarse classification is usually not too large.
[0060] The preset clustering algorithm can be a clustering algorithm pre-configured for re-finely classifying the institutions after coarse classification. The preset clustering algorithm can include but is not limited to: hierarchical clustering algorithm, K-means clustering algorithm, etc.; at the same time, several categories obtained after being processed by the preset clustering algorithm can be denoted as the second category.
[0061] In the embodiment of the present invention, the process of coarsely classifying and clustering the institutions is specifically as follows: First, coarsely classify each institution according to the presence or absence of the corresponding business fields of each institution, that is, divide a group of institutions with the same business fields into one category, and obtain the first category corresponding to each institution after coarse classification; then, count the number of institutions in each first category and compare it with the preset institution number threshold (for example, 2). If the number of institutions corresponding to a certain category of institutions is greater than the preset institution number threshold, call the preset clustering algorithm to cluster the institutions in this category to obtain the target category corresponding to each institution in this category; on the contrary, if the number of institutions corresponding to a certain category of institutions is less than or equal to the preset institution number threshold, there is no need to cluster the institutions in this category, and the first category corresponding to each institution in this category can be directly used as the corresponding target category.
[0062] Among them, the specific process of calling a preset clustering algorithm to cluster this type of institution to obtain the target category corresponding to each institution in this type of institution may include: for the case where the number of institutions corresponding to a certain type of institution is greater than the preset institution number threshold, a preset clustering algorithm such as hierarchical clustering can be called to cluster the historical business samples corresponding to this type of institution at each historical time point. After the clustering is completed, several second categories can be obtained; then, the number of samples of the historical business samples of each institution belonging to each second category can be counted, that is, after all the historical business samples of a certain institution are divided into each second category, the number of historical business samples included in each second category; finally, the second category corresponding to the largest number of samples can be used as the target category of the corresponding institution.
[0063] Further, on the basis of the above-mentioned invention embodiments, the process of obtaining the target internal prediction model set based on internal modeling includes:
[0064] Step 1, obtain the target category corresponding to the target institution, and extract the data items of each business field corresponding to the target category from the first historical business data;
[0065] Step 2, divide the data items corresponding to each business field into training data items and test data items;
[0066] Step 3, sequentially select one of each business field as the current field, use the training data items corresponding to the current field as the model output, and use the training data items corresponding to other business fields as the model input to construct the internal prediction model corresponding to the current field;
[0067] Step 4, respectively evaluate each internal prediction model using the test data items corresponding to each business field to obtain the first prediction accuracy rate of each internal prediction model;
[0068] Step 5, use the first preset number of internal prediction models with the highest first prediction accuracy rate as the target internal prediction models in the target internal prediction model set.
[0069] Among them, the internal prediction model may refer to a model constructed based on the first historical business data of the target institution for predicting the data items corresponding to a certain business field. The internal prediction model is associated with the corresponding business field, that is, for each business field included in the target category corresponding to the target institution, an internal prediction model is correspondingly created.
[0070] In the embodiments of the present invention, the internal modeling process of the target institution specifically includes:
[0071] (1) The target categories obtained by pre-clustering the target institution can be acquired, and the data items corresponding to each business field can be extracted from the obtained first historical business data based on the target categories. Among them, each business field corresponds to data items at multiple historical time points. For example, business field 1 corresponds to a data item in January 2024 and also corresponds to a data item in February 2024. Or, it can also be understood that each historical time point corresponds to data items of multiple business fields, that is, data items of multiple data dimensions.
[0072] (2) The data items corresponding to each business field are divided into training data items and test data items according to a preset data item splitting ratio. Among them, the training data items are used to train and construct an internal prediction model, and the test data items are used to evaluate the prediction performance of the internal prediction model. The preset data item splitting ratio can be set accordingly according to actual needs. For example, it can be 7:3, 8:2, etc. This embodiment does not make specific limitations on this.
[0073] (3) For each business field corresponding to the target institution, one of the business fields is sequentially selected as the current field, the training data items corresponding to the current field are used as the target output for constructing the internal prediction model to be constructed, and the training data items corresponding to the remaining business fields other than the current field are used as the model input, and an internal prediction model corresponding to the current field is constructed and trained. After this step is executed, an internal prediction model corresponding to each business field of the target institution will be established. Among them, the internal prediction model can adopt a deep neural network, such as a transformer model, or can also adopt traditional machine learning algorithms, such as a linear regression algorithm, an XGB model, etc. This embodiment does not make specific limitations on this.
[0074] (4) After the internal prediction models corresponding to each business field are established, the above internal prediction models can be evaluated using the test data items corresponding to each business field, so as to obtain the first prediction accuracy corresponding to each internal prediction model.
[0075] (5) The first preset number of internal prediction models with the highest first prediction accuracy are determined as the target internal prediction models, and the above target internal prediction models are saved to the target internal prediction model set for internal prediction of the data items corresponding to the business fields when it is necessary to verify the business data reported by the target institution. It should be understood that the first preset number is less than or equal to the number of business fields corresponding to the target institution.
[0076] Further, on the basis of the above invention embodiment, the process of obtaining the target external prediction model set based on external modeling includes:
[0077] Step 1: Obtain the target category corresponding to the target institution, and extract the data items of each business field corresponding to the target category from the first historical business data and the second historical business data respectively;
[0078] Step 2: Divide the data items corresponding to each business field into training data items and test data items;
[0079] Step 3: Select one of the business fields in turn as the current field. Based on the training data items of the current field corresponding to each similar institution in the second historical business data, call the preset cluster data feature calculation formula to determine the cluster data training feature corresponding to the current field, and use the training data items corresponding to the current field in the first historical business data as the model output and the cluster data training feature as the model input to construct the external prediction model corresponding to the current field;
[0080] Step 4: Evaluate each external prediction model by using the test data items corresponding to each business field and the cluster data test features respectively to obtain the second prediction accuracy of each external prediction model;
[0081] Step 5: Take the second preset number of external prediction models with the highest second prediction accuracy as the target external prediction models in the target external prediction model set.
[0082] Among them, the cluster data training feature can be understood as the result determined based on the preset cluster data feature calculation formula by using the training data items of the corresponding business field of the institutions of the same type as the target institution for the business field of the target institution. Correspondingly, the cluster data test feature represents the result determined based on the preset cluster data feature calculation formula by using the test data items of the corresponding business field of the institutions of the same type as the target institution for the business field of the target institution.
[0083] In the embodiment of the present invention, the external modeling process of the target institution specifically includes:
[0084] (1) The target category obtained by pre-clustering the target institution can be obtained, and the data items corresponding to each business field are extracted respectively from the obtained first historical business data and second historical business data based on the target category.
[0085] (2) Divide the data items corresponding to each business field into training data items and test data items according to the preset data item division ratio. Among them, the preset data item division ratio adopted in the external modeling and internal modeling stages can be the same or different, and this embodiment does not limit this.
[0086] (3) For each business field corresponding to the target organization, one of the business fields is selected as the current field in turn, and then the test data items of the current field of the same type of organization as the target organization are used to call the preset cluster data feature calculation formula to determine the cluster data training feature corresponding to the current field, where the preset cluster data feature calculation formula is expressed as follows:
[0087]
[0088] In the formula, represents the number of institutions in the i-th category; m represents the m-th target institution; j represents the j-th business field (current field); The data item representing the business field j corresponding to the institution k of the same category i as the target institution m; represents the cluster data characteristics of the target organization m corresponding to the business field j for the i-th organization. Further, if is the training data item, then the formula is the cluster data training feature; similarly, if is the test data item, then the formula That is the cluster data test feature.
[0089] Next, the training data item corresponding to the current field in the first historical business data can be used as the target output of the external prediction model to be constructed, and the cluster data training features calculated for the current field can be used as the model input to construct and train the external prediction model corresponding to the current field. Similarly, after executing this step, a corresponding external prediction model will be established for each business field corresponding to the target organization. It should be understood that the external prediction model and the internal prediction model can adopt the same model construction method or different model construction methods, and this embodiment does not impose specific restrictions on this.
[0090] (4) Based on the method of obtaining the cluster data training features in the aforementioned steps, the corresponding cluster data test features can be determined using the test data items of each business field of the same organization, and then the test data items corresponding to each business field (for the first historical business data) and the cluster data test features can be used to perform model evaluation on the above-mentioned external prediction model, thereby obtaining the second prediction accuracy corresponding to each external prediction model.
[0091] (5) Determine the second preset number of external prediction models with the highest second prediction accuracy as the target external prediction models, and save the above target external prediction models into the target external prediction model set for use in external prediction of data items in corresponding business fields when it is necessary to verify the business data reported by the target institution. It should be understood that the second preset number is less than or equal to the number of business fields corresponding to the target institution, and the second preset number may be the same as or different from the first preset number, which is not limited in this embodiment.
[0092] It should be noted that the above-described internal modeling process and external modeling process can be executed once before verifying the business data reported by each institution each time. In this way, the finally obtained data verification result is the most accurate, but this method requires consuming more computing resources. Under normal circumstances, the operations of the institution's business do not change significantly in the short term, that is, the business fields corresponding to the institution usually do not change in the short term. Therefore, the internal modeling process and the external modeling process can be executed once every preset model training cycle or after receiving the institution's business adjustment notice, so as to avoid executing the internal modeling process and the external modeling process every time data verification is performed, and reduce the consumption of computing resources. In addition, the internal modeling process and the external modeling process can be executed in parallel to improve the efficiency of internal modeling and external modeling.
[0093] Further, on the basis of the above invention embodiments, in S120, based on the business data to be verified and the verified data of each similar institution, the target time point is predicted according to the target internal prediction model set and the target external prediction model set respectively, and the corresponding internal modeling prediction result and external modeling prediction result are obtained, including:
[0094] S1201. Obtain the verified data of each similar institution and the target category corresponding to the target institution, and extract the data items of each business field corresponding to the target category from the business data to be verified and the verified data respectively;
[0095] S1202. Extract the first business fields corresponding to the target internal prediction models from the target internal prediction model set respectively, and extract the second business fields corresponding to the target external prediction models from the target external prediction model set respectively;
[0096] S1203. Select one of the first business fields as the first current field in turn, and use the data items corresponding to the other business fields as the model inputs of the corresponding target internal prediction models to obtain the first target prediction result output by the target internal prediction models;
[0097] S1204. Select one of the second service fields in sequence as the second current field. Based on the data items of the second current field corresponding to each similar institution, call the preset cluster data feature calculation formula to determine the cluster data feature corresponding to the second current field, and use the cluster data feature as the model input of the corresponding target external prediction model to obtain the second target prediction result output by the target external prediction model;
[0098] S1205. Use each first target prediction result as the internal modeling prediction result, and use each second target prediction result as the external modeling prediction result.
[0099] Among them, the first service field represents the service field associated with the target internal prediction model. The second service field represents the service field associated with the target external prediction model.
[0100] In the embodiments of the present invention, the acquisition methods of the internal modeling prediction result and the external modeling prediction result specifically include:
[0101] (1) It is possible to obtain the target category obtained by pre-clustering of the target institution, and obtain the verified passed data reported by the historical similar institutions to which the target institution belongs from a data storage location such as a local or cloud server, and extract the data items of each service field corresponding to the target category from the to-be-verified service data and the verified passed data respectively.
[0102] (2) In the target internal prediction model set and the target external prediction model set respectively, extract the first service fields associated with each target internal prediction model, and the second service fields associated with each target external prediction model.
[0103] (3) For each first service field, select one of the service fields in sequence as the first current field, and use the data items corresponding to the other service fields except the first current field as the model input of the corresponding target internal prediction model (for the first current field) to obtain the first target prediction result output by the target internal prediction model. After this step is executed, the first target prediction results of the target internal prediction models corresponding to each first service field will be obtained.
[0104] (4) Similarly, for each second service field, select one of the service fields in sequence as the second current field, use the data items of the second current field corresponding to each similar institution, call the preset cluster data feature calculation formula to determine the cluster data feature corresponding to the second current field, and then use the cluster data feature as the model input of the corresponding target external prediction model (for the second current field) to obtain the second target prediction result output by the target external prediction model. Similarly, after this step is executed, the second target prediction results of the target external prediction models corresponding to each second service field will be obtained.
[0105] (5) The obtained first target prediction results can be jointly used as the internal modeling prediction results, and the obtained second target prediction results can be jointly used as the external modeling prediction results.
[0106] Furthermore, the obtained business data to be verified, the internal modeling prediction results, and the external modeling prediction results can be jointly input into a pre-trained preset classification model to obtain a data verification result for the business data to be verified.
[0107] Embodiment 2
[0108] Figure 2 The flowchart of a data verification method provided in Embodiment 2 of the present invention. Based on the above embodiment, this embodiment provides an implementation manner of a data verification method, which can model the complex business mechanisms of financial institutions, discover unreasonable data manifestations, and thus accurately identify fallacy data in the business data reported by financial institutions. As Figure 2 shown, a data verification method provided in Embodiment 2 of the present invention specifically includes the following steps:
[0109] S210. Obtain the historical supervision data tables reported by each institution in the past, perform standardization processing on the historical business data in the historical supervision data tables to obtain historical business samples corresponding to each historical time point.
[0110] Among them, the supervision data table contains several business fields, each business field corresponds to a different business type, and each business field corresponds to a data item.
[0111] S220. Based on the presence or absence of business fields in the historical supervision data table, use the historical business samples to cluster each institution to determine the target category corresponding to each institution.
[0112] S230. Based on the data items of each business field corresponding to the target institution in the historical supervision data table, perform internal modeling to obtain a set of target internal prediction models corresponding to the target institution.
[0113] Among them, the internal modeling process of the target institution can refer to the above embodiment and will not be elaborated here.
[0114] S240. Based on the data items of each business field corresponding to the target institution and its peer institutions in the historical supervision data table, perform external modeling to obtain a set of target external prediction models corresponding to the target institution.
[0115] Among them, the external modeling process of the target institution can refer to the above embodiment and will not be elaborated here.
[0116] S250. Obtain the business data to be verified of the target institution for the target time point.
[0117] S260. Based on the business data to be verified and the verified data of each similar institution, use the target internal prediction model set and the target external prediction model set to predict the target time point respectively, and obtain the corresponding internal modeling prediction result and external modeling prediction result respectively.
[0118] S270. Input the obtained business data to be verified, internal modeling prediction result and external modeling prediction result into the pre-trained preset classification model together to obtain the data verification result for the business data to be verified.
[0119] In the embodiment of the present invention, the preset classification model can be used to perform fusion modeling processing on the internal modeling prediction result and the external modeling prediction result, so as to identify the possible fallacy data in the business data to be verified.
[0120] The technical solution of the embodiment of the present invention can fully model the complex business mechanism of the institution through internal modeling and external modeling, and can learn the operating characteristics of the institution itself better than the past rule-based method, so as to reveal the untruthfulness of the data; and then through the fusion modeling processing of the preset classification model, it can accurately identify the possible fallacies in the business data, so as to assist the regulatory agency to discover relevant risks in time.
[0121] Embodiment III
[0122] Figure 3 It is a schematic structural diagram of a data verification device provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:
[0123] An acquisition module 31, configured to acquire the business data to be verified at the target time point of the target institution, the target internal prediction model set and the target external prediction model set; wherein, the target internal prediction model set is obtained by internal modeling based on the first historical business data of the target institution; the target external prediction model set is obtained by external modeling based on the first historical business data and the second historical business data of the similar institutions to which the target institution belongs;
[0124] A prediction module 32, configured to predict the target time point respectively according to the target internal prediction model set and the target external prediction model set based on the business data to be verified and the verified data of each similar institution, and obtain the corresponding internal modeling prediction result and external modeling prediction result respectively;
[0125] A verification module 33, configured to determine the data verification result corresponding to the business data to be verified, the internal modeling prediction result and the external modeling prediction result according to the preset classification model.
[0126] The technical solution of the embodiment of the present invention obtains the business data to be verified at the target time point of the target institution, as well as the target internal prediction model set and the target external prediction model set through an acquisition module; wherein, the target internal prediction model set is internally modeled based on the first historical business data of the target institution; the target external prediction model set is externally modeled based on the first historical business data and the second historical business data of the same type of institutions as the target institution; the prediction module predicts the target time point based on the business data to be verified and the verified data of each same type of institution, and respectively obtains the corresponding internal modeling prediction result and external modeling prediction result according to the target internal prediction model set and the target external prediction model set; the verification module determines the data verification result corresponding to the business data to be verified, the internal modeling prediction result and the external modeling prediction result according to the preset classification model. Through the internal modeling and external modeling methods, this technical solution can fully model the complex business mechanism of the institution, learn the operating characteristics of the institution itself, and then through the fusion modeling process of the preset classification model, it can accurately identify the possible fallacies in the business data, so as to assist the regulatory agency to discover relevant risks in a timely manner.
[0127] Further, on the basis of the above-mentioned embodiment of the invention, the data verification device further includes:
[0128] A historical business sample acquisition module, configured to acquire the historical business data of each institution corresponding to each historical time point, and perform standardization processing on the historical business data to obtain historical business samples;
[0129] A clustering module, configured to cluster each institution based on the business fields in the preset supervision data table by using the historical business samples to determine the target category corresponding to each institution.
[0130] Further, on the basis of the above-mentioned embodiment of the invention, the clustering module includes:
[0131] A rough classification unit, configured to perform rough classification on each institution according to the presence or absence of the business fields corresponding to each institution;
[0132] An institution quantity judgment unit, configured to judge whether the number of institutions included in each first category after rough classification is greater than a preset institution quantity threshold;
[0133] A first target category determination unit, configured to if so, call a preset clustering algorithm to cluster the historical business samples of each institution in the first category, count the number of samples of the historical business samples of each institution belonging to each second category after clustering, and use the second category with the largest corresponding sample quantity as the target category of the corresponding institution;
[0134] A second target category determination unit, configured to if not, use the first category as the target category of the corresponding institution.
[0135] Further, based on the above-described embodiments of the invention, the process of obtaining the target internal prediction model set based on internal modeling includes:
[0136] Obtain the target category corresponding to the target institution, and extract the data items of each service field corresponding to the target category from the first historical service data;
[0137] Divide the data items corresponding to each service field into training data items and test data items;
[0138] Select one of the service fields in sequence as the current field, use the training data items corresponding to the current field as the model output, and use the training data items corresponding to the other service fields as the model input to construct the internal prediction model corresponding to the current field;
[0139] Evaluate each internal prediction model respectively using the test data items corresponding to each service field to obtain the first prediction accuracy rate of each internal prediction model;
[0140] Use the first preset number of internal prediction models with the highest first prediction accuracy rate as the target internal prediction models in the target internal prediction model set.
[0141] Further, based on the above-described embodiments of the invention, the process of obtaining the target external prediction model set based on external modeling includes:
[0142] Obtain the target category corresponding to the target institution, and extract the data items of each service field corresponding to the target category from the first historical service data and the second historical service data respectively;
[0143] Divide the data items corresponding to each service field into training data items and test data items;
[0144] Select one of the service fields in sequence as the current field, based on the training data items of each peer institution corresponding to the current field in the second historical service data, call the preset cluster data feature calculation formula to determine the cluster data training feature corresponding to the current field, and use the training data items corresponding to the current field in the first historical service data as the model output and the cluster data training feature as the model input to construct the external prediction model corresponding to the current field;
[0145] Evaluate each external prediction model respectively using the test data items corresponding to each service field and the cluster data test features to obtain the second prediction accuracy rate of each external prediction model;
[0146] Use the second preset number of external prediction models with the highest second prediction accuracy rate as the target external prediction models in the target external prediction model set.
[0147] Further, based on the above-described invention embodiments, the prediction module 32 includes:
[0148] A data item extraction unit, configured to obtain the verified data of each homogeneous institution and the target category corresponding to the target institution, and extract the data items of each service field corresponding to the target category from the to-be-verified service data and the verified data respectively;
[0149] A first service field and second service field extraction unit, configured to extract the first service fields corresponding to the respective target internal prediction models from the target internal prediction model set respectively, and extract the second service fields corresponding to the respective target external prediction models from the target external prediction model set;
[0150] A first target prediction result determination unit, configured to sequentially select one of the first service fields as the first current field, use the data items corresponding to the other service fields as the model inputs of the corresponding target internal prediction models, and obtain the first target prediction results output by the target internal prediction models;
[0151] A second target prediction result determination unit, configured to sequentially select one of the second service fields as the second current field, call a preset cluster data feature calculation formula based on the data items of the second current field corresponding to each homogeneous institution to determine the cluster data feature corresponding to the second current field, and use the cluster data feature as the model input of the corresponding target external prediction model to obtain the second target prediction results output by the target external prediction models;
[0152] An internal modeling prediction result and external modeling prediction result determination unit, configured to use the first target prediction results as the internal modeling prediction results and use the second target prediction results as the external modeling prediction results.
[0153] The data verification device provided by the embodiments of the present invention can execute the data verification method provided by any embodiment of the present invention, and has the corresponding function modules and beneficial effects for executing the method.
[0154] Embodiment 4
[0155] Figure 4FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0156] As Figure 4 shown, the electronic device 40 includes at least one processor 41, and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0157] A plurality of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0158] The processor 41 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the data verification method.
[0159] In some embodiments, the data verification method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the data verification method described above may be performed. Alternatively, in other embodiments, the processor 41 may be configured to perform the data verification method by any other suitable means (e.g., by means of firmware).
[0160] Various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC) of systems, complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0161] In some embodiments, the data verification method may be implemented as a computer program invisibly embodied in a computer program product. The computer program implements the data verification method of the present invention when executed by a processor. The computer program product may be understood as a software product that mainly implements its solution through a computer program. The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0162] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0163] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0164] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0165] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0166] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0167] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data verification method, characterized in that: The method comprises: Obtaining the business data to be verified at the target time point of the target institution, as well as the target internal prediction model set and the target external prediction model set; wherein the target internal prediction model set is obtained based on the internal modeling of the first historical business data of the target institution; and the target external prediction model set is obtained based on the external modeling of the first historical business data and the second historical business data of the same type of institution as the target institution; Based on the business data to be verified and the verification data of each of the similar institutions, the target time point is predicted according to the target internal prediction model set and the target external prediction model set, and the corresponding internal modeling prediction results and external modeling prediction results are obtained respectively; Determine the data verification results corresponding to the business data to be verified, the internal modeling prediction results, and the external modeling prediction results according to the preset classification model; Among them, based on the business data to be verified and the verification data of each of the similar institutions, the target time point is predicted according to the target internal prediction model set and the target external prediction model set, and the corresponding internal modeling prediction results and external modeling prediction results are obtained respectively, including: Obtaining the verification data of each of the same type of institutions and the target category corresponding to the target institution, and extracting data items of each business field corresponding to the target category from the business data to be verified and the verification data; Extracting the first business field corresponding to each target internal prediction model from the target internal prediction model set, and extracting the second business field corresponding to each target external prediction model from the target external prediction model set; Selecting one of the first business fields in turn as the first current field, and using data items corresponding to other business fields as model inputs of the corresponding target internal prediction model to obtain a first target prediction result output by the target internal prediction model; Selecting one of the second business fields in turn as the second current field, based on the data items of the second current field corresponding to each of the same type of institutions, calling a preset cluster data feature calculation formula to determine the cluster data feature corresponding to the second current field, and using the cluster data feature as the model input of the corresponding target external prediction model to obtain a second target prediction result output by the target external prediction model; Each of the first target prediction results is used as the internal modeling prediction result, and each of the second target prediction results is used as the external modeling prediction result.
2. The method according to claim 1, characterized in that Also includes: Obtaining historical business data corresponding to each historical time point of each institution, and performing standardization processing on the historical business data to obtain historical business samples; Based on the business fields in the preset supervision data table, the institutions are clustered using the historical business samples to determine the target category corresponding to each institution.
3. The method according to claim 2, characterized in that The clustering of the institutions using the historical business samples based on the business fields in the preset supervision data table to determine the target category corresponding to each institution includes: Roughly classify the institutions according to whether they have the corresponding business fields; Determining whether the number of institutions included in each first category after the rough classification is greater than a preset threshold of the number of institutions; If so, calling a preset clustering algorithm to cluster the historical business samples of each of the institutions in the first category, counting the number of samples of the historical business samples of each of the institutions belonging to each second category after clustering, and taking the second category with the largest number of samples as the target category of the corresponding institution; If not, the first category is used as the target category corresponding to the organization.
4. The method according to claim 1, characterized in that: The process of obtaining the target internal prediction model set based on the internal modeling includes: Acquire a target category corresponding to the target organization, and extract data items of each business field corresponding to the target category from the first historical business data; Dividing the data items corresponding to each of the business fields into training data items and test data items; Selecting one of the business fields as the current field in turn, using the training data item corresponding to the current field as the model output, and using the training data items corresponding to other business fields as the model input, so as to construct an internal prediction model corresponding to the current field; Using the test data items corresponding to the business fields to evaluate the internal prediction models, respectively, to obtain a first prediction accuracy rate of the internal prediction models; A first preset number of the internal prediction models with the highest first prediction accuracy are used as target internal prediction models in the target internal prediction model set.
5. The method according to claim 1, characterized in that: The process of obtaining the target external prediction model set based on the external modeling includes: Acquire a target category corresponding to the target organization, and extract data items of each business field corresponding to the target category from the first historical business data and the second historical business data respectively; Dividing the data items corresponding to each of the business fields into training data items and test data items; Select one of the business fields in turn as the current field, and based on the training data items of the current field corresponding to each of the same type of institutions in the second historical business data, call a preset cluster data feature calculation formula to determine the cluster data training features corresponding to the current field, and use the training data items corresponding to the current field in the first historical business data as model outputs, and the cluster data training features as model inputs, so as to construct an external prediction model corresponding to the current field; Using the test data items corresponding to the business fields and the cluster data test features respectively to evaluate the external prediction models, and obtain the second prediction accuracy of the external prediction models; A second preset number of the external prediction models with the highest second prediction accuracy are used as target external prediction models in the target external prediction model set.
6. A data verification device, characterized in that: The device comprises: An acquisition module is used to acquire the business data to be verified at a target time point of a target institution, as well as a target internal prediction model set and a target external prediction model set; wherein the target internal prediction model set is obtained based on an internal modeling of the first historical business data of the target institution; and the target external prediction model set is obtained based on an external modeling of the first historical business data and a second historical business data of a similar institution to which the target institution belongs; A prediction module, used to predict the target time point based on the business data to be verified and the verification data of each of the similar institutions according to the target internal prediction model set and the target external prediction model set, and obtain corresponding internal modeling prediction results and external modeling prediction results respectively; A verification module, used to determine the data verification results corresponding to the business data to be verified, the internal modeling prediction results and the external modeling prediction results according to a preset classification model; Wherein, the prediction module includes: A data item extraction unit, used to obtain the verification data of each of the same type of institutions and the target category corresponding to the target institution, and extract data items corresponding to each business field of the target category from the business data to be verified and the verification data; A first business field and a second business field extraction unit, used to extract the first business field corresponding to each target internal prediction model from the target internal prediction model set, and to extract the second business field corresponding to each target external prediction model from the target external prediction model set; A first target prediction result determination unit, configured to sequentially select one of the first business fields as a first current field, and use data items corresponding to other business fields as model inputs of the corresponding target internal prediction model to obtain a first target prediction result output by the target internal prediction model; A second target prediction result determination unit is used to select one of the second business fields as the second current field in turn, and based on the data items of the second current field corresponding to each of the same type of institutions, call a preset cluster data feature calculation formula to determine the cluster data feature corresponding to the second current field, and use the cluster data feature as a model input of the corresponding target external prediction model to obtain a second target prediction result output by the target external prediction model; The internal modeling prediction result and external modeling prediction result determination unit is used to use each of the first target prediction results as the internal modeling prediction result, and use each of the second target prediction results as the external modeling prediction result.
7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data verification method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data verification method according to any one of claims 1 to 5 when executed.
9. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the data verification method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data verification method and device based on artificial intelligence, equipment and storage medium
CN117421311A
Method, device and equipment for checking salary calculation result based on clustering algorithm
CN118982332A