Banking business data quality analysis processing method and device and electronic equipment

By converting banking business data, identifying and correcting abnormalities using neural network models, and verifying them in combination with context semantics, the problem of low quality of banking business data is solved, and efficient and accurate data correction and rating are achieved.

CN120494626APending Publication Date: 2025-08-15AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510624229.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Inconsistent data formats, incomplete data and dirty data problems lead to low data quality. The existing rely on manual corrections and ratings inefficient and low accuracy.

Method used

By performing data conversion processing on bank business data, using neural network models to identify exception types and perform data corrections, verifying them in combination with context semantics, and finally quality rating is performed based on the abnormal results.

Benefits of technology

It realizes efficient and accurate identification and correction of abnormalities in banking business data, significantly improves data quality, and provides reliable support for business decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494626A_ABST
    Figure CN120494626A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a banking business data quality analysis processing method and device and electronic equipment. The method comprises the following steps: performing data conversion processing on acquired service data to obtain to-be-detected data; the business data represents business transaction data of each system of the bank; processing the to-be-detected data based on the neural network model to obtain an exception type of the to-be-detected data; based on the exception type of the to-be-detected data, performing data correction processing on the to-be-detected data to obtain preliminary detection data; performing data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormal result of the preliminary detection data; based on an abnormal result of the preliminary detection data, performing quality analysis processing on the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data. The method is used for efficiently and accurately performing data correction and quality rating on the business data of the bank.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, and electronic device for quality analysis and processing of banking business data. Background Art

[0002] The sources of a bank's business data are diverse and complex, involving multiple systems and departments. However, there are problems such as inconsistent formats, incomplete data, and dirty data, which seriously affect the quality of business data. There is an urgent need to correct and optimize the business data; and to conduct a quality rating on the corrected business data.

[0003] Existing technology relies on business personnel to manually revise bank business data based on data specifications and business rules. Business personnel then manually rate the quality of the revised business data. However, this approach is inefficient and has low accuracy.

[0004] Therefore, there is an urgent need for a solution that can efficiently and accurately correct and rate the quality of a bank's business data. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, and electronic device for analyzing and processing the quality of banking business data, so as to achieve efficient and accurate data correction and quality rating of banking business data.

[0006] In a first aspect, an embodiment of the present application provides a method for quality analysis and processing of banking business data, comprising:

[0007] Performing data conversion processing on the acquired business data to obtain data to be tested; the business data represents business transaction data of various systems of the bank;

[0008] Processing the data to be detected based on a neural network model to obtain an abnormality type of the data to be detected; performing data correction processing on the data to be detected based on the abnormality type of the data to be detected to obtain preliminary detection data; performing data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormality result of the preliminary detection data;

[0009] Based on the abnormal results of the preliminary detection data, quality analysis processing is performed on the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

[0010] In a second aspect, an embodiment of the present application provides a device for analyzing and processing quality of banking business data, comprising:

[0011] An acquisition module is used to perform data conversion processing on the acquired business data to obtain data to be tested; the business data represents business transaction data of various systems of the bank;

[0012] a processing module configured to process the data to be detected based on a neural network model to obtain an abnormality type of the data to be detected; perform data correction processing on the data to be detected based on the abnormality type of the data to be detected to obtain preliminary detection data; and perform data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormality result of the preliminary detection data;

[0013] The analysis module performs quality analysis on the preliminary detection data based on the abnormal results of the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;

[0015] The memory stores computer-executable instructions;

[0016] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0019] The embodiments of the present application provide a method, device, and electronic device for quality analysis and processing of banking business data, which perform data conversion processing on the acquired business data representing the business transaction data of various systems of the bank to obtain data to be detected; then, the data to be detected is processed based on a neural network model to identify the type of anomaly existing in the data to be detected; then, based on the determined anomaly type, the data to be detected is corrected to eliminate or reduce the anomaly in the data, thereby obtaining preliminary detection data; then, the preliminary detection data is subjected to data verification processing in combination with contextual semantics to comprehensively identify and mark the abnormal results existing in the preliminary detection data; finally, based on the abnormal results of the preliminary detection data, the preliminary detection data is quality analyzed and processed, and finally a quality rating result that can characterize the business data score is obtained, thereby achieving efficient and accurate identification and correction of anomalies in banking business data, significantly improving data quality, and providing more reliable data support for the bank's business decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0021] Figure 1 A flow chart of a method for quality analysis and processing of banking business data provided in an embodiment of the present application Figure 1 ;

[0022] Figure 2 A flow chart of a method for quality analysis and processing of banking business data provided in an embodiment of the present application Figure 2 ;

[0023] Figure 3 A schematic diagram of the process of a quality analysis and processing device for banking business data provided in an embodiment of the present application Figure 1 ;

[0024] Figure 4 A schematic diagram of the process of a quality analysis and processing device for banking business data provided in an embodiment of the present application Figure 2 ;

[0025] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application.

[0026] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0028] Banks' business data comes from a wide range of complex sources, spanning multiple systems and departments. This data may include customer information, transaction records, account balances, credit data, and management data. Due to the varying business requirements and technical architectures of these systems and departments, data formats and structures vary significantly, severely impacting the quality of business data.

[0029] Currently, banks rely primarily on manual operations for data correction and quality rating. These personnel manually review and correct errors and inconsistencies in the data based on data specifications and business rules. This approach is not only time-consuming and labor-intensive, but also prone to inaccurate corrections due to human error. The personnel then perform quality ratings on the corrected data, often based on experience and subjective judgment. This approach lacks standardization and consistency, making it difficult to ensure the accuracy and reliability of the rating results.

[0030] Therefore, the method, device and electronic device for quality analysis and processing of banking business data provided in this application can solve the above problems.

[0031] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0032] Figure 1 A flow chart of a method for quality analysis and processing of banking business data provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the method includes:

[0033] S101. Perform data conversion processing on the acquired business data to obtain data to be tested; the business data represents business transaction data of various systems of the bank.

[0034] For example, data is obtained from business transaction data in various banking systems. These systems may include core banking systems (used to handle basic services such as account management and funds transfers), credit card systems, online banking systems, and mobile banking systems. For example, core banking systems record customer account deposits, withdrawals, transfers, and other transaction information, while credit card systems record credit card purchases, repayments, and other business data. Data conversion processing is performed on the obtained business data to obtain the data to be tested; this data conversion processing includes data integration, data conversion, and data standardization.

[0035] Data integration is the integration of business data from different banking systems (such as core systems and credit systems), and the association of different data sources through unique identifiers (such as customer IDs and transaction serial numbers). Split complex fields into multiple simple fields (such as splitting "name + ID number" into "name" and "ID number"), or merge multiple related fields into one (such as merging "year", "month" and "day" into "transaction date"). Data conversion is the splitting of complex fields into multiple simple fields (such as splitting "name + ID number" into "name" and "ID number"), or merge multiple related fields into one (such as merging "year", "month" and "day" into "transaction date"). Data standardization is the unification of the units of fields such as amount and time (such as converting all amounts into "yuan" and unifying the time format into "YYYY-MM-DD HH:MM:SS").

[0036] S102. Process the data to be detected based on the neural network model to obtain the abnormality type of the data to be detected; based on the abnormality type of the data to be detected, perform data correction processing on the data to be detected to obtain preliminary detection data; perform data verification processing on the preliminary detection data in combination with context semantics to obtain the abnormality result of the preliminary detection data.

[0037] For example, the data to be tested, after data conversion, is input into a neural network model. This data can be structured tabular data containing multiple fields (such as transaction amount, transaction time, account number, etc.). The neural network model identifies anomalies in the data by learning the characteristics and patterns of the data. Model training is usually based on a large amount of historical data that has been labeled with anomaly types (such as data duplication, data missing, data logic anomalies, etc.). The model outputs the anomaly type of the data to be tested. For example, the model may identify a "data duplication anomaly" in a certain record, or a "data logic anomaly in the transaction amount field" in a certain record. Based on the test results, anomalies can be divided into data duplication anomalies, data missing anomalies, and data logic anomalies. Data duplication anomalies refer to the presence of completely duplicate records in the dataset; data missing anomalies refer to missing values in the data (such as empty values, NULL, or unfilled fields); and data logic anomalies refer to data that violates business logic (such as negative transaction amounts, incorrect date sequence, etc.).

[0038] Data correction for duplicate records involves directly deleting the duplicate records (if they have no business significance) or retaining one record and consolidating related information (such as summing up the transaction amounts). For example, if two identical transaction records are detected, one can be deleted or merged into a single record and the amounts added up.

[0039] Data correction for missing data anomalies can be done using two methods: deletion and filling. Deletion involves deleting records containing missing values (applicable to fields with a high percentage of missing values or no business significance). Filling methods include fixed value filling, calculated value filling, and model prediction filling. Fixed value filling involves filling with default values (such as 0 or "UNKNOWN"); calculated value filling involves filling numeric fields with the mean, median, or mode; and categorical fields with the mode. Model prediction filling involves using a model to predict missing values (such as linear regression or KNN) and then filling the missing values.

[0040] Data correction for logical anomalies involves two methods: rule-based correction and predicted data correction. Rule-based correction involves correcting outliers based on pre-set business rules (e.g., correcting negative transaction amounts to 0). Predicted data correction involves using a model to predict data at logical anomalies and then using the predicted data to fill in the gaps.

[0041] After data correction, preliminary test data is obtained. By incorporating contextual semantics, this preliminary test data is verified to identify potential data anomalies (i.e., those that were not fully detected by the previous steps, or data that appears abnormal due to contextual relationships). Contextual semantic verification captures logical relationships, business rules, and consistency between data, thereby improving data quality.

[0042] Based on the preliminary detection data, feature extraction is performed based on contextual semantics to obtain features associated with the contextual semantics. These extracted features are then input into an encoding model for processing, generating a hidden state matrix. Each row of the hidden state matrix represents a high-dimensional representation of a record, capturing key elements of the semantic features of the preliminary detection data as well as contextual information. Encoding models can include autoencoders, variational autoencoders, and embedding layer encoders. Using the classification layer in the neural network model, classification operations are performed on the hidden state matrix to identify anomalous data. Data classified as anomalous is recorded and an anomaly report is generated, providing detailed information about the anomalous record and the classification confidence level.

[0043] S103. Based on the abnormal results of the preliminary detection data, perform quality analysis on the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

[0044] For example, based on abnormal data and abnormal reports, the distribution of abnormal data in the preliminary detection data is determined, such as: the number and proportion of abnormal data records; the distribution of abnormal data in different fields or business scenarios. The formula for calculating the abnormal data proportion is: abnormal data proportion = number of abnormal data records / total number of records in the preliminary detection data × 100%. The calculated abnormal data proportion is compared with the pre-set quality rating standard to determine the quality rating result of the preliminary detection data. For example: the abnormal data proportion is between 0%-10%, indicating that the data quality is good. Although there is a certain proportion of abnormal data, it will not have a serious impact on the business. The abnormal data proportion is greater than or equal to 10%, indicating that there are obvious problems with the data quality, which has or may have a significant impact on the business, and measures need to be taken to improve it.

[0045] The embodiment of the present application provides a quality analysis and processing method for bank business data, which performs data conversion processing on the acquired business data representing the business transaction data of various systems of the bank to obtain data to be detected; then, the data to be detected is processed based on a neural network model to identify the type of anomaly existing in the data to be detected; then, based on the determined anomaly type, the data to be detected is corrected to eliminate or reduce the anomaly in the data, thereby obtaining preliminary detection data; then, the preliminary detection data is subjected to data verification processing in combination with context semantics to comprehensively identify and mark the abnormal results existing in the preliminary detection data; finally, based on the abnormal results of the preliminary detection data, the preliminary detection data is quality analyzed and processed, and finally a quality rating result that can characterize the business data score is obtained, thereby achieving efficient and accurate identification and correction of anomalies in bank business data, significantly improving data quality, and providing more reliable data support for the bank's business decisions.

[0046] Figure 2 A flow chart of a method for quality analysis and processing of banking business data provided in an embodiment of the present application Figure 2 ,like Figure 2 As shown, this embodiment Figure 1 Based on the embodiment, a method for quality analysis and processing of banking business data is described in detail. The method includes:

[0047] S201. Determine the data conversion method corresponding to the business data based on the data source information of the business data and the preset data conversion rules; wherein the data source information represents the system information and format information of the business data; the preset data conversion rules indicate the data conversion method for the business data under different system information and different format information; perform standardized text processing on the business data according to the data conversion method corresponding to the business data to obtain the data to be tested.

[0048] For example, the system information in the data source information indicates the source system of the business data. Different systems may have different data storage, processing, and transmission mechanisms. Format information describes the specific storage format of the business data. Common data formats include CSV (Comma Separated Values), JSON (JavaScript Object Notation), XML (Extensible Markup Language), and database tables. Different formats differ in data structure, field definitions, data types, and other aspects. The preset data conversion rules are a set of predefined guidelines that indicate the data conversion method to be used for business data with different system information and different format information. The data conversion method includes the specific execution process for converting business data from the original format to the target format. This may involve a series of operations such as data parsing, field mapping, data cleansing, and format conversion. For example, when converting CSV to JSON, you first need to read the CSV file, parse the data line by line, then map the field values of each line to the corresponding properties of the JSON object, and finally serialize the JSON object into a JSON string. During the conversion process, data standardization is required to ensure data consistency and accuracy. For example, standardize field naming rules (e.g., unify the "customer name" field in different systems to "customer_name"), remove redundant spaces, format dates and numbers (e.g., unify dates to "YYYY-MM-DD" format), etc. After conversion and standardization, the data obtained is the data to be tested.

[0049] S202: Based on the abnormality type of the data to be detected, perform data correction processing on the data to be detected to obtain preliminary detection data.

[0050] For example, this step may refer to the above-mentioned step S102 and will not be described in detail.

[0051] In one example, the anomaly type of the data to be detected is a data duplication anomaly. Based on the data to be detected and the duplicate data in the data to be detected, the cosine similarity of the duplicate data in the data to be detected is determined; if the cosine similarity is greater than or equal to a preset threshold, the duplicate data in the data to be detected is deleted to obtain preliminary detection data.

[0052] For example, duplicate data refers to two or more pieces of data in the data set to be detected that are exactly the same or highly similar in content. This duplication may be caused by duplicate data on different systems. Cosine similarity is an indicator that measures the similarity between two vectors, and its value range is [-1,1]. In text processing, it is often used to measure the similarity between two texts. For two text data, first convert them into vector representations (such as TF-IDF vectors), and then calculate the cosine similarity between them. Convert the text data to be compared into vectors. The TF-IDF (term frequency-inverse document frequency) method can be used to represent the text as a weighted combination of word vectors. Use the cosine similarity formula to calculate the similarity between two vectors. The calculation formula is:

[0053]

[0054] Among them, C(A,B) represents the cosine similarity between vector A and vector B; A and B represent two similar text vectors respectively; A·B represents the dot product of vector A and vector B; ||A|| represents the norm of vector A; ||B|| represents the norm of vector B.

[0055] Based on business needs and data characteristics, set a cosine similarity threshold. For example, setting a threshold of 0.9 means that two data points are considered duplicates when their cosine similarity is greater than or equal to 0.9. If the cosine similarity is greater than or equal to the preset threshold, the data is considered duplicate and one of the data points is deleted (usually, only one is retained). After deleting the duplicate data, update the dataset to obtain preliminary test data.

[0056] For example, opening bank accounts through different channels may generate duplicate records. Data 1: Customer A's information (branch account opening): Zhang Wei, 135****5678, Engineer; Data 2: Customer A's information (mobile banking): Zhang Wei, 135****5678, Software Engineer. Data 1 and 2 are converted into vector representations and the cosine similarity between the two vectors is calculated. If the cosine similarity is greater than or equal to 0.9, they are considered duplicates. Data 1 is deleted to obtain preliminary test data.

[0057] In one example, the abnormality type of the data to be detected is a data missing abnormality, and the position of the missing data in the data to be detected is masked; based on the context data of the mask at the position of the data to be detected and the missing data, the missing data at the position of the missing data is predicted to obtain a set of predicted data; wherein the context data represents the data within a preset length on the left and right sides of the mask; the masked predicted data set includes at least one first predicted data, and the first predicted data represents the missing data predicted at the position of the missing data in the detection data; based on the data to be detected and the predicted data set, the perplexity of the data to be detected is determined; the perplexity represents the degree of matching between the first predicted data in the predicted data set and the data to be detected; the first predicted data with the smallest perplexity in the predicted data set is determined to be the missing data in the data to be detected, and preliminary detection data is obtained.

[0058] For example, when the anomaly type of the data to be detected is missing data, the missing data position in the data to be detected is replaced with a specific mask tag (such as [MASK], None, or other custom tags). For example, in the user information table above, the "age" position in the record where the "age" field is empty is replaced with [MASK].

[0059] Context data is defined as data within a preset length on both sides of the mask. This preset length can be determined based on the data characteristics and business requirements. For example, when processing natural language text, the preset length might be five words to the left and right of the mask position; when processing tabular data, the preset length might be three fields before and after the row containing the mask position. Context data is extracted from the data to be detected, centered around the mask position.

[0060] Predictive models (such as deep learning-based language models and machine learning algorithms) are used to predict missing data based on contextual data. The predictive model generates multiple possible prediction results, which constitute the prediction data set.

[0061] Perplexity is a metric that measures the performance of a language model or the degree of match between predicted data and the original data. In data-missing exception handling, perplexity indicates the degree of match between the first predicted data in a prediction dataset and the data to be tested. Lower perplexity indicates a higher degree of match between the predicted data and the original data.

[0062] For each first prediction data in the prediction data set, replace it with the missing position of the data to be tested to form a complete data sequence. Use a language model or other appropriate model to calculate the probability of the complete data sequence. The perplexity calculation formula is:

[0063]

[0064] Among them, PP(W) represents the perplexity, W represents the entire data sequence; P(w1,w2,···,w N ) represents the data sequence w1,w2,···,w N This probability is calculated by the language model based on the relationship between each word in the sequence;

[0065] Compare the perplexities of all the first predictions in the prediction data set and find the first prediction with the lowest perplexity. The lowest perplexity means that the prediction best matches the context of the data to be tested and is most likely the missing real data. Fill the missing position of the data to be tested with the first prediction with the lowest perplexity to obtain the preliminary test data.

[0066] For example, in a loan application, a customer may not fill out the "Occupation" field. The data to be tested is "Data 3: Customer Information: Name: Zhang Wei, Age: 28, Contact: 135****5678, Occupation: , Annual Income: 1,000,000 RMB, Credit Score: 800;." First, [MASK] is used to mark the missing data, resulting in "Data 3: Customer Information: Name: Zhang Wei, Age: 28, Contact: 135****5678, Occupation: [MASK], Annual Income: 1,000,000 RMB, Credit Score: 800;." Centered around the [MASK] position, the corresponding contextual data is extracted from Data 3. A predictive model is used to predict the missing data based on this contextual data. The model makes predictions based on the common relationships between "age" and "income" and "occupation." For example, a customer aged 28 with an annual income of 1,000,000 RMB might have an occupation such as teacher or software engineer. The model might predict "teacher" or "software engineer" as the first prediction data, which constitutes the predicted data set for the "Occupation" field. When predicting "Occupation" as "Software Engineer", we fully record "Data 3: Customer Information: Name: Zhang Wei, Age: 28, Contact: 135****5678, Occupation: Software Engineer, Annual Income: 1 million yuan, Credit Score: 800;". The frequency of occurrence in the frequency table is 0.25, the probability P = 0.25, the sequence length N = 5, and the perplexity When predicting "occupation" as "teacher", the complete record is "Data 3: Customer Information: Name: Zhang Wei, Age: 28, Contact: 135****5678, Occupation: Teacher, Annual Income: 1 million yuan, Credit Score: 800;". The frequency of occurrence in the frequency table is 0.15, the probability P = 0.15, the sequence length N = 5, and the perplexity Comparing the perplexity of different prediction data, the perplexity of "software engineer" corresponding to 1.32 is the smallest, so the preliminary test data is "Data 3: Customer information: Name: Zhang Wei, Age: 28, Contact: 135****5678, Occupation: Software Engineer, Annual Income: 1 million yuan, Credit Score: 800;".

[0067] In one example, the anomaly type of the data to be detected is a data logic anomaly. The data to be detected is encoded and semantic feature extracted in sequence to obtain the semantic features of the data to be detected; based on the semantic features and the context information of the logically abnormal data in the data to be detected, the logically abnormal data in the data to be detected is predicted to obtain second predicted data corresponding to the logically abnormal data in the data to be detected; the second predicted data is used to replace the logically abnormal data in the data to be detected to obtain preliminary detection data.

[0068] Exemplarily, the data to be detected is subjected to encoding processing and semantic feature extraction processing in sequence to obtain the semantic features of the data to be detected; wherein, the encoding processing methods may include one-hot encoding, word embedding encoding, etc. Semantic features can reflect the semantic information of the data to be detected. For example, in bank customer data, semantic features may include encoded numerical features such as the customer's age, occupation, credit score, and word vector representation of customer feedback text. Determine the context scope of logical anomaly data. The context can be the previous and next data points in the time series (such as the previous and next transactions of an abnormal transaction in the transaction record), spatially adjacent data (such as the surrounding area data of an abnormal point in the geographic location data) or logically associated data (such as other relevant information of a customer in the customer information). Extract the context information and convert it into the same encoding form as the semantic feature so that it can be input into the prediction model together with the semantic feature matrix.

[0069] The prediction model is used to predict the logically abnormal data in the data to be detected based on the semantic features and context information of the data to be detected to obtain second predicted data. The second predicted data is used to replace the logically abnormal data in the data to be detected to obtain preliminary detection data.

[0070] S203. Perform semantic segmentation on the preliminary detection data in combination with the contextual semantics to obtain subwords of the preliminary detection data; encode the subwords using an encoder based on the anomaly detection model to obtain a hidden state matrix corresponding to the subwords; wherein the hidden state matrix represents the semantic encoding of the subwords in the context of the preliminary detection data; perform classification processing on the hidden state matrix based on the classification layer of the anomaly detection model to obtain anomaly results of the preliminary detection data.

[0071] For example, the preliminary detection data is first segmented based on the contextual semantics to obtain subwords. The subwords are then encoded using the anomaly detection model's encoder to obtain the corresponding hidden state matrix. Finally, the hidden state matrix is classified using the anomaly detection model's classification layer to obtain anomaly results for the preliminary detection data.

[0072] When processing preliminary test data, the word segmentation model is first used to segment the data using contextual semantic information, thereby obtaining a subword sequence for the preliminary test data. For Chinese data, dictionary-based word segmentation methods can be used. These methods rely on a preset dictionary to identify words and are highly accurate and efficient. Alternatively, statistical-based word segmentation methods such as Hidden Markov Models (HMMs) or Conditional Random Fields (CRFs) can be used. These methods automatically learn word boundaries based on statistical patterns in the corpus and are suitable for processing large-scale text data. For English data, in simple cases, word segmentation can be performed based on spaces or punctuation marks. This method is suitable for text with relatively regular structure. When more sophisticated processing is required, more complex word segmentation methods such as Byte Pair Encoding (BPE) or WordPiece can be used. These methods can effectively handle variants and rare words in English and improve the model's lexical coverage.

[0073] After word segmentation, the resulting subword sequence is input into the Transformer encoder for processing. Through its self-attention mechanism and multi-layered structure, the Transformer encoder fully captures the semantic information of subwords in context and generates a hidden state matrix. Each row of this matrix corresponds to the hidden state representation of a subword. These hidden states not only incorporate the semantic features of the subword itself but also incorporate the influence of its context, providing rich semantic information for subsequent classification tasks. A classification layer is then added to the Transformer encoder, typically consisting of a fully connected layer and a softmax function. After the hidden state matrix undergoes a linear transformation in the fully connected layer, the softmax function converts it into a classification probability distribution, outputting the probability value for each category. Ultimately, the category with the highest probability is selected as the classification result for the data, determining whether it is "normal" or "abnormal."

[0074] S204. Determine the abnormal ratio of the abnormal data in the preliminary detection data to the data to be detected based on the abnormal result of the preliminary detection data; and determine the quality rating result based on the value of the abnormal ratio.

[0075] For example, based on the anomaly results of the preliminary detection data, the proportion of anomaly data in the data to be detected is calculated. The anomaly results of the preliminary detection data are traversed, and the number of data items marked as anomalies is counted, denoted as M anomaly . The total number of data items to be detected is determined, denoted as M total . The anomaly ratio P anomaly is calculated as: P anomaly = (N anomaly / N total) × 100%. Based on the value of the anomaly ratio, a quality rating result of the data to be detected is determined to determine the data quality.

[0076] S205. If the score represented by the quality rating result is less than or equal to the first threshold, a first-level alarm message is issued to the data source of the business data; if the score represented by the quality rating result is greater than the first threshold, a second-level alarm message is issued to the data source of the business data.

[0077] For example, if the score represented by the quality rating result is less than or equal to 10%, it indicates that the data quality is good. Although there is a certain proportion of abnormal data, it will not have a serious impact on the business. In this case, a level 1 alarm message is issued to the data source of the business data. If the score represented by the quality rating result is greater than 10%, it indicates that there are obvious problems with the data quality, which has or may have a significant impact on the business, and measures need to be taken to improve it. Based on this, a level 2 alarm message should be issued to the data source of the business data.

[0078] There are various ways to send alerts. The system can promptly notify relevant responsible persons via email, providing detailed abnormality information and handling suggestions. Secondly, it can quickly notify executives or teams through SMS services to ensure that key personnel are aware of abnormal situations immediately. In addition, the system can also send notifications via instant messaging tools, facilitating real-time communication and collaboration among team members.

[0079] The embodiment of the present application provides a quality analysis and processing method for banking business data. Based on the data source information (including system and format information) of the business data and preset data conversion rules, the data conversion method is determined, and the business data is subjected to standardized text processing to obtain the data to be detected; for data duplication anomalies, the cosine similarity of the duplicate data is calculated and the duplicate data with high similarity is deleted; for data missing anomalies, the missing data is predicted through masking processing and context data, and the predicted data is selected based on the perplexity; for data logical anomalies, semantic features are extracted and logical anomaly data is predicted and replaced in combination with context information; finally, the preliminary detection data is subjected to semantic segmentation, encoding and classification processing in combination with context semantics to obtain anomaly results, and the quality rating result is determined according to the anomaly ratio, so as to achieve accurate identification and efficient processing of various anomalies in banking business data, significantly improve data accuracy and consistency, and provide banks with high-quality and reliable data support.

[0080] Figure 3A schematic diagram of a device for analyzing and processing quality of banking business data provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the present embodiment provides a banking business data quality analysis and processing device 30 comprising:

[0081] The acquisition module 301 is used to perform data conversion processing on the acquired business data to obtain data to be tested; the business data represents the business transaction data of various systems of the bank;

[0082] The processing module 302 is configured to process the data to be detected based on the neural network model to obtain an abnormality type of the data to be detected; perform data correction processing on the data to be detected based on the abnormality type of the data to be detected to obtain preliminary detection data; perform data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormality result of the preliminary detection data;

[0083] The analysis module 303 performs quality analysis on the preliminary detection data based on the abnormal results of the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

[0084] This embodiment provides a device for analyzing and processing quality of banking business data, which can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0085] Figure 4 A schematic diagram of a device for analyzing and processing quality of banking business data provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the present embodiment provides a banking business data quality analysis and processing device 40 comprising:

[0086] The acquisition module 401 is used to perform data conversion processing on the acquired business data to obtain data to be detected; the business data represents the business transaction data of various systems of the bank;

[0087] Processing module 402 is configured to process the data to be detected based on the neural network model to obtain an anomaly type of the data to be detected; perform data correction processing on the data to be detected based on the anomaly type of the data to be detected to obtain preliminary detection data; perform data verification processing on the preliminary detection data in combination with contextual semantics to obtain an anomaly result of the preliminary detection data;

[0088] The analysis module 403 performs quality analysis on the preliminary detection data based on the abnormal results of the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

[0089] In a possible implementation, the acquisition module 401 includes:

[0090] Determine the data conversion method corresponding to the business data based on the data source information of the business data and the preset data conversion rules; wherein the data source information represents the system information and format information of the business data; the preset data conversion rules indicate the data conversion method for business data with different system information and different format information;

[0091] According to the data conversion method corresponding to the business data, the business data is processed into standardized text to obtain the data to be tested.

[0092] In a possible implementation, the processing module 402 includes:

[0093] The abnormal type of the data to be detected is data duplication abnormality, and the cosine similarity of the duplicate data in the data to be detected is determined according to the data to be detected and the duplicate data in the data to be detected;

[0094] If the cosine similarity is greater than or equal to the preset threshold, the duplicate data in the data to be detected are deleted to obtain preliminary detection data.

[0095] In a possible implementation, the processing module 402 includes:

[0096] The anomaly type of the data to be detected is data missing anomaly, and the position of the missing data in the data to be detected is masked;

[0097] Predicting missing data at the location of the missing data based on context data of the mask at the location of the data to be detected and the missing data to obtain a set of predicted data; wherein the context data represents data within a preset length on both sides of the mask; the set of masked predicted data includes at least one first predicted data, the first predicted data representing the missing data predicted at the location of the missing data in the detected data;

[0098] Based on the data to be detected and the predicted data set, the perplexity of the data to be detected is determined; the perplexity represents the degree of matching between the first predicted data in the predicted data set and the data to be detected; the first predicted data with the smallest perplexity in the predicted data set is determined to be the missing data in the data to be detected, and preliminary detection data is obtained.

[0099] In a possible implementation, the processing module 402 includes:

[0100] The anomaly type of the data to be detected is data logic anomaly. The data to be detected is encoded and semantic feature extracted in sequence to obtain the semantic features of the data to be detected.

[0101] Predicting the data with logical anomalies in the data to be detected based on the semantic features and context information of the data with logical anomalies in the data to be detected to obtain second predicted data corresponding to the data with logical anomalies in the data to be detected;

[0102] The second prediction data is used to replace the logically abnormal data in the data to be detected to obtain preliminary detection data.

[0103] In a possible implementation, the processing module 402 further includes:

[0104] Perform semantic segmentation on the preliminary detection data in combination with the context semantics to obtain the subwords of the preliminary detection data;

[0105] The encoder based on the anomaly detection model encodes the subwords to obtain a hidden state matrix corresponding to the subwords; wherein the hidden state matrix represents the semantic encoding of the subwords in the context of the preliminary detection data;

[0106] Based on the classification layer of the anomaly detection model, the hidden state matrix is classified and processed to obtain the preliminary anomaly results of the detection data.

[0107] In a possible implementation, the analysis module 403 includes:

[0108] According to the abnormal results of the preliminary test data, determine the abnormal ratio of the abnormal data in the preliminary test data to the abnormal data to be tested;

[0109] The quality rating result is determined based on the value of the abnormal ratio.

[0110] In a possible implementation, the analysis module 403 further includes:

[0111] If the score represented by the quality rating result is less than or equal to the first threshold, a first-level alarm message is issued to the data source of the business data;

[0112] If the score represented by the quality rating result is greater than the first threshold, a second-level alarm message is issued to the data source of the business data.

[0113] This embodiment provides a device for analyzing and processing quality of banking business data, which can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0114] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 5As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.

[0115] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that the at least one processor 501 performs the above method.

[0116] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0117] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.

[0118] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0119] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0120] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0121] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0122] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0123] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0124] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0125] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0126] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0127] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0128] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0129] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A method for quality analysis and processing of banking business data, characterized in that: include: Performing data conversion processing on the acquired business data to obtain data to be tested; the business data represents business transaction data of various systems of the bank; Processing the data to be detected based on a neural network model to obtain an abnormality type of the data to be detected; Based on the abnormality type of the data to be detected, performing data correction processing on the data to be detected to obtain preliminary detection data; Performing data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormal result of the preliminary detection data; Based on the abnormal results of the preliminary detection data, quality analysis processing is performed on the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

2. The method according to claim 1, characterized in that The business data includes data source information; Perform data conversion on the acquired business data to obtain the data to be tested, including: Determining a data conversion method corresponding to the business data based on data source information of the business data and preset data conversion rules; wherein the data source information represents system information and format information of the business data; and the preset data conversion rules indicate data conversion methods for business data with different system information and different format information; According to the data conversion method corresponding to the business data, the business data is subjected to standardized text processing to obtain data to be detected.

3. The method according to claim 1, characterized in that The abnormality type of the data to be detected is a data duplication abnormality. Based on the abnormality type of the data to be detected, data correction processing is performed on the data to be detected to obtain preliminary detection data, including: Determining the cosine similarity of the repeated data in the data to be detected according to the data to be detected and the repeated data in the data to be detected; If the cosine similarity is greater than or equal to a preset threshold, duplicate data in the data to be detected is deleted to obtain preliminary detection data.

4. The method according to claim 1, wherein The abnormality type of the data to be detected is a data missing abnormality. Based on the abnormality type of the data to be detected, data correction processing is performed on the data to be detected to obtain preliminary detection data, including: Masking the positions of missing data in the data to be detected; Predicting missing data at the location of the missing data based on the data to be detected and the context data of the mask at the location of the missing data to obtain a set of predicted data; wherein the context data represents data within a preset length on both sides of the mask; and the set of masked predicted data includes at least one first predicted data, wherein the first predicted data represents the missing data predicted at the location of the missing data in the detected data; Based on the data to be detected and the predicted data set, the perplexity of the data to be detected is determined; the perplexity represents the degree of matching between the first predicted data in the predicted data set and the data to be detected; the first predicted data with the smallest perplexity in the predicted data set is determined to be the missing data in the data to be detected, and the preliminary detection data is obtained.

5. The method according to claim 1, wherein The abnormality type of the data to be detected is a data logic abnormality. Based on the abnormality type of the data to be detected, data correction processing is performed on the data to be detected to obtain preliminary detection data, including: Performing encoding processing and semantic feature extraction processing on the data to be detected in sequence to obtain semantic features of the data to be detected; performing prediction processing on the logically abnormal data in the data to be detected based on the semantic features and context information of the logically abnormal data in the data to be detected to obtain second predicted data corresponding to the logically abnormal data in the data to be detected; The second prediction data is used to replace the logically abnormal data in the data to be detected to obtain the preliminary detection data.

6. The method according to claim 1, characterized in that Performing data verification processing on the preliminary detection data in combination with context semantics to obtain abnormal results of the preliminary detection data includes: Performing semantic segmentation processing on the preliminary detection data in combination with context semantics to obtain subwords of the preliminary detection data; An encoder based on an anomaly detection model encodes the subword to obtain a hidden state matrix corresponding to the subword; wherein the hidden state matrix represents the semantic encoding of the subword in the context of the preliminary detection data; Based on the classification layer of the anomaly detection model, the hidden state matrix is classified to obtain an anomaly result of the preliminary detection data.

7. The method according to any one of claims 1 to 6, characterized in that Based on the abnormal results of the preliminary test data, the preliminary test data is subjected to quality analysis and processing to obtain a quality rating result, including: Determining, based on the abnormal results of the preliminary detection data, the abnormal proportion of the abnormal data in the preliminary detection data to the abnormal data to be detected; The quality rating result is determined based on the value of the abnormal ratio.

8. The method according to claim 7, characterized in that The method comprises: If the score represented by the quality rating result is less than or equal to a first threshold, a first-level alarm message is issued to the data source of the business data; If the score represented by the quality rating result is greater than a first threshold, a second-level alarm message is issued to the data source of the business data.

9. A quality analysis and processing device for banking business data, characterized in that: include: The acquisition module is used to perform data conversion processing on the acquired business data to obtain the data to be detected; The business data represents the business transaction data of various systems of the bank; A processing module, configured to process the data to be detected based on a neural network model to obtain an abnormality type of the data to be detected; Based on the abnormality type of the data to be detected, performing data correction processing on the data to be detected to obtain preliminary detection data; Performing data verification processing on the preliminary detection data in combination with context semantics to obtain an abnormal result of the preliminary detection data; The analysis module performs quality analysis on the preliminary detection data based on the abnormal results of the preliminary detection data to obtain a quality rating result; wherein the quality rating result represents the score of the business data.

10. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8.