An artificial intelligence system for data standardization
The AI system addresses data standardization issues in banking systems by using Z-score normalization and similarity comparison to correct and reorganize borrower data, ensuring accurate credit evaluations.
Patent Information
- Application Number
- CN202411806863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The data format and content of borrower information in the banking system are complex and diverse, with inconsistencies and errors, which lead to data confusion and affect borrower credit assessment.
The data processing module is used to obtain information through the API interface protocol, and the format string and string replacement method are used to unify the format, combine the Z-score standardization method and similarity comparison method to judge and re-divider abnormal data to ensure the accuracy of data standardization processing.
Effectively eliminate data format differences, reduce errors, ensure the accuracy of data cleaning, avoid credit assessment deviations caused by data confusion, and realize accurate analysis of borrower information and credit assessment.
Smart Images

Figure CN119648389B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data standardization. Specifically, it relates to an artificial intelligence system for data standardization. Background Art
[0002] In the information system of a bank, a vast amount of borrower information is stored. This information is crucial for business processes such as risk assessment, loan approval, and credit monitoring. Borrower information includes various types of data, such as personal identity information, financial status, credit records, loan history, etc.
[0003] However, the formats and contents of the above data are usually extremely complex and diverse, and there are inconsistencies and errors in data from different sources. If these data cannot be standardized, it will have a negative impact on the assessment of borrower information.
[0004] Meanwhile, during the process of standardizing data, the data sources in the bank information system are mainly the information reserved by borrowers when handling business, and the content entered into the bank system by bank customer service after receiving borrower information;
[0005] But there are some problems in this process. On the one hand, due to the different handwriting of different borrowers, it is easy to have confusion between numbers and similar characters. On the other hand, when a large amount of data is manually entered, data may be confused in the data table due to lack of line breaks. If these two types of data confusion situations cannot be reasonably judged in a timely manner, it will affect the assessment of the borrower's credit. In view of this, we propose an artificial intelligence system for data standardization. Summary of the Invention
[0006] The purpose of the present invention is to solve the situation where, when assessing the borrower's credit, due to data confusion problems in the bank system, abnormal borrower information data appears during the data standardization process;
[0007] To achieve the above purpose, the present invention provides an artificial intelligence system for data standardization that can reclassify abnormal borrower information data during data standardization, including a data processing module, a data standardization module, a data confusion analysis module, and a data classification module;
[0008] The data processing module obtains borrower information through the API interface protocol, unifies the borrower information respectively through the format string and the string replacement method. The data standardization module is used to establish a data set corresponding to the borrower information, standardize the data set through the Z-score standardization method, and judge the abnormal values in the data set by using the abnormal value definition method;
[0009] The data confusion analysis module uses the similarity comparison method to determine whether the abnormal value is caused by the string replacement method in the data processing module. If it is caused by the string replacement method, it indicates that the string replacement method confuses characters and numbers at this time, and the abnormal value will be replaced again. If it is not caused by the string replacement method, it indicates that the abnormal value in the data standardization module is normal;
[0010] The data division module includes an unconfused data judgment unit and a re-division unit;
[0011] The unconfused data judgment unit is used to receive the number of abnormal values judged by the data confusion analysis module that are not caused by the string replacement method. Then, it judges whether the abnormal values are adjacent in the data set based on the index position, sets a quantity threshold, and when the number of adjacent abnormal values > the quantity threshold, it indicates that the data in the data set is confused;
[0012] The re-division unit is used to receive the data confusion signal, re-divide the confused data through the abnormal value definition method in the data standardization module. After re-division, the data standardization process is performed again, and the standardized data is output to the staff.
[0013] The data processing module is used to send a data request and an API key to the bank information system. The data request includes the borrower's account information, the borrower's credit report, the date format information, and the financial information submitted by the borrower himself. After receiving the request, the bank information system compares the API key with the legitimate key. If the API key = the legitimate key, it indicates that the verification is successful. At this time, the bank information system retrieves the borrower's account information, the borrower's credit report, the date format information, and the financial information submitted by the borrower himself according to the borrower's identity information.
[0014] The steps for the data processing module to unify the date format information and numeric information through format strings and the string replacement method are as follows:
[0015] The borrower information is divided into date format information and numeric information;
[0016] Receive the date format information in the bank information system. Through the format string, the date string in the date format information is grouped and understood according to the definition of the format string. Then, the format string searches for the corresponding year, month, and date parts in the defined order to uniformly process the date format information;
[0017] The example is as follows:
[0018] Assume the date string is "2024 / 01 / 02" and the format string is "%Y / %m / %d". The format string identifies that "2024" corresponds to "%Y" representing the year, "01" corresponds to "%m" representing the month; "02" corresponds to "%d" representing the date, thus completing the mapping between the date string and each part of the date;
[0019] Parse the date format information in the order from left to right according to the format string. For example, for the format string "%Y / %m / %d" and the date string "2024 / 01 / 02", first find the part corresponding to "%Y", that is, the year "2024", and obtain that "2024" in the date format information is the year. After finding the year, then find the corresponding month and date parts in the order of the format string, thus uniformly processing the date format information;
[0020] The steps to unify numerical information through the string replacement method are as follows:
[0021] Establish a replacement set for unnecessary units and characters;
[0022] Compare each character starting from the beginning of the characters in the numerical information with the characters in the replacement set one by one,
[0023] If the character in the numerical information = the character in the replacement set, then remove the character in the numerical information;
[0024] If the character in the numerical information ≠ the character in the replacement set, then retain the character in the numerical information.
[0025] The steps for the data standardization module to standardize the borrower information through the Z-score standardization method are as follows:
[0026] Receive the borrower account information, borrower credit, date format information, and financial information submitted by the borrower in the data processing module, sort them in the order of the date format information, and respectively establish data sets corresponding to the date format information;
[0027] Receive the data set: , where a certain data is , and the calculation formula of Z-score is:
[0028] ;
[0029] Among them, is the standardized Z-score value, is the value of the original data point, n is the subscript used to distinguish the values of different original data points, is the mean of this data set, The calculation formula of
[0030] ;
[0031] is the standard deviation of the data set, representing the degree of dispersion of the data. The calculation formula is as follows:
[0032] .
[0033] The steps for the data standardization module to determine whether there are abnormal values in the standardized data set are as follows:
[0034] Receive the Z-score value, mean value, and corresponding standard deviation after the data set is standardized;
[0035] Abnormal value definition method:
[0036] Set the definition of abnormal values;
[0037] Calculate the mean value At the standard deviation of the degree of dispersion, set the discrimination value to ;
[0038] If of the data points fall within the range of the mean value plus or minus times the standard deviation , then the interval is ;
[0039] If of the data points fall within the range of the mean value plus or minus times the standard deviation , then the interval is ;
[0040] Conversely, data points outside this range are regarded as abnormal values;
[0041] Abnormal value judgment:
[0042] Take the interval ;
[0043] When , the corresponding value is an abnormal value.
[0044] The present invention fully considers the situation that there are differences in the handwriting of different borrowers. Such differences are extremely likely to lead to the problem of confusion between numbers and similar characters. For example, the number "0" may be indistinguishable from the letter "O" due to the writing style, and the number "1" may also be confused with the lowercase letter "l". This kind of confusion frequently appears in the information filled in by borrowers with diverse handwriting styles, and thus will affect the assessment of the borrowers' credit;
[0045] The steps of the data obfuscation analysis module using the similarity comparison method are as follows:
[0046] Receive the data corresponding to the abnormal value judged in the data standardization module and the digital characters replaced in the data processing module, set a similarity threshold based on the learning model, and judge whether the replaced character is similar to the value;
[0047] When they are similar, convert the replaced character into a value, and then judge whether the value is abnormal through the data standardization module. Similarly, it can be judged whether the abnormal value is caused by the similarity between the value and the character;
[0048] Set the similarity threshold based on the learning model The steps are as follows:
[0049] Receive known similar and dissimilar characters and their corresponding values ;
[0050] For similar and dissimilar respectively and the number of pixels with the same position being the same: ;
[0051] Among them, and are respectively and at row and column pixel values, is an indicator function, if the pixel values are the same, it is , otherwise it is ;
[0052] The character shape similarity calculation formula is: ;
[0053] Then, compare the similarity distribution of the corresponding characters and values of similar and dissimilar to determine the similarity threshold ;
[0054] Receive the data corresponding to the abnormal value judged in the data standardization module and the characters replaced corresponding to the data in the data processing module ;
[0055] According to the similarity of the character shape similarity and similarity;
[0056] Receive the set similarity threshold ;
[0057] If ≥ The output will convert the character to be replaced into a numerical value or convert the numerical value into a character;
[0058] If < it indicates that the character replacement of the data processing module is not confused.
[0059] In view of the scenario of manually inputting a large amount of data, the present invention takes into account such a risk: when a staff member enters information into a data table, due to negligence or operational errors, the line break operation may not be performed, which will cause data confusion problems. For example, data that should belong to different rows and represent different meanings may be wrongly connected together, making the original independent borrower information segments mixed with each other. This not only destroys the original structure of the data, but may also lead to serious deviations in subsequent data processing, having a great negative impact on the accurate analysis of borrower information and credit assessment;
[0060] The step for the non-confused data judgment unit to judge whether the non-confused data is adjacent in its corresponding data set based on the index position is as follows:
[0061] The data set of the data standardization module is sorted in the order of date format information. For multiple abnormal data, one abnormal data is selected as the index If the remaining abnormal data are respectively respectively 、 ;
[0062] it indicates that multiple abnormal data are adjacent data in the data set of the data standardization module (200);
[0063] Set a quantity threshold, and the quantity threshold ≥ 2;
[0064] When the number of adjacent abnormal values > the quantity threshold, it indicates that the data in the data set is confused;
[0065] The steps for the re-partitioning unit to re-partition abnormal data through the abnormal value delimitation method in the data standardization module are as follows:
[0066] Receive a data confusion signal, obtain the historical information of the corresponding borrower through the API interface protocol in the data processing module, and unify the historical information respectively through the format string and string replacement method in the data processing module;
[0067] Calculate the delimitation value of the historical data corresponding to the abnormal data through the abnormal value delimitation method of the data standardization module;
[0068] If the abnormal data is within the delimitation value, re-partition the abnormal data, and then partition the abnormal data that has not been re-partitioned in turn;
[0069] The division is completed, and the data is standardized again by the data standardization module and output to the staff for the staff to evaluate the borrower.
[0070] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0071] 1. In the artificial intelligence system for data standardization, a borrower information set is established through the data standardization module, and then the Z-score standardization method is used to standardize the data set. The Z-score standardization method can convert borrower information data of different magnitudes and dimensions into a comparable standard form, enabling the relationships between various features to be more accurately analyzed and measured during credit assessment. Before that, the borrower information is unified by the data processing module to remove redundant characters and unify the date format, which can eliminate data format differences caused by different input methods or sources, and helps reduce errors introduced due to chaotic data formats.
[0072] 2. In the artificial intelligence system for data standardization, after the data set is standardized, the data confusion analysis module uses the similarity comparison method to judge whether the data is removed as a character due to character and data confusion when removing redundant characters, or whether a character is retained as data. If the above situation occurs, a similarity threshold for characters and numbers is set based on the learning model, and the confused data is re-divided, effectively avoiding the situation of accidentally deleting valid data or retaining invalid characters during the process of removing redundant characters, ensuring the accuracy of data cleaning, so that the retained data are all truly valuable data and key information will not be lost due to incorrect cleaning operations.
[0073] 3. In the artificial intelligence system for data standardization, if the data confusion analysis module uses the similarity comparison method to judge that there is no confusion between characters and data, the unconfused data judgment unit judges whether the abnormal data is adjacent data. If it is adjacent data, it means that there may be a possibility of adjacent data mixing at this time, avoiding blind processing of the entire data set and concentrating on solving problems that may be caused by adjacent data mixing. At this time, the re-division unit re-divides the abnormal data through the abnormal value definition method in the data standardization module. Brief Description of the Drawings
[0074] Figure 1 It is the overall module schematic diagram of the present invention.
[0075] The meanings of the various labels in the figure are as follows:
[0076] 100, Data Processing Module; 200, Data Standardization Module; 300, Data Obfuscation Analysis Module; 400, Data Partitioning Module; 410, Unobfuscated Data Judgment Unit; 420, Re-partitioning Unit. Detailed Implementation Manner
[0077] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0078] An artificial intelligence system for data standardization includes a data processing module 100, a data standardization module 200, a data obfuscation analysis module 300, and a data partitioning module 400;
[0079] The data processing module 100 obtains borrower information through the API interface protocol, and unifies the borrower information respectively through the format string and the string replacement method. The API interface follows strict data transmission specifications, which can ensure the accuracy of data during the transmission process, reduce data errors caused by manual intervention or irregular transmission. In the past, for the problem of inconsistent data formats, some practices may only be simple manual visual recognition and manual modification, with extremely low efficiency and it is difficult to ensure the consistency of the unified standard. There are also some that may use relatively complex scripts for processing, but the generality is not strong. Once there are new changes in the data format, a large amount of code needs to be rewritten for adjustment. For example, in the face of borrower names in different formats, some are all in uppercase, some have the first letter in uppercase and the rest in lowercase, and some have spaces in the middle, etc. If modified one by one manually, the workload is huge, and traditional fixed-rule scripts are difficult to handle the changing actual formats.
[0080] The data processing module 100 is used to send a data request and an API key to the bank information system. The data request includes borrower account information, borrower credit report, date format information, and financial information submitted by the borrower himself. After receiving the request, the bank information system compares the API key with the legal key. If the API key = legal key, it indicates that the verification is successful. At this time, the bank information system retrieves the borrower account information, borrower credit report, date format information, and financial information submitted by the borrower himself according to the borrower identity information.
[0081] The steps for the data processing module 100 to unify the date format information and numerical information through the format string and the string replacement method are as follows:
[0082] The borrower information is divided into date format information and numerical information;
[0083] Receive the date format information in the bank information system. Group and understand the date string in the date format information according to the definition of the format string. Then, the format string searches for the corresponding year, month, and date parts in the defined order to uniformly process the date format information;
[0084] The example is as follows:
[0085] Suppose the date string is "2024 / 01 / 02" and the format string is "%Y / %m / %d". The format string recognizes that "2024" corresponds to "%Y" and represents the year, "01" corresponds to "%m" and represents the month; "02" corresponds to "%d" and represents the date, thus completing the mapping of the date string to each part of the date;
[0086] Parse the date format information in the order from left to right of the format string. For example, for the format string "%Y / %m / %d" and the date string "2024 / 01 / 02", first find the part corresponding to "%Y", that is, the year "2024", and obtain that "2024" in the date format information is the year. After finding the year, then find the corresponding month and date parts in the order of the format string, so as to uniformly process the date format information;
[0087] The steps to unify numerical information through the string replacement method are as follows:
[0088] Establish a replacement set for unnecessary units and characters;
[0089] Compare each character starting from the beginning of the characters in the numerical information with the characters in the replacement set one by one,
[0090] If the character in the numerical information = the character in the replacement set, then remove the character in the numerical information;
[0091] If the character in the numerical information ≠ the character in the replacement set, then retain the character in the numerical information.
[0092] The data standardization module 200 is used to establish a data set corresponding to the borrower information, perform standardization processing on the data set through the Z-score standardization method, use the abnormal value definition method to judge the abnormal values in the data set, and the format string provides a template for formatting data, which can flexibly handle various formats of borrower information. By defining a good format template, such as stipulating that the format of the name is "the first letter is capitalized, the rest are in lowercase, separated by spaces in the middle", it can adapt to name data in different initial formats for unified conversion. The string replacement method can perform precise replacement for some specific characters, symbols, etc., such as uniformly replacing " / " in the date format with "-", etc. The operation is simple and can quickly adapt to the subtle differences of different data, with strong versatility, and there is no need to write complex customized code for each specific situation.
[0093] Regardless of the original range and magnitude of the borrower information data, the Z-score standardization can effectively process it to meet the general standardization requirements. It does not depend on specific data range settings or subjective empirical values. As long as the data meets certain conditions for calculating the mean and standard deviation, which are usually met by numerical data, it can play a good role and is widely applicable to various scenarios involving borrower information analysis, modeling, etc., and helps with the integration and comparison between different data sets;
[0094] The steps for the data standardization module 200 to standardize the borrower information through the Z-score standardization method are as follows:
[0095] Receive the borrower account information, borrower credit, date format information, and financial information submitted by the borrower in the data processing module 100, sort them in the order of the date format information, and establish data sets corresponding to the date format information respectively;
[0096] Receive the data set: , where a certain data is , and the calculation formula of Z-score is:
[0097] ;
[0098] Among them, is the standardized Z-score value, is the value of the original data point, is the mean of this data set, The calculation formula of
[0099] ;
[0100] is the standard deviation of this data set, indicating the degree of dispersion of the data. The calculation formula is as follows:
[0101] 。
[0102] The steps for the data standardization module 200 to determine whether there are abnormal values in the standardized data set are as follows:
[0103] Receive the Z-score value, mean, and corresponding standard deviation after the data set is standardized;
[0104] Method for defining abnormal values:
[0105] Set the definition of abnormal values;
[0106] Calculate the mean At the standard deviation Of the degree of dispersion, set the discrimination value to ;
[0107] If Of the data points fall within the mean Plus or minus Times the standard deviation Range, then the interval is ;
[0108] If Of the data points fall within the mean Plus or minus Times the standard deviation Range, then the interval is ;
[0109] Conversely, data points outside this range are regarded as abnormal values;
[0110] Judgment of abnormal values:
[0111] Take the interval ;
[0112] When , the corresponding Value is an abnormal value.
[0113] Traditional data anomaly handling methods often focus on simple corrections of the abnormal values themselves or directly marking and removing them, etc., and rarely delve deeply into tracing the specific reasons for the anomalies, especially analyzing whether they are caused by specific data processing steps such as the string replacement method;
[0114] For example, after finding that the borrower's income value is abnormal, it only simply judges that it exceeds the conventionally set range, and then either manually adjusts the value or corrects it according to fixed rules, without exploring whether data confusion was caused by string replacement operations in the previous preprocessing steps such as data format unification, resulting in the inability to solve the problem at the root, and similar anomalies may occur repeatedly in the future;
[0115] The data confusion analysis module 300 uses the similarity comparison method to determine whether the abnormal value is caused by the string replacement method in the data processing module 100. If it is caused by the string replacement method, it means that the string replacement method confuses characters and numbers at this time, and the abnormal value will be replaced again. If it is not caused by the string replacement method, it means that the abnormal value in the data standardization module 200 is normal. By using the similarity comparison method to specifically determine whether the abnormal value is caused by the string replacement method, the source of the abnormality can be accurately traced;
[0116] For example, if the phone number field in the borrower information has an abnormal format, such as characters being mixed into a originally pure - number phone number, by comparing the data characteristics and formats before and after replacement using the similarity comparison method, it can be determined whether, during the operation of the string replacement method, due to the mis - replacement of some similar characters, such as accidentally replacing the number "0" with the letter "O", an abnormality has occurred. This kind of confusion frequently appears in the information filled in by borrowers with diverse handwriting, which will then affect the assessment of the borrower's credit;
[0117] The steps for the data confusion analysis module 300 to use the similarity comparison method are as follows:
[0118] Receive the data corresponding to the abnormal value judged in the data standardization module 200 and the digital characters replaced in the data processing module 100, set a similarity threshold based on the learning model, and determine whether the replaced characters are similar to the values;
[0119] When they are similar, convert the replaced characters into values, and then use the data standardization module 200 to determine whether the values are abnormal. Similarly, it can be determined whether the abnormal value is caused by the similarity between the value and the character resulting in a numerical abnormality;
[0120] Set the similarity threshold based on the learning model The steps are as follows:
[0121] Receive known similar and dissimilar characters and their corresponding values ;
[0122] For similar and dissimilar ones respectively and count the number of pixels with the same value at the same position: ;
[0123] Among them, and are respectively and the pixel values at the row and the column, Is an indicator function, and is , otherwise it is ;
[0124] The character shape similarity calculation formula is: ;
[0125] Then, compare the similarity and dissimilarity of the corresponding characters and the numerical similarity distribution to determine the similarity threshold ;
[0126] Receive the data corresponding to the abnormal value in the data normalization module 200 And the characters corresponding to the data replaced in the data processing module 100 ;
[0127] According to the character shape similarity And Similarity;
[0128] Receive the set similarity threshold ;
[0129] If ≥ , then output to convert the character to be replaced into a numerical value or replace the numerical value with a character;
[0130] If < , it means that the character replacement in the data processing module 100 is not confused.
[0131] The present invention takes into account that in the scenario of manually inputting a large amount of data, there is such a risk: when the staff enters information into the data table, they may, due to negligence or operational errors, fail to perform a line break operation, which will cause data confusion problems. For example, data that should belong to different rows and represent different meanings are wrongly connected together, making the original independent borrower information fragments mixed with each other, not only destroying the original structure of the data, but also possibly causing serious deviations in subsequent data processing links, having a great negative impact on the accurate analysis of borrower information and credit assessment and other work;
[0132] The non-confused data judgment unit 410 judges based on the index position. The steps for whether the non-confused data is adjacent in its corresponding data set are as follows:
[0133] Sort the data set of the data normalization module 200 in the order of date format information. For multiple abnormal data, select one abnormal data as the index , if the remaining abnormal data are respectively Are respectively , ;
[0134] It indicates that multiple abnormal data in the data set of the data standardization module (200) are adjacent data;
[0135] Set a quantity threshold, where the quantity threshold ≥ 2;
[0136] When the quantity of adjacent abnormal values > the quantity threshold, it indicates data confusion in the data set;
[0137] By judging whether abnormal values are adjacent in the data set based on the index position and considering the quantity situation of adjacent abnormal values, it is possible to discover abnormalities from the perspective of the overall layout and relevance of the data;
[0138] For example, in a data set related to the credit scores of borrowers, if a certain indicator of multiple consecutive borrowers, such as the debt ratio, shows abnormalities, and these abnormal values are arranged adjacent to each other in the data set, it may imply the existence of some systematic data confusion or error source. Perhaps a data collection device failure, an error in a specific business process link, etc. have caused problems with this segment of data. This type of analysis from the perspective of position correlation helps to more deeply dig out the underlying problems hidden behind the data, avoiding simply dealing with individual abnormal values on the surface and enhancing the comprehensiveness of data quality control;
[0139] The steps for the re - partitioning unit 420 to re - partition abnormal data through the abnormal value definition method in the data standardization module 200 are as follows:
[0140] Receive the data confusion signal, obtain the historical information of the corresponding borrower through the API interface protocol in the data processing module 100, and unify the historical information respectively through the format string and string replacement method in the data processing module 100;
[0141] Calculate the definition value of the historical data corresponding to the abnormal data through the abnormal value definition method of the data standardization module 200;
[0142] If the abnormal data is within the definition value, re - partition the abnormal data, and then sequentially partition the abnormal data that has not been re - partitioned;
[0143] After the partitioning is completed, re - standardize the data through the data standardization module 200 and output it to the staff for the staff to evaluate the borrower;
[0144] When receiving the data confusion signal, by using the existing abnormal value definition method in the data standardization module 200 to re - partition the confused data, it is possible to trace back to the link where the data has problems, sort out the confused data according to a scientific and reasonable definition standard, re - distinguish it according to reasonable logic and business rules, find out the data part that was originally mis - partitioned or interfered with, and achieve precise repair;
[0145] For example, regarding the data related to the borrower's repayment ability, if the data such as income and expenditure are confused due to certain external factors, the abnormal value definition method can be used to re-determine the reasonable value ranges of each part, correctly divide them, restore the accuracy and usability of the data, and avoid the information loss caused by simply discarding the data.
[0146] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence system for data standardization, characterized in that, It includes a data processing module (100), a data standardization module (200), a data confusion analysis module (300), and a data partitioning module (400); The data processing module (100) obtains borrower information through the API interface protocol, and unifies the borrower information respectively by the format string and the string replacement method. The data standardization module (200) is used to establish a data set corresponding to the borrower information, perform standardization processing on the data set by the Z-score standardization method, and use the abnormal value definition method to judge the abnormal values in the data set; The data confusion analysis module (300) uses the similarity comparison method to judge whether the abnormal value is caused by the string replacement method in the data processing module (100). If it is caused by the string replacement method, it means that the string replacement method confuses characters and numbers at this time, and the abnormal value will be replaced again. If it is not caused by the string replacement method, it means that the abnormal value in the data standardization module (200) is normal; The steps of the data confusion analysis module (300) using the similarity comparison method are as follows: Receive the data corresponding to the abnormal value judged in the data standardization module (200) and the digital characters replaced in the data processing module (100), set a similarity threshold based on the learning model, and judge whether the replaced characters are similar to the values; When they are similar, convert the replaced characters into values, and then judge whether the values are abnormal through the data standardization module (200). Similarly, it can be judged whether the abnormal value is caused by the similarity between the value and the character; The steps of setting the similarity threshold k based on the learning model are as follows: Receive the known similar and dissimilar characters A and their corresponding values B; The number of pixels at the same positions of A and B that are the same in respectively similar and dissimilar cases: where A ij and B ij are the pixel values of A and B at the i-th row and j-th column respectively, and I is an indicator function which is 1 if the pixel values are the same and 0 otherwise; The calculation formula for the similarity of character shapes is as follows: Compare the similarity distribution of the corresponding characters and values of similarity and dissimilarity to determine the similarity threshold k; Receive the data A1 corresponding to the abnormal value judged in the data standardization module (200) and the characters B1 replaced corresponding to the data in the data processing module (100); According to the similarity of the character shape similarities of A1 and B1; Receive the set similarity threshold k; If Similarity shape ≥ k, then output the character to be replaced converted to a numerical value or the numerical value replaced with a character; If Similarity shape < k, it indicates that the replacement characters of the data processing module (100) are not confused; The data partitioning module (400) includes an unconfused data judgment unit (410) and a re-partitioning unit (420); The unconfused data judgment unit (410) is used to receive the number of abnormal values judged by the data confusion analysis module (300) that are not caused by the string replacement method, and judge whether the abnormal values are adjacent in the data set based on the index position, and set a quantity threshold. When the number of adjacent abnormal values > the quantity threshold, it means that the data in the data set is confused; The re-partitioning unit (420) is used to receive the data confusion signal, re-partition the confused data by the abnormal value definition method in the data standardization module (200), and after re-partitioning, perform data standardization processing again and output the standardized data to the staff.
2. The artificial intelligence system for data standardization according to claim 1, characterized in that: The data processing module (100) is used to send a data request and an API key to the bank information system. The data request includes borrower account information, borrower credit reports, date format information, and financial information submitted by the borrower himself. After receiving the request, the bank information system compares the API key with the legitimate key. If the API key = legitimate key, it indicates successful verification. At this time, the bank information system retrieves the borrower account information, borrower credit reports, date format information, and financial information submitted by the borrower himself according to the borrower identity information.
3. The artificial intelligence system for data standardization according to claim 1, characterized in that: The steps for the data processing module (100) to unify the date format information and numerical information through format strings and string replacement methods are as follows: Borrower information is divided into date format information and numerical information; Receive the date format information in the bank information system, and use format strings to format the date strings in the date format information; Group and understand according to the definition of format strings; The format string searches for the corresponding year, month, and date parts in the defined order; Unify the processing of date format information; The steps for unifying numerical information through string replacement methods are as follows: Establish a replacement set for unnecessary units and characters; Compare each character starting from the beginning of the characters in the numerical information with the characters in the replacement set one by one, If the character in the numerical information = the character in the replacement set, then remove the character in the numerical information; If the character in the numerical information ≠ the character in the replacement set, then retain the character in the numerical information.
4. The artificial intelligence system for data standardization according to claim 1, wherein: The steps for the data standardization module (200) to standardize borrower information through the Z-score standardization method are as follows: Receive the borrower account information, borrower credit, date format information, and financial information submitted by the borrower himself in the data processing module (100), sort them in the order of date format information, and establish data sets corresponding to the date format information respectively; Received data set: X = {x1, x2, … x n}, where a certain data is x, and the calculation formula for Z-score is: Among them, z is the standardized Z-score value, x is the value of the original data point, μ is the mean of this data set, and the calculation formula of μ is as follows: σ is the standard deviation of this data set, indicating the degree of dispersion of the data, and the calculation formula is as follows:
5. The artificial intelligence system for data standardization according to claim 1, characterized in that: The steps for the data standardization module (200) to determine whether there are abnormal values in the standardized data set are as follows: Receive the Z-score value, mean, and corresponding standard deviation after standardizing the data set; Abnormal value definition method: Set the definition of abnormal values; Calculate the degree of dispersion of the mean μ within the standard deviation σ, and set the discrimination value to 90%; If 90% of the data points fall within the range of the mean μ plus or minus 3 times the standard deviation σ, the interval is (μ - 3σ, μ + 3σ), If 90% of the data points fall within the range of the mean μ plus or minus 2 times the standard deviation σ, the interval is (μ - 2σ, μ + 2σ), Conversely, data points outside this range are regarded as abnormal values; Abnormal value judgment: Take the interval (μ - 3σ, μ + 3σ); When |z| > 3, the corresponding x value is an abnormal value.
6. The artificial intelligence system for data standardization according to claim 1, wherein: The steps for the unconfused data judgment unit (410) to judge whether abnormal values are adjacent in the corresponding data set based on the index position are as follows: The data normalization module (200) sorts the data set in the order of date format information; For multiple abnormal data, select one abnormal data as the index h; If the remaining abnormal data are h+1… and h-1… respectively with h; It indicates that multiple abnormal data in the data set of the data normalization module (200) are adjacent data; Set a quantity threshold, where the quantity threshold ≥ 2; When the quantity of adjacent abnormal values > the quantity threshold, it indicates data confusion in the data set.
7. The artificial intelligence system for data standardization according to claim 1, characterized in that: The steps for the re-partitioning unit (420) to re-partition abnormal data through the abnormal value definition method in the data normalization module (200) are as follows: Receive a data confusion signal, obtain the historical information of the corresponding borrower through the API interface protocol in the data processing module (100), and unify the historical information respectively through the format string and string replacement method in the data processing module (100); Calculate the defined value of the historical data corresponding to the abnormal data through the abnormal value definition method of the data normalization module (200); If the abnormal data is within the defined value, re-partition the abnormal data, and then partition the un-re-partitioned abnormal data in sequence; After the partitioning is completed, re-normalize the data through the data normalization module (200) and output it to the staff for the staff to evaluate the borrower.
Citation Information
Patent Citations
Data analysis system based on intelligent water affair management platform
CN118193734A
Express delivery channel safety supervision system based on end-side cloud architecture
CN118822252A