Abnormality processing method and device, storage medium, electronic equipment and product
By using a multi-dimensional anomaly detection model and comprehensive calculation methods, the problem of insufficient coverage and misjudgment in the anomaly detection of credit data in existing technologies has been solved, and accurate quantitative evaluation and high-quality credit assessment of credit data have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QIANTANG CREDIT INFORMATION CO LTD
- Filing Date
- 2026-04-09
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods for detecting anomalies in credit data are insufficient to effectively cover different forms of anomalies, leading to missed detections and misjudgments, and failing to meet the stringent data quality requirements of financial institutions.
By acquiring credit reporting data and using anomaly detection models for multi-dimensional detection, including logical contradiction detection, time series detection, user behavior detection, and group attribute detection, and combining the target weights of each detection dimension for comprehensive calculation, we can achieve comprehensive coverage and accurate quantitative assessment of different types of anomalies.
It improves the accuracy and reliability of credit data anomaly detection, meets the stringent data quality requirements of financial institutions, and ensures the accuracy and reliability of credit assessment.
Smart Images

Figure CN122022992A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to one or more embodiments in the field of computer technology, and more particularly to an exception handling method, apparatus, storage medium, electronic device and product. Background Technology
[0002] With the rapid development of credit reporting services, the requirements for data credibility in credit assessment have significantly increased. Accurate and reliable credit data has become an important foundation for financial institutions to conduct lending business and prevent credit risks. Credit data typically comes from a wide range of sources and is of complex and diverse types, making anomaly detection of data provided by data source institutions a core aspect of ensuring the reliability of credit assessments for credit reporting agencies.
[0003] However, due to the complexity of credit data, the forms of anomalies are also hidden and diverse, making it difficult for existing anomaly detection methods to effectively cover different forms of anomalies, and even causing missed detections and misjudgments of anomalies. This results in low accuracy of anomaly detection and makes it difficult to meet the stringent data quality requirements of credit reporting business. Summary of the Invention
[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, an exception handling method is proposed, applied to a credit reporting system corresponding to a credit reporting agency, comprising: Obtain the data to be tested provided by the data provider for credit reporting services; The data to be detected is input into a preset anomaly detection model to obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs. The target weights corresponding to each detection dimension are determined, and the anomaly detection results of each detection dimension are comprehensively calculated based on the target weights to obtain the comprehensive detection results for the data to be detected. Anomaly handling is then performed based on the comprehensive detection results.
[0005] According to a second aspect of one or more embodiments of this specification, an exception handling apparatus is provided, comprising: The acquisition module is used to acquire the data to be tested provided by the data provider; The detection module is used to input the data to be detected into a preset anomaly detection model and obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs. The processing module is used to determine the target weights corresponding to each detection dimension, and to perform comprehensive calculations on the abnormal detection results of each detection dimension based on the target weights to obtain a comprehensive detection result for the data to be detected, and to perform abnormal processing based on the comprehensive detection result.
[0006] According to a third aspect of one or more embodiments of this specification, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method described above by executing the executable instructions.
[0007] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method described above.
[0008] According to a fifth aspect of one or more embodiments of this specification, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method described above.
[0009] As can be seen from the above embodiments, after obtaining the data to be detected, this specification can use an anomaly detection model to perform anomaly detection on it in multiple dimensions. Among them, logical contradiction detection can accurately identify explicit contradictions in the logic of data fields, outlier detection can capture data mutations in the time dimension, and abnormal user behavior detection can accurately detect user behaviors that do not conform to normal behavioral logic. Then, based on the target weights corresponding to each detection dimension, the anomaly detection results of each detection dimension are comprehensively calculated, so as to achieve comprehensive coverage and accurate quantitative evaluation of different types of anomalies, thereby improving the accuracy and reliability of credit data anomaly detection and fully meeting the stringent requirements of credit reporting business for data quality. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the architecture of an exception handling service system provided in an exemplary embodiment; Figure 2 This is a schematic diagram of the architecture of a credit scoring system provided in an exemplary embodiment; Figure 3 This is a flowchart of an exception handling method provided in an exemplary embodiment; Figure 4 This is an exemplary embodiment of an overall flowchart of exception handling; Figure 5 This is a schematic diagram of the structure of a device provided in an exemplary embodiment; Figure 6 This is a block diagram of an exception handling apparatus provided in an exemplary embodiment. Detailed Implementation
[0011] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0012] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0013] With the rapid expansion of the credit reporting industry, financial institutions are placing increasingly stringent requirements on the credibility of credit data when conducting lending business and controlling credit risk. Accurate and reliable credit data has long been a crucial support for such work. However, the sources of credit data are already very broad, covering numerous financial institutions, public utilities, and internet platforms, and the data formats and quality vary considerably. This makes anomaly detection of data sources a core element for credit reporting agencies to ensure the reliability of credit assessments.
[0014] However, due to the complexity of credit data, current anomaly detection methods have significant shortcomings: First, traditional rules are insufficient in covering complex anomalies, which can easily lead to false positives and false negatives; second, existing detection methods can only identify surface-level explicit errors and are difficult to capture deep logical contradictions.
[0015] Specifically, taking the static threshold rule "query count exceeding 10 times is considered abnormal" as an example, although this rule can perform preliminary screening for explicit operations such as high-frequency queries, it often fails to identify malicious forgery within a reasonable range (such as forging 6 consecutive months of small-amount repayments); for another example, while existing methods can accurately identify explicit errors such as "age entered as 200", they lack effective identification capabilities when faced with logical contradictions that defy common sense, such as "a 22-year-old fresh graduate with a monthly mortgage payment of 500,000".
[0016] Based on this, this specification provides an anomaly handling method. By using an anomaly detection model to perform anomaly detection on the data to be detected in multiple dimensions, and then comprehensively calculating the anomaly detection results of each detection dimension according to the target weights corresponding to each detection dimension, it is possible to achieve comprehensive coverage and accurate quantitative evaluation of different types of anomalies.
[0017] The technical solutions described in the various embodiments of this specification will be explained in detail below with reference to the accompanying drawings.
[0018] Figure 1 This is a schematic diagram of the architecture of an exception handling service system provided in an exemplary embodiment. For example... Figure 1 As shown, the system may include a server 11, a network 12, and several electronic devices, such as a personal computer (PC) 13, a mobile phone 14, etc.
[0019] Server 11 can be a physical server containing an independent host, or it can be a virtual server hosted in a host cluster. During operation, server 11 can run server-side programs for a certain application to implement the relevant functions of that application. For example, when server 11 runs an exception handling service program, it can be implemented as a corresponding exception handling service platform.
[0020] PC13 and mobile phone14 are just some of the types of electronic devices that users can use. In reality, users can obviously also use electronic devices such as tablets, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smartwatches, etc.), etc., and one or more embodiments in this specification do not limit this. During operation, the electronic device can run a client-side program of an application to implement the relevant functions of that application. For example, when the electronic device runs an exception handling service program, it can act as a client for that exception handling service. The client application of the aforementioned exception handling service can be launched and run on the electronic device. This client-side program can be a native application installed on the electronic device, or it can be a mini-program, quick app, or other similar form. Of course, when using web technologies such as HTML5 or similar, the relevant functions can be implemented through a page displayed by a browser. This browser can be a standalone browser application or a browser module embedded in some applications.
[0021] As for the network 12 that enables interaction between electronic devices such as PC13 and mobile phone 14 and server 11, communication can be achieved using either wired or wireless networks, depending on the communication methods supported by the respective electronic devices. This specification does not impose any restrictions on this. For example, PC13 can support both wired and wireless communication, so it can use either wired or wireless networks as needed. Mobile phone 14 typically only supports wireless communication, so it can use a wireless network for communication.
[0022] In the financial services sector, financial institutions typically need to leverage credit reporting agencies to obtain multi-dimensional credit-related data in order to achieve accurate risk assessment and business decision-making. Specifically, the information flow path is as follows: financial institutions, as data requesters, initiate data query requests for target users to credit reporting agencies; the credit reporting agency, as an intermediary, receives the request and, based on its own business system, connects with multiple data source institutions to obtain various credit-related business data such as the user's credit records, repayment behavior, and repayment history. The credit reporting agency can then use this data to conduct risk assessments of the user and feed the risk assessment results back to the financial institution, providing decision support for the financial institution's credit approval, risk control, and other business operations.
[0023] Figure 2 This is a schematic diagram of the architecture of a credit reporting system provided in an exemplary embodiment.
[0024] like Figure 2 As shown, the credit reporting service system includes data requesters, the credit reporting system itself, and data providers.
[0025] Among them, the data demanders can be the business processing systems of financial institutions such as banks, securities companies, and insurance companies, which are used to provide financial services such as deposits and withdrawals, transfers and remittances, credit issuance, securities trading, insurance application and claims.
[0026] A credit reporting system can be a specialized platform operated by credit reporting agencies to integrate, verify, analyze, and process relevant query requests and multi-channel credit data (such as credit records and default information) to form standardized credit assessment results.
[0027] Data providers can be data platforms of data source institutions such as e-commerce platforms, social networks, telecommunications operators, and government departments (i.e., data providers), used to store cross-domain data such as user consumption behavior records, social interaction data, communication trajectories and locations, enterprise business information, and tax records.
[0028] Because anomaly detection relies on multi-channel data integration capabilities, standardized verification systems, and professional analytical computing power, and because the credit reporting system, as a core node in data flow, directly connects data providers and requesters, it can control the entire data collection, processing, and evaluation process. Therefore, in this specification, the executing entity for the anomaly detection method can be the credit reporting system corresponding to the credit reporting agency, specifically referring to the servers, cloud computing platforms, distributed nodes, and terminal devices of the credit reporting business processing system. For ease of description, the following will use the credit reporting system as the executing entity to illustrate the graph neural network-based anomaly detection method provided in this specification.
[0029] Figure 3 This is a flowchart illustrating an exception handling method provided in an exemplary embodiment, including the following steps: S300: Obtain the data to be tested provided by the data provider.
[0030] In this specification, the credit reporting system may obtain the data to be tested at various times, including: When a data provider, after passing the qualification review, connects to the credit reporting system for the first time, the anomaly detection process is triggered, and the first batch of data to be detected is synchronized from the data source of that institution. When a data provider that has been connected to the credit reporting system undergoes critical changes such as data updates or interface iterations, the anomaly detection process is triggered, and the updated data to be detected is extracted from the corresponding data source. The credit reporting system can periodically perform data synchronization tasks according to a preset cycle. When the set time node is reached, the anomaly detection process is triggered, and anomaly detection requests are sent to various data providers to obtain the data to be detected.
[0031] The data to be detected may include text data and image data. For text data, it contains structured fields (such as loan amount and number of overdue payments) and semi-structured logs (such as query records and authorization information). For image data (such as scanned copies of user ID cards, images of credit contracts, and copies of asset certificates), the text information can be extracted through technologies such as Optical Character Recognition (OCR) and Natural Language Processing (NLP) during the anomaly handling process.
[0032] S302: Input the data to be detected into a preset anomaly detection model to obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs.
[0033] After obtaining the data to be detected, the credit scoring system can further construct prompt words based on the data to be detected and the preset detection conditions corresponding to each detection dimension.
[0034] For example, a credit scoring system can split the data to be detected by field type and embed it into a preset prompt word template. The prompt word template has been pre-configured with the detection conditions corresponding to each detection dimension, thereby obtaining structured prompt words that are adapted to the anomaly detection model, providing a clear analytical basis for subsequent anomaly identification.
[0035] Then, a prompt word containing the data to be detected can be input into a preset anomaly detection model, so that the anomaly detection model can output anomaly detection results in multiple detection dimensions under the prompt word.
[0036] In this specification, the anomaly detection model can be a trained Large Language Model (LLM), and the detection dimensions for anomaly detection can include: Perform logical contradiction checks on the fields contained in the data to be tested: The client can use an anomaly detection model to identify and extract key credit reporting fields related to credit reporting business from the data to be detected. Then, the extracted key credit reporting fields are combined to construct a structured semantic description text. Based on this semantic description text, the client can detect whether there are logical contradictions between the fields, thereby achieving cross-field detection of the data to be detected.
[0037] For example, the semantic description text constructed based on the extracted fields is: "22-year-old college student, monthly income of 80,000 yuan, 3 houses under his name, and social security paid continuously for 2 years". Based on this semantic description text, it can be concluded that the combination of information is unreasonable (college students usually do not have stable high-salary income and have not paid social security, which is obviously contrary to the asset status of multiple properties), so it is found that there is a significant logical contradiction.
[0038] In addition, during the process of logical contradiction detection by the anomaly detection model, besides utilizing the knowledge learned during training, reference data related to the target user group (such as industry average income, income reports for each age group, typical logical contradiction cases, etc.) can be retrieved from a pre-set information database. Then, the prompt words are reinforced based on the benchmark reference data, thereby using the retrieved reference data as a reference to further improve the accuracy of logical contradiction judgment.
[0039] The output of this detection dimension may include: an anomaly score determined based on the degree to which the logical contradiction deviates from the common sense benchmark and / or the degree of correlation with typical logical contradiction cases, as well as the reason for obtaining the score (e.g., there is a logical contradiction between xxx and xxx).
[0040] Among them, the score is positively correlated with the degree of deviation from common sense, and also positively correlated with the degree of correlation with typical logical contradiction cases.
[0041] Outlier detection is performed on the data to be detected based on the time data series corresponding to the data to be detected: The client can use an anomaly detection model to extract the correlation features between credit indicators of the same user at different time points in the data to be detected. Then, based on the extracted correlation features and timestamp information, a time data sequence is constructed. Subsequently, time-series pattern recognition and anomaly mining are performed on the time data sequence to identify outliers (such as deviations corresponding to a sudden drop in repayment frequency, deviations corresponding to a cliff-like increase in debt ratio, etc.).
[0042] In practical applications, there are various methods for identifying time-series anomalies. For example, a fitting function can be used to fit the time series data, thereby capturing the normal trend of data change and locating abnormal fluctuations that deviate from the trend. Another example is that for each data point in the time series data, the value of the data point can be predicted based on the previous data values, and the deviation between the predicted value and the actual value of the data point can be used to determine whether the data point is an outlier. Furthermore, time-series anomaly identification can also introduce a sliding window algorithm, directly identifying abrupt changes in the data series based on the statistical characteristics (such as mean and variance) of the data within the window.
[0043] The output results corresponding to this detection dimension may include: the anomaly score obtained based on time series feature analysis and deviation calculation, and the reason for obtaining the score (such as outlier data that does not conform to the data distribution pattern at xxx).
[0044] Detecting abnormal user behavior based on user behavior characteristics in the data to be detected: The client can use an anomaly detection model to identify user behavior data in the data to be detected, and determine the user's behavior characteristics in the user's corresponding personal information domain. If the deviation between the current behavior characteristics reflected by the behavior data and the user behavior characteristics is greater than a preset deviation, then it is determined that there is abnormal user behavior in the data to be detected.
[0045] Taking a user's credit report query behavior data as an example, if it is determined from this behavior data that the number of queries for the user in previous months has remained stable in the range of 1-2 times, and suddenly the number of queries in a certain month surges to more than 40 times, then it can be determined that the user's query behavior is obviously abnormal.
[0046] The output results corresponding to this detection dimension may include: the anomaly score determined by the quantitative analysis of the deviation between the user's current behavior and historical behavior characteristics, and the reason for obtaining the score (such as the number of credit inquiry in xxx month not conforming to historical patterns).
[0047] Detecting abnormal user behavior based on the group attribute characteristics of the user's user group: The client can use an anomaly detection model to identify user behavior data within the data to be detected, and to determine the group attribute characteristics of the user's user group. If, based on the group attribute characteristics, the behavior data does not conform to the normal behavior characteristics of its user group, then abnormal user behavior is determined to exist in the data to be detected.
[0048] The aforementioned user groups can be determined based on one or more user information such as gender, age, and occupation. Their group attribute characteristics are used to characterize the common patterns and statistical characteristics of the group in business. Taking credit reporting business as an example, these group attribute characteristics are mainly used to characterize users' income level, debt structure ratio, credit application frequency, repayment performance rate, etc. in dimensions such as credit reporting, consumption, and credit.
[0049] The aforementioned normal behavioral characteristics are used to characterize the behavioral benchmarks and reasonable ranges exhibited by the vast majority of users within this group in the corresponding business.
[0050] For example, if a user's user group is determined to be "programmers aged 25-30," whose group characteristics are: "99% monthly income > 10,000 and < 50,000," then it can be concluded that the normal behavioral characteristics for personal income reporting should be within the reporting range of 10,000-50,000. If the user's reported monthly income is "user reported monthly income of 300,000," then it can be determined that the user's behavioral data does not conform to the normal behavioral characteristics of their user group.
[0051] By leveraging group attribute characteristics, we can quickly define the behavioral baseline range of target users, thereby establishing differentiated anomaly detection thresholds and improving the accuracy and efficiency of cross-group user behavior anomaly detection.
[0052] The output of this detection dimension may include: the abnormality score determined by the quantitative calculation of the deviation between the user's behavior and the normal behavior characteristics corresponding to the attribute characteristics of the group to which the user belongs, and the reason for obtaining the score (such as the amount of the claim involved in the user's xxx behavior is inconsistent with the normal income of the xxx group to which the user belongs).
[0053] The client can select any two or more of the above detection dimensions to combine and obtain the anomaly detection results corresponding to each detection dimension.
[0054] S304: Determine the target weights corresponding to each detection dimension, and perform comprehensive calculations on the abnormal detection results of each detection dimension based on the target weights to obtain comprehensive detection results for the data to be detected, and perform anomaly handling based on the comprehensive detection results.
[0055] In this specification, for each detection dimension, the target weight corresponding to that detection dimension can be a dynamic weight determined based on its anomaly detection results. Specifically, the base weight pre-set for that detection dimension can be dynamically updated based on the anomaly score contained in the anomaly detection results of that detection dimension, thereby obtaining the target weight corresponding to that detection dimension. In this way, the weight ratio of high anomaly contribution dimensions in the overall judgment can be strengthened, and the interference effect of low anomaly contribution dimensions can be weakened, thereby improving the accuracy and reliability of the overall anomaly judgment results.
[0056] Specifically, if the anomaly score of the data to be detected in this detection dimension is greater than the first preset score, it indicates that the abnormal features in this dimension are significant and have high reference value for the overall anomaly judgment. Therefore, the basic weight corresponding to this detection dimension can be increased to obtain the target weight. If the anomaly score of the data to be detected in this detection dimension is less than the second preset score, it indicates that there is no obvious anomaly in this dimension and its reference value for overall anomaly judgment is low. Therefore, the basic weight corresponding to this detection dimension can be reduced to obtain the target weight. If the anomaly score of the data to be detected in this detection dimension is not greater than the first preset score and not less than the second preset score, it indicates that the detection result in this dimension is in the normal fluctuation range and there is no need to adjust the weight. Therefore, the basic weight can be determined as the target weight. Among them, the first preset score is greater than the second preset score. When the anomaly score is greater than the first preset score or less than the second preset score, the target weight is positively correlated with the anomaly score.
[0057] In this specification, the first and second preset scores can be set according to actual conditions. For example, if the total score ranges from 0 to 100, the second preset score can be 10 and the first preset score can be 90. For any detection dimension, if the anomaly score is > 90, the base weight of that dimension is increased, and the higher the score, the greater the increase in weight. If the anomaly score is < 10, the base weight of that dimension is decreased, and the lower the score, the greater the decrease in weight. If 10 ≤ anomaly score ≤ 90, the base weight of that dimension remains unchanged.
[0058] This allows the weight allocation of each detection dimension to be precisely matched with the actual degree of anomaly, thereby improving the objectivity and reliability of the comprehensive anomaly judgment results.
[0059] Of course, in practical applications, a fixed weight can be set for each testing dimension based on past testing experience and the requirements of existing testing standards for risk control.
[0060] Furthermore, the client can classify the risk level of the detection results based on the final determined anomaly score, thereby triggering the corresponding risk handling strategy.
[0061] For example, low risk (<30 points): only record, for continuous model learning; medium risk (30-70 points): generate an alarm ticket and push it to data governance personnel for review; high risk (>70 points): automatically block the data and suspend subsequent reporting from the corresponding data source.
[0062] In practical applications, the anomaly handling strategies may also include: suspending the data access permissions of the problematic organization, triggering historical data resampling and cleaning, updating the quality inspection rule base, and preventing the recurrence of similar problems. The anomaly score corresponding to different handling strategies can be set according to the actual situation, and this manual does not make specific limitations on this.
[0063] In addition, to improve the standardization and timeliness of anomaly handling, standardized anomaly alarm information can be automatically generated.
[0064] Specifically, the client can determine the anomaly handling strategy that matches the comprehensive detection results, and identify reference alarm information (such as high-risk credit anomaly handling guidelines and group behavior deviation alarm templates) that match the comprehensive detection results from the preset alarm information database. Then, based on the reference alarm information, anomaly handling strategy, and comprehensive detection results, the client can construct prompt words and input the prompt words into the preset information generation model. The prompt information generation model generates anomaly alarm information for the comprehensive detection results based on the anomaly handling strategy, using the reference alarm information as a reference. Then, anomaly handling is performed based on the anomaly alarm information (such as sending the anomaly alarm information to risk control auditors, or creating anomaly verification work order based on the anomaly alarm information, and then tracking the handling progress and results).
[0065] In addition, in order to accurately locate the source of abnormal data and quickly block the transmission of risks, an abnormal data source map can be constructed and a closed-loop handling mechanism can be triggered.
[0066] Specifically, the client can construct an anomaly data tracing graph based on historical data (which may include previously detected anomalous and non-anomalous data) and the participating entities involved in the data flow process. These participating entities are used to represent all flow nodes and responsible entities of the data from the source to the application stage. An example of the knowledge graph branch link determined based on the data flow relationship is: such as anomaly field → its data table → ETL task → API interface → data provider.
[0067] If the data to be tested is determined to be abnormal based on the comprehensive test results, the participating objects associated with the data source of the data to be tested are identified in the abnormal data source map, and the participating objects are handled as abnormalities.
[0068] For example, if a high-risk anomaly is detected in the "housing provident fund contribution amount" field of a user's credit data, the system can locate the data table to which the anomaly field belongs through the source map, and then associate it with the ETL task responsible for data processing and the API interface that provides data to the outside world, and finally locate the source organization of the data as a human resources service company.
[0069] In addition, to further improve the accuracy of subsequent anomaly detection results, the model parameters can be optimized and precisely iterated.
[0070] Specifically, for each detection dimension, if the deviation between the abnormal detection result and the comprehensive detection result under that detection dimension is greater than a preset deviation, then the loss value corresponding to that detection dimension is determined based on the deviation between the abnormal detection result and the comprehensive detection result under that detection dimension, and then the model parameters related to that detection dimension in the detection model are adjusted with the optimization objective of minimizing the loss value.
[0071] The aforementioned model to be detected can adopt a decoupled architecture with multi-dimensional parameter isolation. Therefore, the adjusted parameters can be sub-network parameters specific to that detection dimension. For example, a temporal feature detection sub-network for temporal detection, an attention sub-network for semantic detection, a behavior pattern mining sub-network for individual behavior features, and a group pattern statistics sub-network for group attribute features, etc.
[0072] This architecture enables independent updates of single-dimensional parameters, avoiding interference from cross-dimensional parameters, thereby improving model iteration efficiency and anomaly detection accuracy.
[0073] To facilitate understanding, this manual provides an overall flowchart for exception handling, such as... Figure 4 As shown.
[0074] Figure 4 This is an exemplary embodiment of an overall flowchart of exception handling.
[0075] The entire anomaly handling process can be divided into two stages. In the anomaly detection stage, anomalies can be detected in multiple dimensions, including logical contradiction detection, temporal mutation detection, individual abnormal behavior detection, and group deviation behavior detection (see step S302 for details). Then, based on the target weights corresponding to each detection dimension, the anomaly detection results of each detection dimension are comprehensively calculated to obtain the comprehensive detection result.
[0076] During the anomaly handling phase, various risk handling tasks can be carried out based on the comprehensive detection results, such as risk level classification, generating anomaly alarm information, tracing the source of risks based on the knowledge graph, and automatically creating and dispatching verification / processing work orders (see step 304 for details).
[0077] The purpose of this solution is to screen and process the raw credit data provided by the data provider before executing actual credit reporting business, so as to ensure that the data used by the credit reporting system in subsequent actual credit reporting business is authentic, complete and correct, thereby outputting accurate and reliable user credit assessment results and providing financial institutions with high-quality decision-making basis.
[0078] Figure 5 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 5As shown, device 500 mainly consists of a communication interface 502, a user interface 504, a processor 506, and a data storage 508. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 510. The communication interface 502 enables device 500 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 502 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 502 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 502 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 502 may also include multiple physical communication interfaces, such as Wi-Fi interfaces, Bluetooth interfaces, and wide-area wireless interfaces.
[0079] User interface 504 includes receiving user input and providing output to the user. Therefore, user interface 504 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 504 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 504 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 500 may support remote access from other devices via communication interface 502 or another physical interface (not shown). User interface 504 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 504 may also be configured as a display device for rendering or displaying text fragments.
[0080] Processor 506 may contain one or more general-purpose processors and / or special-purpose processors.
[0081] Data storage 508 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 506. Data storage 508 may include removable and non-removable components.
[0082] Processor 506 is capable of executing program instructions 518 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 508 to perform the various functions described herein. Data storage 508 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 500, enable device 500 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 518 by processor 506 may result in processor 506 using data 512.
[0083] For example, program instructions 518 may include an operating system 522 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 500 and one or more applications 520 (e.g., a browser, social application, or game application). Similarly, data 512 may include operating system data 516 and application data 514. Operating system data 516 is primarily accessible to the operating system 522, while application data 514 is primarily accessible to one or more applications 520. Application data 514 may reside in a file system visible or hidden from the user of device 500.
[0084] Application 520 can communicate with operating system 522 through one or more application programming interfaces (APIs). These APIs help application 520 read and / or write application data 514, transmit or receive information via communication interface 502, receive or display information on user interface 504, etc.
[0085] In some terminology, application 520 may be simply referred to as "app". Furthermore, application 520 can be downloaded to device 500 through one or more online app stores or app markets. However, applications can also be installed on device 500 in other ways, such as through a web browser or a physical interface on device 500 (e.g., a USB port).
[0086] Please refer to Figure 6 Data processing devices can be applied to, for example Figure 5 The device shown is used to implement the technical solution of this specification. The data processing apparatus may include: The acquisition module 600 is used to acquire the data to be detected provided by the data provider. The detection module 602 is used to input the data to be detected into a preset anomaly detection model and obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs. The processing module 604 is used to determine the target weights corresponding to each detection dimension, and to perform comprehensive calculations on the abnormal detection results of each detection dimension based on the target weights to obtain a comprehensive detection result for the data to be detected, and to perform abnormal processing based on the comprehensive detection result.
[0087] Optionally, the detection module 602 is specifically used to determine user behavior data in the data to be detected, and to determine the user behavior characteristics in the user's corresponding personal information domain; if the deviation between the current behavior characteristics reflected by the behavior data and the user behavior characteristics is greater than a preset deviation, then it is determined that there is abnormal user behavior in the data to be detected.
[0088] Optionally, the detection module 602 is specifically used to determine the user's behavior data in the data to be detected, and to determine the group attribute characteristics corresponding to the user group to which the user belongs; if the behavior data does not conform to the normal behavior characteristics of the user group to which it belongs, it is determined that there is abnormal user behavior in the data to be detected.
[0089] Optionally, the detection module 602 is specifically used to update the basic weights pre-set for each detection dimension based on the abnormality score contained in the abnormality detection results of that detection dimension, so as to obtain the target weights corresponding to that detection dimension.
[0090] Optionally, the detection module 602 is specifically configured to: if the anomaly score of the data to be detected in the detection dimension is greater than a first preset score, increase the basic weight corresponding to the detection dimension to obtain the target weight; if the anomaly score of the data to be detected in the detection dimension is less than a second preset score, decrease the basic weight corresponding to the detection dimension to obtain the target weight; if the anomaly score of the data to be detected in the detection dimension is neither greater than the first preset score nor less than the second preset score, determine the basic weight as the target weight; wherein, the first preset score is greater than the second preset score, and when the anomaly score is greater than the first preset score or less than the second preset score, the target weight is positively correlated with the anomaly score.
[0091] Optionally, the device further includes: The adjustment module 606 is used to determine the loss value corresponding to the detection dimension based on the deviation between the abnormal detection result and the comprehensive detection result under each detection dimension if the deviation is greater than a preset deviation. The model parameters related to the detection dimension in the model to be detected are adjusted with the goal of minimizing the loss value.
[0092] Optionally, the processing module 604 is specifically configured to: determine an anomaly handling strategy matching the comprehensive detection result; and determine reference alarm information matching the comprehensive detection result in a preset alarm information database; construct prompt words based on the reference alarm information, the anomaly handling strategy, and the comprehensive detection result, and input the prompt words into a preset information generation model to prompt the information generation model to: generate anomaly alarm information for the comprehensive detection result based on the anomaly handling strategy, using the reference alarm information as a reference; and perform anomaly handling based on the anomaly alarm information.
[0093] Optionally, the processing module 604 is specifically used to construct an abnormal data source map based on the participants involved in the data flow process of historical data; if it is determined that the data to be detected is abnormal based on the comprehensive detection results, then the participants associated with the data source of the data to be detected are identified in the abnormal data source map, and the participants are subjected to abnormal processing.
[0094] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0095] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.
[0096] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0097] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0098] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.
[0099] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.
[0100] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0101] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0102] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0103] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0104] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.
Claims
1. An exception handling method, comprising: Obtain the data to be tested provided by the data provider; The data to be detected is input into a preset anomaly detection model to obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs. The target weights corresponding to each detection dimension are determined, and the anomaly detection results of each detection dimension are comprehensively calculated based on the target weights to obtain the comprehensive detection results for the data to be detected. Anomaly handling is then performed based on the comprehensive detection results.
2. The method as described in claim 1, wherein abnormal user behavior detection is performed on the data to be detected based on the user's user behavior characteristics, specifically including: The user's behavioral data is determined from the data to be detected, and the user's behavioral characteristics are determined from the user's corresponding personal information field; If the deviation between the current behavioral characteristics reflected by the behavioral data and the user behavioral characteristics is greater than a preset deviation, then it is determined that there is abnormal user behavior in the data to be detected.
3. The method as described in claim 1, wherein abnormal user behavior detection is performed on the data to be detected based on the group attribute characteristics of the user's user group, specifically including: The user's behavioral data is determined from the data to be detected, and the group attribute characteristics corresponding to the user group to which the user belongs are determined. If, based on the group attribute characteristics, it is determined that the behavioral data does not conform to the normal behavioral characteristics of its user group, then it is determined that there is abnormal user behavior in the data to be detected.
4. The method as described in claim 1, determining the target weights corresponding to each detection dimension, specifically includes: For each detection dimension, the pre-set base weights for that detection dimension are updated based on the anomaly score contained in the anomaly detection results of that detection dimension, thus obtaining the target weights corresponding to that detection dimension.
5. The method as described in claim 4, wherein updating the pre-set base weights for the detection dimension to obtain the target weights corresponding to the detection dimension specifically includes: If the anomaly score of the data to be detected in this detection dimension is greater than the first preset score, then the basic weight corresponding to this detection dimension is increased to obtain the target weight; If the anomaly score of the data to be detected in this detection dimension is less than the second preset score, then the base weight corresponding to this detection dimension is reduced to obtain the target weight; If the anomaly score of the data to be detected in this detection dimension is not greater than the first preset score and not less than the second preset score, then the basic weight is determined as the target weight. Wherein, the first preset score is greater than the second preset score, and when the anomaly score is greater than the first preset score or less than the second preset score, the target weight is positively correlated with the anomaly score.
6. The method of claim 1, further comprising: For each detection dimension, if the deviation between the abnormal detection result under that detection dimension and the comprehensive detection result is greater than a preset deviation, then the loss value corresponding to that detection dimension is determined based on the deviation between the abnormal detection result under that detection dimension and the comprehensive detection result. With the goal of minimizing the loss value, the model parameters related to the detection dimension in the model to be detected are adjusted.
7. The method as described in claim 1, wherein anomaly handling is performed based on the comprehensive detection results, specifically including: Determine an anomaly handling strategy that matches the comprehensive detection results, and identify reference alarm information that matches the comprehensive detection results from a preset alarm information database; Based on the reference alarm information, the anomaly handling strategy, and the comprehensive detection result, a prompt word is constructed, and the prompt word is input into a preset information generation model to prompt the information generation model to generate an anomaly alarm information for the comprehensive detection result based on the anomaly handling strategy, with the reference alarm information as a reference. Perform anomaly handling based on the aforementioned abnormal alarm information.
8. The method as described in claim 1, wherein before performing anomaly processing based on the comprehensive detection results, the method further comprises: Based on the participants involved in the data flow process in historical data, construct an anomaly data tracing map; Anomaly handling is performed based on the comprehensive detection results, specifically including: If the data to be tested is determined to be abnormal based on the comprehensive detection results, then the participating objects associated with the data source of the data to be tested are identified in the abnormal data tracing map, and the participating objects are subjected to abnormal processing.
9. An anomaly handling device, comprising: The acquisition module is used to acquire the data to be tested provided by the data provider; The detection module is used to input the data to be detected into a preset anomaly detection model and obtain anomaly detection results of multiple detection dimensions output by the anomaly detection model; wherein, the multiple detection dimensions are selected from: logical contradiction detection of fields contained in the data to be detected, outlier detection of the data to be detected based on the time data sequence corresponding to the data to be detected, abnormal user behavior detection of the data to be detected based on user behavior characteristics, and abnormal user behavior detection of the data to be detected based on the group attribute characteristics of the user group to which the user belongs. The processing module is used to determine the target weights corresponding to each detection dimension, and to perform comprehensive calculations on the abnormal detection results of each detection dimension based on the target weights to obtain a comprehensive detection result for the data to be detected, and to perform abnormal processing based on the comprehensive detection result.
10. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-8 by executing the executable instructions.
11. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-8.
12. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-8.