Large-amount credit business risk monitoring system based on PDF (Portable Document Format) data analysis

Through a system based on PDF data analysis, combined with OCR technology and text mining technology, the risk information inside and outside the bank is integrated, and the problem of inefficient integration and utilization of risk data in the existing technology is solved, and multi-dimensional risk monitoring and efficient risk control for large-scale credit customers is realized.

CN119991280APending Publication Date: 2025-05-13SHANGHAI RURAL COMML BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411993213.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively integrate and utilize different types of risk data inside and outside the industry, especially unstructured data, resulting in inefficient risk monitoring.

Method used

A system based on PDF data analysis is adopted, combining OCR technology, text mining technology and RPA technology, and integrating risk information inside and outside the industry, transforming unstructured data into structured data, and risk analysis is carried out through artificial intelligence and knowledge graph models.

Benefits of technology

It has realized multi-dimensional risk monitoring for large-scale credit customers, improved the richness of risk analysis data and the scalability of risk control methods, and significantly improved the efficiency of risk monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991280A_ABST
    Figure CN119991280A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence application, and discloses a PDF (Portable Document Format) data analysis-based large-amount credit granting business risk monitoring system, which integrates internal and external risk information through an OCR (Optical Character Recognition) technology. Large-amount customer risk monitoring reports are summarized from multiple dimensions of industrial and commercial tax affairs, judicial complaints, financial information, in-bank fund flow, checking and deducting data, external news, public opinion information and the like, and first-line and strip-line personnel are assisted to carry out risk control tracking through one-stop reading; financial science and technology means are fully utilized, risk analysis data content is enriched, risk monitoring means are expanded, and risk monitoring work efficiency is improved; the method comprises the following steps of: identifying and judging unstructured data through an OCR (Optical Character Recognition) technology, converting various acquired image data into text contents by virtue of an optical character identification technology, a text mining technology, an RPA (Registered Packet Access) technology and the like, extracting important key information by taking the text mining technology as a carrier, performing operation by utilizing models such as artificial intelligence and a knowledge graph, and generating a judgment result; and important risk information is prompted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence application technology, and in particular to a large-amount credit business risk monitoring system based on PDF data analysis. Background Art

[0002] Faced with the triple constraints of "capital constraints, financial constraints, and risk constraints", the trend of strong external supervision continues. The banking industry must continue to pay attention to risk control in key areas to further ensure the high-quality development of the business. For large-amount credit customers (this project defines customers with a loan balance of ≥50 million yuan as large-amount credit customers), because of their wide data sources and high credibility, they should fully combine the new direction of "digital transformation" and use financial technology to better integrate the risk data available inside and outside the bank.

[0003] At present, different types of risk data inside and outside the bank are scattered in various lines, and the utilization rate of unstructured data is low, and the data value of unstructured information such as paper and images has not been fully explored. Frontline personnel need to spend a lot of time and energy to obtain risk data through different channels. Unstructured information needs to be manually judged based on existing experience, and the risk control tools and means of users are relatively basic. Fragmented information data requires frontline and risk monitoring personnel to spend a lot of time and energy to identify substantive risks, and risk control efficiency needs to be improved urgently.

[0004] Therefore, we need to propose a large-amount credit business risk monitoring system based on PDF data parsing, implement the risk warning process from the technology of data collection, data structuring processing and risk model generation of risk reports, and convert unstructured risk warning data into structured data based on OCR technology for risk warning. Summary of the invention

[0005] The purpose of the present invention is to provide a large-amount credit business risk monitoring system based on PDF data analysis, integrate internal and external risk information through OCR technology, summarize large-amount customer risk monitoring reports from multiple dimensions such as industrial and commercial taxation, judicial litigation, financial information, internal fund flow, frozen and withheld data, external news, and public opinion information, and assist front-line and line personnel to carry out risk control tracking through one-stop reading; make full use of financial technology means to enrich risk analysis data content, expand risk monitoring means, and improve risk monitoring work efficiency; identify and judge unstructured data through OCR technology, and use optical character recognition technology, text mining technology, and RPA technology to realize intelligent information collection. Convert the collected various types of image data into text content, and then use text mining technology as a carrier to extract important key information, use artificial intelligence, knowledge graphs and other models for calculation, generate judgment results, and prompt important risk information to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a large amount credit business risk monitoring system based on PDF data analysis, comprising:

[0007] Data preparation module for data cleaning of structured and unstructured data;

[0008] A data preprocessing module for processing structured data and unstructured data and obtaining feature words;

[0009] A risk monitoring module that uses the data and feature words obtained by the data preparation module and the data preprocessing module to perform risk model calculations and generate risk disclosure results;

[0010] Structured data includes page input data, file import data, and data from the data sharing support platform; unstructured data includes PDF data, image data, and data collected from online forums;

[0011] The data preprocessing module includes:

[0012] OCR recognition unit: recognize text information in unstructured data;

[0013] Data association and comparison unit: selects major risk indicators and compares them with the established vocabulary to identify and evaluate potential risk information in the text;

[0014] The risk monitoring module includes a model operation unit and a risk disclosure unit.

[0015] Preferably, structured data is in-bank data, including basic customer information, customer deposit and loan details, and other customer information; unstructured data is out-of-bank data, including borrower's license and certificate, other bank account statements, invoice information, online public opinion information, collateral information, and casual photo information.

[0016] Preferably, the basic customer information includes the customer's name, ID number, contact information, home address, and company address; the customer's deposit and loan details include the customer's deposit records and loan records; and other customer information includes the customer's transaction records, credit scores, and historical risks.

[0017] Preferably, the borrower's license documents include the customer's business license, ID card, and driver's license; other bank account statements are the borrower's transaction records in other banks; invoice information is invoice data used to analyze customer consumption behavior and transaction patterns; online public opinion information is the customer's comments and comments on the Internet obtained through a third party; collateral information is collateral information provided by the borrower for value assessment; and casual picture information is picture information containing customers or related scenes that is used for uploading or obtained through social media channels.

[0018] Preferably, the OCR recognition unit uses OCR recognition technology to convert text in bills, newspapers, pictures, PDFs, and books into text data in the same format as the structured data;

[0019] The text extraction process of OCR recognition technology is as follows:

[0020] A1. Image preprocessing: first input the image data and do grayscale processing to simplify the image complexity, then do binarization and denoising to improve the clarity of the text;

[0021] A2. Text area detection: determine the area containing text in the image and analyze the outline of the text area;

[0022] A3. Character segmentation: Segment the detected text area into individual characters and perform feature extraction on the individual characters;

[0023] A4. Character recognition: Convert the character results recognized by the classifier into recognizable character codes.

[0024] Preferably, the process of the data association and comparison unit identifying and evaluating the potential risk information of the text generated by the OCR recognition unit is as follows:

[0025] B1. Data resource integration: Receive data from different sources and in different formats, integrate the data into a unified database, and use unique identifiers to match the data to ensure data consistency;

[0026] B2. Feature matching: Compare the features extracted from unstructured data with key indicators in structured data, identify correlations, and build a network between customers and their transactions, social relationships, and financial status;

[0027] B3. Compare the extracted key indicators with the preset risk thresholds, evaluate the multi-dimensional data performance, and analyze the customer's performance in different dimensions;

[0028] B4. Feature analysis: Compare the characteristics of the data set, compare the customer's financial status with the industry average, and identify potential risks.

[0029] Preferably, the data association and comparison unit also uses AI semantic analysis technology to perform semantic analysis on the converted text data, and extract keywords and topics in the text, identify the customer's emotional tendencies and opinions in online public opinion information, and reveal the customer's potential risks.

[0030] Preferably, the model operation unit includes:

[0031] Anti-money laundering model: Identify and monitor customers and behaviors involved in money laundering activities;

[0032] Relationship map: Build a relationship map between customers, reveal potential association risks, and analyze the kinship and cooperation relationships between customers.

[0033] Preferably, the risk disclosure unit generates a risk disclosure report based on the calculation results of the risk model, the content of the risk disclosure report includes risk type, risk level, involved customers, recommended measures, and risk prevention and control measures are taken according to the risk disclosure report information.

[0034] Preferably, the training process of unstructured data is as follows:

[0035] S1. Data collection: Automatically collect online forum data through RPA, collect key risk vocabulary, and query credit customer information from the database;

[0036] S2. Data preprocessing: sort out the unstructured data that credit grantors are concerned about and perform data normalization;

[0037] S3, unstructured data to structured data: Use OCR recognition and intelligent voice to convert unstructured data into text data;

[0038] S4. Word segmentation: Word segmentation is achieved through hidden Markov model, backward maximum matching algorithm and conditional random multiple algorithms, and multiple matching results are deduplicated;

[0039] S5. Obtain model input and output parameters: The Word2vec word embedding method matches the cosine similarity between word vectors of the existing structured data, extracts the key risk words, takes the key risk words as input, and uses the existing structured model to predict the risk value as output;

[0040] S6. Model training: Use key risk words as input and the results of the existing structured data model as output, and use the TextCNN model for training.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1. The present invention integrates internal and external risk information through OCR technology, summarizes large-value customer risk monitoring reports from multiple dimensions such as industrial and commercial taxation, judicial litigation, financial information, internal fund flow, frozen and withheld data, external news, and public opinion information, and assists front-line and line personnel in risk control tracking through one-stop reading; fully utilizes financial technology means to enrich risk analysis data content, expand risk monitoring means, and improve risk monitoring work efficiency;

[0043] 2. The present invention uses OCR technology to identify and judge unstructured data, and uses optical character recognition technology, text mining technology, and RPA technology to achieve intelligent information collection. The collected various types of image data are converted into text content, and then text mining technology is used as a carrier to extract important key information. Artificial intelligence, knowledge graphs and other models are used for calculation to generate judgment results and prompt important risk information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a system block diagram of the present invention;

[0045] Figure 2 This is a training flow chart of unstructured data of the present invention;

[0046] Figure 3 It is a flow chart of the risk monitoring of the system of the present invention;

[0047] Figure 4 It is a flow chart of text extraction of OCR recognition technology of the present invention;

[0048] Figure 5 A flow chart for generating text potential risk information for the data association and comparison unit of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] See also Figure 1-5 The present invention provides a technical solution: a large-amount credit business risk monitoring system based on PDF data analysis, which uses OCR technology to identify and judge unstructured data, and uses optical character recognition technology, text mining technology, and RPA technology to achieve intelligent information collection. The collected various types of image data are converted into text content, and then the text mining technology is used as a carrier to extract important key information, and the artificial intelligence, knowledge graph and other models are used for calculation to generate judgment results and prompt important risk information;

[0051] include:

[0052] Data preparation module for data cleaning of structured and unstructured data;

[0053] The data preparation module uses rich data sources from inside and outside the bank to clean structured and unstructured data, providing a high-quality data foundation for subsequent data analysis and risk monitoring.

[0054] Data cleaning includes:

[0055] Deduplication: Remove duplicate data to ensure data uniqueness.

[0056] Completion: Complete missing data, such as through data interpolation, prediction, etc.

[0057] Standardization: Convert data into a unified format and standard to facilitate subsequent processing and analysis.

[0058] Outlier processing: Identify and process outliers to ensure the accuracy and rationality of data.

[0059] Structured data refers to in-bank data, including basic customer information, customer deposit and loan details, and other customer information;

[0060] Customer basic information includes the customer's name, ID number, contact information, home address, and company address; customer deposit and loan details include the customer's deposit records and loan records; other customer information includes the customer's transaction records, credit score, and historical risks.

[0061] Unstructured data refers to off-bank data, including the borrower’s license and credentials, other bank account statements, invoice information, online public opinion information, collateral information, and casually taken pictures.

[0062] The borrower's license documents include the customer's business license, ID card, and driver's license. The account statements of other banks are the borrower's transaction records in other banks. The invoice information is the invoice data used to analyze the customer's consumption behavior and transaction patterns. The online public opinion information is the customer's comments and evaluations on the Internet obtained through a third party. The collateral information is the collateral information provided by the borrower for value assessment. The casual picture information is the picture information containing the customer or related scenes for uploading or obtained through social media channels.

[0063] A data preprocessing module is used to process structured data and unstructured data and obtain feature words; the data preprocessing module is responsible for further processing and conversion of the original data in order to extract useful features and information from it to provide support for subsequent risk monitoring.

[0064] The data preprocessing module includes:

[0065] OCR recognition unit: Identifies text information in unstructured data and converts it into a machine-readable format to facilitate subsequent data analysis; optionally, combines manual review and machine learning algorithms to improve recognition accuracy and ensure the reliability of extracted information.

[0066] The OCR recognition unit uses OCR recognition technology to convert text in bills, newspapers, pictures, PDFs, and books into text data in the same format as structured data;

[0067] The text extraction process of OCR recognition technology is as follows:

[0068] A1. Image preprocessing: first input the image data and do grayscale processing to simplify the image complexity, then do binarization and denoising to improve the clarity of the text;

[0069] Specifically, image input: The OCR system inputs paper documents or pictures into the computer through devices such as scanners, digital cameras, and mobile phones;

[0070] Grayscale: Convert color images into grayscale images to simplify image complexity;

[0071] Binarization: The grayscale image is further converted into a binary image, that is, the text part in the image becomes black and the background becomes white, which helps to simplify the image information and facilitate subsequent text extraction and recognition;

[0072] Denoising: removes cluttered information in the image, such as noise, stains, etc., to improve text clarity;

[0073] Deskew: Adjust the image orientation to ensure text is aligned horizontally.

[0074] A2. Text area detection: Through image analysis and edge detection algorithms, determine the area containing text in the image, and analyze the outline of the text area to locate the text area more accurately;

[0075] A3. Character segmentation: Segment the detected text area into individual characters and extract features from the individual characters; commonly used methods include projection, template matching, neural networks, etc. The extracted features include shape, angle, texture, etc., for subsequent classification and recognition.

[0076] A4. Character recognition: Use the trained model or algorithm as a classifier to classify the extracted character features and convert the character results recognized by the classifier into recognizable character codes.

[0077] Data association and comparison unit: selects major risk indicators (such as credit scores and post-loan management indicators) and compares them with the established vocabulary (such as industry terms and risk warning words) to identify and evaluate potential risk information in the text;

[0078] Key risk indicators include: credit score, repayment history, delinquency rate, financial ratios (such as current ratio, debt-to-asset ratio), income stability, etc.; and specific risk thresholds are determined based on industry standards and historical data to facilitate subsequent analysis.

[0079] Compare information extracted from unstructured data with key risk indicators to identify correlations;

[0080] Integrate the extracted information with existing structured data to create a comprehensive data model for more comprehensive analysis, assign weights to each indicator based on its importance, perform weighted scoring, and calculate the overall credit risk score;

[0081] When an indicator exceeds the set threshold, an alarm is automatically triggered, prompting risk managers to conduct further review.

[0082] The process by which the data association and comparison unit identifies and evaluates the potential risk information of the text generated by the OCR recognition unit is as follows:

[0083] B1. Data resource integration: Receive data from different sources and in different formats, integrate the data into a unified database, and use unique identifiers to match the data to ensure data consistency;

[0084] B2. Feature matching: Compare the features extracted from unstructured data with key indicators in structured data, identify correlations, and build a network between customers and their transactions, social relationships, and financial status;

[0085] B3. Compare the extracted key indicators with the preset risk thresholds, evaluate the multi-dimensional data performance, and analyze the customer's performance in different dimensions;

[0086] B4. Feature analysis: Compare the characteristics of the data set, compare the customer's financial status with the industry average, and identify potential risks.

[0087] The data association and comparison unit extracts characteristic words related to customers or risks from various data sources. These characteristic words serve as important inputs for the subsequent risk monitoring module.

[0088] The data association and comparison unit also uses AI semantic analysis technology to perform semantic analysis on the converted text data, extract keywords and topics in the text, identify customers' emotional tendencies and opinions in online public opinion information, and reveal customers' potential risks.

[0089] Structured data includes page input data, file import data, and data from the data sharing support platform; unstructured data includes PDF data, image data, and data collected from online forums;

[0090] A risk monitoring module that uses the data and feature words obtained by the data preparation module and the data preprocessing module to perform risk model calculations and generate risk disclosure results;

[0091] The risk monitoring module includes a model operation unit and a risk disclosure unit.

[0092] The model operation unit includes:

[0093] Anti-money laundering model: Identify and monitor customers and behaviors involved in money laundering activities;

[0094] Relationship map: Build a relationship map between customers, reveal potential association risks, and analyze the kinship and cooperation relationships between customers.

[0095] The risk disclosure unit generates a risk disclosure report based on the calculation results of the risk model. The content of the risk disclosure report includes risk type, risk level, involved customers, and recommended measures, and risk prevention and control measures are taken based on the information in the risk disclosure report.

[0096] This system integrates internal and external risk information through OCR technology, and summarizes large-value customer risk monitoring reports from multiple dimensions such as industrial and commercial taxation, judicial litigation, financial information, internal fund flow, frozen and withheld data, external news, and public opinion information. Through one-stop reading, it assists front-line and line personnel in conducting risk control tracking. Make full use of financial technology to enrich risk analysis data content, expand risk monitoring methods, and improve risk monitoring efficiency.

[0097] The training process for unstructured data is as follows:

[0098] S1. Data collection: Automatically collect online forum data through RPA, collect key risk vocabulary, and query credit customer information from the database;

[0099] S2. Data preprocessing: sort out the unstructured data that credit grantors are concerned about and perform data normalization;

[0100] S3, unstructured data to structured data: Use OCR recognition and intelligent voice to convert unstructured data into text data;

[0101] S4. Word segmentation: Word segmentation is achieved through hidden Markov model, backward maximum matching algorithm and conditional random multiple algorithms, and multiple matching results are deduplicated;

[0102] S5. Obtain model input and output parameters: The Word2vec word embedding method matches the cosine similarity between word vectors of the existing structured data, extracts the key risk words, takes the key risk words as input, and uses the existing structured model to predict the risk value as output;

[0103] S6. Model training: Use key risk words as input and the results of the existing structured data model as output, and use the TextCNN model for training.

[0104] When the system is running, it includes the following processes:

[0105] C1. Data collection: Automatically collect online forum data through RPA, collect key risk vocabulary, and query credit customer information from the database;

[0106] C2. Data preprocessing: sort out the unstructured data that credit grantors are concerned about and perform data normalization;

[0107] C3. Convert unstructured data to structured data: Use OCR recognition and intelligent voice to convert unstructured data into text data;

[0108] C4, word segmentation: word segmentation is achieved through hidden Markov model, backward maximum matching algorithm and conditional random multiple algorithms, and multiple matching results are deduplicated;

[0109] C5. Keyword matching: The Word2vec word embedding method matches the cosine similarity between word vectors of the existing structured data to extract keyword risk vocabulary;

[0110] C6. Model calculation: For non-structured models, the keywords are vectorized and input into TextCNN with structured data for model calculation; for existing structured models, the structured data is input into the existing model to calculate the risk threshold;

[0111] C7. Comprehensive results: Integrate structured and unstructured data to add comprehensive risk scoring indicators;

[0112] C8. Output result report: The report displays the risk scores and comprehensive risk scores calculated by the unstructured model and the structured model respectively; the keywords in the various types of unstructured data used for reference are highlighted according to the influence of each keyword in the unstructured model on the model results (the weight in the model parameters);

[0113] C9. Credit review personnel use the report as a reference for credit indicators.

[0114] Existing risk warning technologies all use structured data for risk warning judgment. This application applies OCR technology to the large-scale credit risk detection system for the first time, fully tapping the value of unstructured data and effectively improving the risk detection effect.

[0115] By integrating OCR and risk warning models to form a unified risk warning system, and by optimizing the system design, it is able to generate risk warning reports for large-scale credit customers in real time, simplify the work flow of business personnel, and improve the warning effect.

[0116] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large amount credit business risk monitoring system based on PDF data analysis, characterized in that: include: Data preparation module for data cleaning of structured and unstructured data; A data preprocessing module for processing structured data and unstructured data and obtaining feature words; A risk monitoring module that uses the data and feature words obtained by the data preparation module and the data preprocessing module to perform risk model calculations and generate risk disclosure results; Structured data includes page input data, file import data, and data from the data sharing support platform; unstructured data includes PDF data, image data, and data collected from online forums; The data preprocessing module includes: OCR recognition unit: recognize text information in unstructured data; Data association and comparison unit: selects major risk indicators and compares them with the established vocabulary to identify and evaluate potential risk information in the text; The risk monitoring module includes a model operation unit and a risk disclosure unit.

2. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized by: Structured data refers to in-bank data, including basic customer information, customer deposit and loan details, and other customer information; unstructured data refers to out-bank data, including borrower's license and credentials, other bank account statements, invoice information, online public opinion information, collateral information, and casual photo information.

3. According to claim 2, a large amount of credit business risk monitoring system based on PDF data analysis is characterized in that: Customer basic information includes the customer's name, ID number, contact information, home address, and company address; customer deposit and loan details include the customer's deposit records and loan records; other customer information includes the customer's transaction records, credit score, and historical risks.

4. According to claim 2, a large amount of credit business risk monitoring system based on PDF data analysis is characterized in that: The borrower's license documents include the customer's business license, ID card, and driver's license. The account statements of other banks are the borrower's transaction records in other banks. The invoice information is the invoice data used to analyze the customer's consumption behavior and transaction patterns. The online public opinion information is the customer's comments and evaluations on the Internet obtained through a third party. The collateral information is the collateral information provided by the borrower for value assessment. The casual picture information is the picture information containing the customer or related scenes for uploading or obtained through social media channels.

5. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized in that: The OCR recognition unit uses OCR recognition technology to convert text in bills, newspapers, pictures, PDFs, and books into text data in the same format as structured data; The text extraction process of OCR recognition technology is as follows: A1. Image preprocessing: first input the image data and do grayscale processing to simplify the image complexity, then do binarization and denoising to improve the clarity of the text; A2. Text area detection: determine the area containing text in the image and analyze the outline of the text area; A3. Character segmentation: Segment the detected text area into individual characters and perform feature extraction on the individual characters; A4. Character recognition: Convert the character results recognized by the classifier into recognizable character codes.

6. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized in that: The process by which the data association and comparison unit identifies and evaluates the potential risk information of the text generated by the OCR recognition unit is as follows: B1. Data resource integration: Receive data from different sources and in different formats, integrate the data into a unified database, and use unique identifiers to match the data to ensure data consistency; B2. Feature matching: Compare the features extracted from unstructured data with key indicators in structured data, identify correlations, and build a network between customers and their transactions, social relationships, and financial status; B3. Compare the extracted key indicators with the preset risk thresholds, evaluate the multi-dimensional data performance, and analyze the customer's performance in different dimensions; B4. Feature analysis: Compare the characteristics of the data set, compare the customer's financial status with the industry average, and identify potential risks.

7. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized in that: The data association and comparison unit also uses AI semantic analysis technology to perform semantic analysis on the converted text data, extract keywords and topics in the text, identify customers' emotional tendencies and opinions in online public opinion information, and reveal customers' potential risks.

8. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized in that: The model operation unit includes: Anti-money laundering model: Identify and monitor customers and behaviors involved in money laundering activities; Relationship map: Build a relationship map between customers, reveal potential association risks, and analyze the kinship and cooperation relationships between customers.

9. According to claim 1, a large amount credit business risk monitoring system based on PDF data analysis is characterized in that: The risk disclosure unit generates a risk disclosure report based on the calculation results of the risk model. The content of the risk disclosure report includes risk type, risk level, involved customers, and recommended measures, and risk prevention and control measures are taken based on the information in the risk disclosure report.

10. The large amount credit business risk monitoring system based on PDF data analysis according to claim 1 is characterized by: The training process for unstructured data is as follows: S1. Data collection: Automatically collect online forum data through RPA, collect key risk vocabulary, and query credit customer information from the database; S2. Data preprocessing: sort out the unstructured data that credit grantors are concerned about and perform data normalization; S3, unstructured data to structured data: Use OCR recognition and intelligent voice to convert unstructured data into text data; S4. Word segmentation: Word segmentation is achieved through hidden Markov model, backward maximum matching algorithm and conditional random multiple algorithms, and multiple matching results are deduplicated; S5. Obtain model input and output parameters: The Word2vec word embedding method matches the cosine similarity between word vectors of the existing structured data, extracts the key risk words, takes the key risk words as input, and uses the existing structured model to predict the risk value as output; S6. Model training: Use key risk words as input and the results of the existing structured data model as output, and use the TextCNN model for training.