Credit text data processing method and system based on NLP
Through the NLP-based credit text data processing method, word segmentation and reorganization of phrases, the establishment of a credit fraud intention word set, the screening of suspected credit fraud texts, and the combination of user categories and bad records solved the problem of inaccurate credit fraud detection and achieved more accurate fraud identification.
Patent Information
- Application Number
- CN202511046329.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing credit fraud detection methods cannot accurately identify fraudsters who use typos or unconventional words to evade detection, and cannot distinguish information sharing between normal users, resulting in inaccurate detection results.
Adopting the credit text data processing method based on NLP, through word segmentation, phrase reorganization, and the establishment of a credit fraud intention word set, combined with context and user relationships, suspected credit fraud texts are screened, and the degree of fraud is determined based on user behavior records for accurate identification.
It improves the accuracy of credit fraud detection, avoids the possibility of fraudsters evading detection, accurately identifies user credit fraud behavior, and reduces misjudgment of normal users.
Smart Images

Figure CN120542422B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a credit text data processing method and system based on NLP. Background Art
[0002] Credit fraud can lead to deterioration in asset quality and affect genuine credit assessments, necessitating detection. Credit text data, such as loan applications, credit reports, and proof of income, contains clues to fraud, so analysis of this data is often used to identify credit fraud.
[0003] Currently, fraud-related vocabulary is generally defined manually, and credit texts are matched with fraud intent words in the fraud-related vocabulary. In combination with the user relationships between the users corresponding to the credit texts (credit fraudsters may steal the identities of multiple users and share information), fraud intent is identified, and then warnings are issued for credit texts with higher fraud risks.
[0004] However, in actual applications, fraudsters may use typos or unconventional words to evade detection, and there is a possibility of frequent changes in wording. In addition, normal users may also share information, such as family members, or there may be false associations that interfere with detection. Therefore, current methods cannot accurately distinguish the above situations, resulting in inaccurate detection results of credit fraud. Summary of the Invention
[0005] In order to solve the technical problem of inaccurate credit fraud detection, the present invention aims to provide a method and system for processing credit text data based on NLP. The technical solutions adopted are as follows:
[0006] In a first aspect, the present invention provides a method for processing credit text data based on NLP, the method comprising:
[0007] Perform word segmentation on multiple credit texts to be tested to obtain word segmentation results;
[0008] Recombining each phrase in each single sentence in the word segmentation result to obtain multiple recombined words;
[0009] Establishing a set of credit fraud intent words based on the association between each of the recombined words and high-frequency fraud intent words, historical normal credit texts, and context;
[0010] Screening suspected credit fraud texts based on the consistency and contextual connection between the credit texts to be tested that contain the credit fraud intention words;
[0011] For each user category, determining the possibility of fraud in the user category based on the relationship between the suspected credit fraud texts corresponding to the users in the user category;
[0012] The extent to which the user meets the criteria for credit fraud is determined based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs.
[0013] According to the NLP-based credit text data processing method provided by the present invention, the method establishes a set of credit fraud intent words based on the association between each of the recombined words and high-frequency fraud intent words, historical normal credit texts, and context, including:
[0014] For each of the recombined words, determine the likelihood that the recombined word is a credit fraud intent word based on its similarity to various high-frequency fraud intent words, its frequency of appearance in historical normal credit texts, its similarity to various high-frequency phrases in the historical normal credit texts, and its similarity to the remaining words in the credit text to be tested, excluding the recombined word.
[0015] Determining the credit fraud intention words in each of the recombined words based on the likelihood that each of the recombined words corresponds to the credit fraud intention word;
[0016] A set of credit fraud intention words is established based on the high-frequency fraud intention words and the credit fraud intention words in each of the reorganized words.
[0017] According to the NLP-based credit text data processing method provided by the present invention, screening suspected credit fraud texts based on the consistency and contextual connection between the credit texts to be tested containing the credit fraud intention words includes:
[0018] screening target credit texts to be tested containing the credit fraud intention words from the credit texts to be tested;
[0019] determining the fuzziness level of each of the credit fraud intention words appearing in each of the target credit documents to be tested;
[0020] For each target credit text to be tested, determine the credit fraud degree of the target credit text to be tested based on the growth of the frequency of the credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text to be tested, the fuzziness of each credit fraud intention word appearing in the target credit text to be tested, the number of target paragraphs containing the credit fraud intention words in the target credit text to be tested, and the frequency of the triples in the target credit text to be tested in the remaining target credit texts to be tested of the corresponding user;
[0021] Based on the credit fraud degrees corresponding to the target credit texts to be tested, suspected credit fraud texts are screened from the target credit texts to be tested.
[0022] According to the NLP-based credit text data processing method provided by the present invention, determining the fuzziness level of each credit fraud intention word appearing in each target credit text to be tested includes:
[0023] For each credit fraud intention word appearing in the target credit text to be tested, the fuzziness of the credit fraud intention word is determined based on the difference in similarity between the credit fraud intention word and normal words in each paragraph, and the similarity between the word vectors corresponding to the credit fraud intention word in each paragraph.
[0024] According to the NLP-based credit text data processing method provided by the present invention, determining the possibility of fraud in the user category based on the connection between the suspected credit fraud texts corresponding to the users in the user category includes:
[0025] determining the likelihood that each triple in each of the suspected credit fraud texts in the user category belongs to a fixed template of the user category;
[0026] determining a fixed triplet of the user category according to the likelihood that each triplet in the user category corresponds to a fixed template of the user category;
[0027] The possibility of fraud in the user category is determined based on the number of users in the user category, the number of fixed triples in the user category, the degree of credit fraud of each suspected credit fraud text in the user category, and the time interval between the release of each suspected credit fraud text in the user category.
[0028] According to the NLP-based credit text data processing method provided by the present invention, determining the possibility that each triple in each suspected credit fraud text in the user category belongs to the fixed template of the user category includes:
[0029] For each triple in each suspected credit fraud text in the user category, determine the possibility that the triple belongs to the fixed template of the user category based on the frequency of occurrence of the triple in each suspected credit fraud text in the user category, the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text in the user category and the discrete degree of the number of characters, and the difference in the proportion of single sentences of the triple in each suspected credit fraud text in the user category.
[0030] According to the NLP-based credit text data processing method provided by the present invention, the determination of the degree to which a user meets the criteria for credit fraud based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs includes:
[0031] Determine the ratio between the number of bad repayment records in the history of the user and the number of bad repayment records in the history of the user corresponding to each of the credit documents to be tested, and obtain the bad credit behavior rate of the user;
[0032] The degree to which the user meets the credit fraud criteria is determined based on the user's own bad credit behavior rate and the possibility of fraud in the user category to which the user belongs.
[0033] According to the NLP-based credit text data processing method provided by the present invention, after determining the extent to which the user meets the credit fraud criteria based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs, the method further includes:
[0034] Determining a priority value for alerting the user based on the degree to which the user meets the criteria for credit fraud, the number of suspected credit fraud texts posted by the user, and the time since the user last posted the suspected credit fraud text;
[0035] According to the priority value, the user is warned of the suspected credit fraud text.
[0036] According to the NLP-based credit text data processing method provided by the present invention, alerting the user of the suspected credit fraud text according to the priority value includes:
[0037] Determine the target user whose priority value is greater than a preset priority value threshold;
[0038] The suspected credit fraud text of each target user is warned in descending order of the priority value.
[0039] In a second aspect, the present invention provides an NLP-based credit text data processing system, the system comprising a memory and a processor; the memory is used to store executable program code; the processor is used to call and run the executable program code from the memory to implement the NLP-based credit text data processing method provided by the present invention.
[0040] The present invention has the following advantageous effects: Each phrase in each single sentence from the word segmentation results of multiple credit texts to be tested is reorganized to obtain multiple reorganized words. Based on the association of each reorganized word with high-frequency fraud intent words, historical normal credit texts, and context, a set of credit fraud intent words can be accurately established. Compared to methods that manually define a credit fraud intent vocabulary, this method can prevent fraudsters from evading detection by using typos or unconventional words and frequently changing their terms, which results in the credit fraud intent vocabulary failing to include credit fraud intent words that may appear in the credit texts to be tested. Suspected credit fraud texts are then preliminarily screened based on the consistency and contextual connections between the credit texts to be tested containing credit fraud intent words. The likelihood of fraud within a user category is then determined based on the connections between the suspected credit fraud texts corresponding to each user within the user category. This method accurately identifies connections between user credit fraud behaviors and avoids misjudgments of legitimate users who may share information. Finally, based on the user's own record of poor credit behavior and the likelihood of fraud within the user category to which the user belongs, the degree to which the user qualifies for credit fraud is determined, thereby improving the accuracy of credit fraud detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 A flowchart of a method for processing credit text data based on NLP provided by one embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a process for screening suspected credit fraud texts provided by one embodiment of the present invention;
[0044] Figure 3 A schematic diagram of a process for alerting a user according to an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of the overall process of a credit text data processing method based on NLP provided by one embodiment of the present invention;
[0046] Figure 5 A schematic structural diagram of an NLP-based credit text data processing system provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0047] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of an NLP-based credit document data processing method and system proposed by the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0048] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0049] The following describes in detail a specific scheme of a credit text data processing method and system based on NLP provided by the present invention with reference to the accompanying drawings.
[0050] NLP (Natural Language Processing) is a key research area in the field of artificial intelligence. It integrates knowledge from multiple disciplines, including linguistics, computer science, machine learning, mathematics, and cognitive psychology. It is an interdisciplinary field that integrates computer science, artificial intelligence, and linguistics. It encompasses two main aspects: natural language understanding and natural language generation. Its research covers multiple levels, including characters, words, phrases, sentences, paragraphs, and texts. It serves as a bridge between machine language and human language. It aims to enable machines to understand, interpret, and generate human language, enabling effective communication between humans and machines and enabling computers to perform tasks such as language translation, sentiment analysis, and text summarization. Basic NLP techniques include text preprocessing, word embedding, syntactic analysis, semantic analysis, and text generation. Text preprocessing includes text cleaning, word segmentation, and part-of-speech tagging.
[0051] See also Figure 1 , which shows a flow chart of a method for processing credit text data based on NLP according to an embodiment of the present invention, including the following steps:
[0052] Step 101: segment multiple credit texts to be tested to obtain segmentation results.
[0053] The word segmentation results include multiple phrases in each single sentence in each credit text to be tested.
[0054] In one embodiment, the credit text to be tested may include at least one of credit application materials, historical credit records, interactive communication records, and third-party public data, etc. The credit application materials may include at least one of the credit application forms, statement documents, and identity document attachments submitted by users. The historical credit records may include at least one of historical loan contracts, credit report texts, and historical repayment remarks. The interactive communication records may include at least one of customer service conversation logs, online chat records, and correspondence emails. The third-party public data may include publicly proven fraud and credit default records.
[0055] In one embodiment, the credit text to be tested may be the credit text of a financial institution within the most recent first preset time period. For example, the credit text of a financial institution within the most recent quarter.
[0056] In one embodiment, the credit text to be tested may be first cleaned and standardized to screen out redundant and irrelevant information and unify the text formats from different sources. Then, the cleaned and standardized credit text to be tested is segmented to obtain a segmentation result. During the segmentation process, the credit text to be tested may be first clause-separated according to punctuation marks, and then the Jieba segmentation tool is used to segment and label the词性 of each sentence.
[0057] In one embodiment, the unique user fields may be screened, and the method of regular expressions is used for the unique user fields to match the credit texts to be tested belonging to the same user in each credit text to be tested.
[0058] Step 102, recombine each phrase in each single sentence in the segmentation result to obtain multiple recombined words.
[0059] It can be understood that due to the Jieba segmentation technology, there is a situation where a single complete term is segmented into multiple short phrases. For example, "package salary statement" is segmented into "package", "salary", "statement"; and fraudsters change their narrative methods with a high frequency and in various forms, and there is a situation where new uncollected variant phrases are generated. At this time, Jieba segmentation may also have the possibility of incorrect segmentation. Therefore, it is necessary to recombine each phrase in each single sentence in the segmentation result to obtain multiple recombined words, and then screen out credit fraud intention words from the multiple recombined words.
[0060] During recombination, the original adjacent relationship of each phrase in the single sentence is not changed, and only the number of phrases included in the recombined words is changed. For example, assume that the segmentation result of a certain single sentence is [a, b, c, d], where a, b, c, d are each a phrase. During recombination, the single phrase is sequentially recombined with the subsequent phrases, and the multiple recombined words obtained are: [a], [ab], [abc], [abcd], [b], [bc], [bcd], [c], [cd], [d].
[0061] Step 103 : establishing a set of credit fraud intention words based on the association between each recombined word and high-frequency fraud intention words, historical normal credit texts, and context.
[0062] In one embodiment, high-frequency fraud intent words are high-frequency phrases in historical credit fraud texts within the last second preset time period. For example, high-frequency phrases in historical credit fraud texts detected by financial institutions within the past six months are selected as high-frequency fraud intent words.
[0063] In one embodiment, historical credit texts in a financial institution other than historical credit fraud texts detected by the financial institution itself are regarded as historical normal credit texts.
[0064] It can be understood that if the frequency of a phrase in a historical credit fraud text is much higher than that of other words, the possibility that the phrase is a high-frequency fraud intent word is greater. Therefore, in one embodiment, for each phrase in a historical credit fraud text, if the frequency of the phrase in all historical credit fraud texts is greater than the average frequency of each phrase in all historical credit fraud texts, the phrase is a high-frequency fraud intent word. That is, When Phrases are marked as high-frequency fraud intent words. Among them, Indicates the The frequency of each phrase appearing in all historical credit fraud texts. Represents the mean frequency of each phrase in all historical credit fraud texts.
[0065] In one embodiment, the possibility of each recombined word being a credit fraud intention word is determined based on the association of each recombined word with high-frequency fraud intention words, historical normal credit texts, and context. Then, based on the possibility of each recombined word corresponding to a credit fraud intention word, the credit fraud intention word in each recombined word is determined. Based on the high-frequency fraud intention words and the credit fraud intention words in each recombined word, a set of credit fraud intention words is established.
[0066] In one embodiment, after establishing a set of credit fraud intent words, the word segmentation results of each credit document to be tested can be adjusted based on the set of credit fraud intent words. Entity annotation is then performed based on the word segmentation results in conjunction with a financial dictionary. Triples are extracted based on the entity annotation results. Step 104 and subsequent steps are then performed based on the adjusted word segmentation results and the extracted triples. Triples are used to represent logical relationships between entities.
[0067] For example, for a single sentence in the credit text to be tested, "The applicant applied for a 500,000 yuan business loan from a certain bank."
[0068] Word segmentation and part-of-speech tagging examples:
[0069] The applicant (nn) applied for a business loan (v) of 500,000 (m) from a certain bank (nt) (p) (nz)
[0070] Example of entity annotation using a financial dictionary:
[0071] Applicant (nn) / applied for (v) / 500,000 (m) / business loan (nz) # loan product from (p) / a certain bank (nt) # institution name
[0072] Example of a triple:
[0073] Subject: Applicant
[0074] Relationship: Application
[0075] Target: 500,000 yuan business loan
[0076] Step 104 , screening suspected credit fraud texts based on the consistency and contextual connection between the credit texts to be tested that contain credit fraud intention words.
[0077] In one embodiment, target credit texts to be tested containing credit fraud intention words are screened from various credit texts to be tested, and the degree of fuzziness of each credit fraud intention word appearing in each target credit text to be tested is determined. Based on the degree of fuzziness of each credit fraud intention word appearing in the target credit text to be tested, the number of target paragraphs containing credit fraud intention words in the target credit text to be tested, and the consistency between each target credit text to be tested, the degree of credit fraud of the target credit text to be tested is determined. Based on the degree of credit fraud corresponding to each target credit text to be tested, suspected credit fraud texts are screened from each target credit text to be tested.
[0078] Step 105 , for each user category, based on the connections between the suspected credit fraud texts corresponding to the users in the user category, determines the possibility of fraud in the user category.
[0079] In one embodiment, users may be classified according to their basic information to obtain the user category to which each user belongs. The basic user information may include at least one of the user's phone number, address, bank account number, and device identification.
[0080] Step 106 , based on the user's bad credit behavior record and the possibility of fraud in the user category to which the user belongs, determines the extent to which the user meets the credit fraud criteria.
[0081] In one embodiment, after determining the extent to which the user meets the credit fraud criteria, a priority value for alerting the user may be determined, and the user may be alerted based on the priority value.
[0082] In one embodiment, target users with priority values greater than a preset priority value threshold are determined, and alerts are issued to the target users for suspected credit fraud texts in descending order of priority values.
[0083] The NLP-based credit text data processing method reorganizes the individual phrases within each sentence in the word segmentation results of multiple credit texts to be tested to obtain multiple reorganized words. Based on the association of each reorganized word with high-frequency fraud intent words, historical normal credit texts, and context, a set of credit fraud intent words can be accurately established. Compared to manually defining a credit fraud intent vocabulary, this method can prevent fraudsters from evading detection by using typos or unconventional words and frequently changing their terms, which results in the credit fraud intent vocabulary failing to encompass the credit fraud intent words that may appear in the credit texts to be tested. Suspected credit fraud texts are then initially screened based on the consistency and contextual connections between the credit texts to be tested containing credit fraud intent words. The likelihood of fraud within a user category is then determined by combining the connections between the suspected credit fraud texts corresponding to each user within the user category. This method accurately identifies connections between user credit fraud behaviors and avoids misjudgments of legitimate users who may share information. Finally, based on the user's own record of poor credit behavior and the likelihood of fraud within the user category to which they belong, the degree to which the user qualifies as a credit fraud suspect is determined, thereby improving the accuracy of credit fraud detection.
[0084] In one embodiment, a set of credit fraud intent words is established based on the association of each recombined word with a high-frequency fraud intent word, a historical normal credit text, and a context, including: for each recombined word, based on the similarity of the recombined word with each high-frequency fraud intent word, the frequency of its appearance in the historical normal credit text, the similarity with each high-frequency phrase in the historical normal credit text, and the similarity with the remaining words in the credit text to be tested except the recombined word, determining the credit fraud intent word in each recombined word based on the likelihood of each recombined word corresponding to the credit fraud intent word; and establishing a set of credit fraud intent words based on the high-frequency fraud intent words and the credit fraud intent words in each recombined word.
[0085] In one embodiment, the method for determining high-frequency phrases in historically normal credit documents is similar to the method for determining high-frequency fraudulent intent words in historically fraudulent credit documents. For each phrase in historically normal credit documents, if the frequency of its occurrence in all historically normal credit documents is greater than the average frequency of all other phrases in all historically normal credit documents, the phrase is considered a high-frequency phrase in historically normal credit documents.
[0086] It can be understood that the higher the similarity between a recombined word and high-frequency fraudulent intent words, the lower the similarity with high-frequency phrases in historically normal credit texts, or the lower the frequency of its occurrence in historically normal credit texts, the greater the likelihood that the recombined word is a fraudulent intent word. Furthermore, since fraudulent intent words are often variant phrases with altered narrative styles, they are not closely connected to the context. Therefore, the lower the similarity between a recombined word and the rest of the words in the test credit text, the greater the likelihood that the recombined word is a fraudulent intent word.
[0087] In one embodiment, the likelihood that a recombined word is a fraudulent term is positively correlated with the similarity between the recombined word and various high-frequency fraudulent terms. The likelihood that a recombined word is a fraudulent term is negatively correlated with the frequency of the recombined word in historically normal credit documents. The likelihood that a recombined word is a fraudulent term is negatively correlated with the similarity between the recombined word and various high-frequency phrases in historically normal credit documents. The likelihood that a recombined word is a fraudulent term is negatively correlated with the similarity between the recombined word and the remaining words in the credit document to be tested, excluding the recombined word.
[0088] In one embodiment, the similarity in each embodiment of the present invention can be obtained by calculating the cosine similarity between word vectors.
[0089] In one embodiment, the average of the similarities between the recombined word and each high-frequency fraudulent intention word, the average of the similarities between the recombined word and each high-frequency phrase in the historical normal credit text, and the average of the similarities between the recombined word and the remaining words in the credit text to be tested except the recombined word can be determined, and the product of the frequency of occurrence of the recombined word in the historical normal credit text, the average of the similarities between the recombined word and each high-frequency fraudulent intention word, and the average of the similarities between the recombined word and the remaining words in the credit text to be tested except the recombined word can be calculated. Based on the ratio of the average of the similarities between the recombined word and each high-frequency fraudulent intention word to the product, the possibility that the recombined word is a credit fraud intention word is determined.
[0090] In one embodiment, the possibility that a recombined word belongs to a credit fraud intent word can be determined according to the following formula:
[0091]
[0092] in, Indicates the The possibility that the recombined word belongs to a credit fraud intention word. Indicates the The average similarity between the recombined words and each high-frequency fraud intent word. Indicates the The frequency of occurrence of the recombined words in historical normal credit texts. Indicates the The average similarity between the recombined words and each high-frequency phrase in the historical normal credit text. Indicates the The average similarity between the recombined word and the remaining words in the credit text to be tested except the recombined word. is a normalization function with a range of (0, 1). It should be noted that to ensure meaningful calculation results, when performing fractional operations in the embodiments of the present invention, if the denominator is 0, a parameter adjustment factor greater than 0 must be added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer based on actual conditions and is not specifically limited in this application.
[0093] In one embodiment, if the recombined word is a credit fraud intent word, If the probability is greater than a preset threshold, the recombined word is determined to be a credit fraud intent word. Based on the high-frequency fraud intent words and the credit fraud intent words in each recombined word, a set of credit fraud intent words is established. In one embodiment, the preset probability threshold can be set to 0.5.
[0094] Understandably, traditional methods for obtaining a collection of credit fraud intent words are typically based on historical cases or publicly available fraud keywords. However, these methods may not fully cover variant words and homophonic words, and fraudsters often intentionally change narrative styles or use typos to evade detection. Therefore, it is necessary to analyze the likelihood that a particular word is a detection-avoiding term based on the differences in the frequency of occurrence of each phrase in the credit text to be tested, thereby determining a collection of credit fraud intent words.
[0095] Therefore, in the above embodiment, for each recombined word, the possibility that the recombined word belongs to a credit fraud intention word can be accurately determined based on the similarity between the recombined word and each high-frequency fraud intention word, the frequency of its appearance in historical normal credit texts, the similarity with each high-frequency phrase in historical normal credit texts, and the similarity with the remaining words in the credit text to be tested except the recombined word. According to the possibility that each recombined word belongs to a credit fraud intention word, the credit fraud intention word in each recombined word can be accurately determined. Based on the high-frequency fraud intention words and the credit fraud intention words in each recombined word, a set of credit fraud intention words can be accurately established. Compared with the method of manually defining a credit fraud intention word library, this method can avoid the problem that fraudsters use typos or unconventional words to evade detection and frequently change terms, resulting in the credit fraud intention word library being unable to include credit fraud intention words that may appear in the credit text to be tested.
[0096] In one embodiment, see Figure 2, based on the consistency and contextual connection between the credit texts to be tested that contain credit fraud intent words, screening suspected credit fraud texts includes the following steps:
[0097] Step 201 : Filter target credit texts to be tested that contain credit fraud intention words from various credit texts to be tested.
[0098] Step 202: Determine the fuzziness level of each credit fraud intention word appearing in each target credit document to be tested.
[0099] The fuzziness level represents the degree of semantic fuzziness of the credit fraud intention words.
[0100] Step 203, for each target credit text to be tested, determine the degree of credit fraud of the target credit text to be tested based on the growth of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text, the fuzziness of each credit fraud intention word appearing in the target credit text to be tested, the number of target paragraphs containing credit fraud intention words in the target credit text to be tested, and the frequency of the triples in the target credit text to be tested appearing in the other target credit texts to be tested of the corresponding user.
[0101] In one embodiment, the growth of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text to be tested includes the growth value of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text to be tested.
[0102] In one embodiment, the growth value can be determined based on the ratio of the frequency of credit fraud intent words appearing in the target credit text to the frequency of credit fraud intent words appearing in the target credit text last published by the user corresponding to the target credit text. The formula is as follows:
[0103]
[0104] in, Represents the growth value of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the z-th target credit text to be tested. Represents the frequency of credit fraud intention words in the z-th target credit text to be tested. Indicates the frequency of credit fraud intention words appearing in the z-1th target credit text to be tested posted by the user.
[0105] It can be understood that for a single user, the more target paragraphs there are in the target credit text to be tested corresponding to the user, the greater the frequency growth trend of credit fraud intent words appearing in the target credit text to be tested corresponding to the user compared to the credit text previously published by the user, the higher the degree of vague descriptions appearing in the target credit text to be tested corresponding to the user, or the lower the similarity between the target credit texts to be tested corresponding to the user, the greater the possibility that the target credit text to be tested contains credit fraud. Among them, since the growth value of the frequency of credit fraud intent words appearing in the target credit text to be tested published by the user corresponding to the target credit text indicates that the greater the frequency growth trend of credit fraud intent words appearing in the target credit text to be tested compared to the credit text previously published by the user, the greater the possibility that the target credit text to be tested contains credit fraud, therefore, the degree of credit fraud in the target credit text to be tested is positively correlated with the growth value of the frequency of credit fraud intent words appearing in the target credit text to be tested published by the user corresponding to the target credit text. Since the higher the frequency of the triples in the target credit text to be tested appearing in the remaining target credit texts to be tested of the corresponding user, the higher the similarity between the target credit texts to be tested corresponding to the user, and the lower the possibility of credit fraud in the target credit text to be tested, the degree of credit fraud of the target credit text to be tested is negatively correlated with the frequency of the triples in the target credit text to be tested appearing in the remaining target credit texts to be tested of the corresponding user.
[0106] Therefore, in one embodiment, the degree of credit fraud of the target credit text to be tested is positively correlated with the growth value of the frequency of credit fraud intent words appearing in the target credit text to be tested published by the user corresponding to the target credit text to be tested. The degree of credit fraud of the target credit text to be tested is positively correlated with the degree of ambiguity of each credit fraud intent word appearing in the target credit text to be tested. The degree of credit fraud of the target credit text to be tested is positively correlated with the number of target paragraphs containing credit fraud intent words in the target credit text to be tested. The degree of credit fraud of the target credit text to be tested is negatively correlated with the frequency of appearance of triples in the target credit text to be tested in the remaining target credit texts to be tested of the corresponding user.
[0107] In one embodiment, the product of the growth value of the frequency of credit fraud intent words appearing in the target credit text published by the corresponding user is calculated, the average value of the fuzziness level of each credit fraud intent word appearing in the target credit text, and the number of target paragraphs in the target credit text containing credit fraud intent words is calculated. The credit fraud degree of the target credit text is determined based on the ratio of this product to the average value of the frequency of the triples in the target credit text appearing in the remaining target credit texts of the corresponding user. The formula is as follows:
[0108]
[0109] in, Indicates the credit fraud degree of the zth target credit text to be tested. Represents the growth value of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the z-th target credit text to be tested. It represents the mean value of the fuzziness of each credit fraud intention word appearing in the z-th target credit text to be tested. The number of target paragraphs containing credit fraud intent words in the z-th target credit text to be tested. The average frequency of the triples in the z-th target credit text to be tested in the remaining target credit texts to be tested of the corresponding user.
[0110] Step 204 : Screening suspected credit fraud texts from the target credit texts according to the credit fraud degrees corresponding to the target credit texts to be tested.
[0111] In one embodiment, the credit fraud degree corresponding to each target credit document to be tested is normalized. If the normalized credit fraud degree is greater than a preset credit fraud degree threshold, the corresponding target credit document to be tested is determined to be a suspected credit fraud document. The preset credit fraud degree threshold can be set to 0.5.
[0112] In the above embodiment, since the multi-source text data in normal credit texts describes the same facts, the similarity between the multi-source text data is high, and the frequency of credit fraud terms appearing in the text is low. However, the multi-source data in fraud texts may contain narrative conflicts or contain a large number of ambiguous narratives to evade detection. Moreover, fraudsters often update their speech templates at a high frequency, resulting in a sudden increase in the frequency of fraudulent terms in recently submitted credit texts. Suspected credit fraud texts can be preliminarily screened based on the above logic. Therefore, based on the increase in the frequency of credit fraud intent terms appearing in the target credit texts published by the user corresponding to the target credit text, the degree of ambiguity of each credit fraud intent term appearing in the target credit text, the number of target paragraphs containing credit fraud intent terms in the target credit text, and the frequency of occurrence of triples in the target credit text in the corresponding user's other target credit texts, the degree of credit fraud of the target credit text can be accurately determined. Then, based on the credit fraud degree corresponding to each target credit text, suspected credit fraud texts can be accurately screened from each target credit text.
[0113] In one embodiment, the fuzziness level of each credit fraud intention word appearing in each target credit text to be tested is determined, including: for each credit fraud intention word appearing in the target credit text to be tested, the fuzziness level of the credit fraud intention word is determined based on the difference in similarity between the credit fraud intention word and normal words in each paragraph in which it is located, and the similarity between the word vectors corresponding to the credit fraud intention word in each paragraph.
[0114] In one embodiment, for each credit fraud intention word that appears in the target credit text to be tested, the paragraphs containing the credit fraud intention word are input into the MobileBERT model (a lightweight BERT model) to obtain the word vector of the credit fraud intention word in each paragraph.
[0115] It can be understood that semantically ambiguous words often exhibit polysemy or semantic drift in different contexts. Therefore, for each credit fraud intent word appearing in the target credit document, the greater the difference in similarity between the credit fraud intent word and the remaining words in different paragraphs, the more obvious the semantic differences between the credit fraud intent word in different contexts, and thus the higher the degree of ambiguity of the credit fraud intent word. The smaller the mean similarity between the word vectors corresponding to the credit fraud intent word in each paragraph, the greater the difference in the word vectors of the credit fraud intent word in each paragraph, and thus the higher the degree of ambiguity of the credit fraud intent word.
[0116] Therefore, in one embodiment, the ambiguity of a credit fraud intent word is positively correlated with the difference in similarity between the credit fraud intent word and normal words in each paragraph. The ambiguity of a credit fraud intent word is negatively correlated with the mean similarity between the word vectors corresponding to the credit fraud intent word in each paragraph.
[0117] In one embodiment, for each credit fraud intent word appearing in the target credit document to be tested, the fuzziness of the credit fraud intent word is determined based on the ratio of the extreme difference in similarity between the credit fraud intent word and normal words in each paragraph in which it is located to the mean similarity between the word vectors corresponding to the credit fraud intent word in each paragraph. The formula is as follows:
[0118]
[0119] in, Indicates the fuzziness of the kth credit fraud intention word in the zth target credit text to be tested. It represents the extreme difference in similarity between the kth credit fraud intention word in the zth target credit text to be tested and the normal words in each paragraph. It represents the mean similarity between the word vectors corresponding to the k-th credit fraud intention word in each paragraph in the z-th target credit text to be tested.
[0120] In the above embodiment, for each credit fraud intention word appearing in the target credit text to be tested, the degree of ambiguity of the credit fraud intention word can be accurately determined based on the difference in similarity between the credit fraud intention word and the normal words in the respective paragraphs in which it is located, and the similarity between the word vectors corresponding to the credit fraud intention word in the respective paragraphs.
[0121] In one embodiment, the possibility of fraud in a user category is determined based on the connection between each suspected credit fraud text corresponding to each user in the user category, including: determining the possibility that each triple in each suspected credit fraud text in the user category belongs to a fixed template of the user category; determining the fixed triples of the user category based on the possibility that each triple in the user category corresponds to the fixed template of the user category; determining the possibility of fraud in the user category based on the number of users in the user category, the number of fixed triples of the user category, the degree of credit fraud of each suspected credit fraud text in the user category, and the time interval between the release of each suspected credit fraud text in the user category.
[0122] In one embodiment, the likelihood of each triplet in a user category corresponding to a fixed template belonging to the user category is normalized. If the normalized likelihood of the fixed template belonging to the user category is greater than a preset fixed template likelihood threshold, the corresponding triplet is determined as the fixed triplet for the user category. The preset fixed template likelihood threshold can be set to 0.5.
[0123] It's understandable that since the number of users associated with normal users is realistic and limited, while fraudulent activity generally doesn't follow this pattern, the more users a user category contains, the greater the likelihood that the user category is fraudulent. The greater the number of fixed triples in a user category, the greater the likelihood that the user category is fraudulent. The higher the credit fraud level of each suspected credit fraud text within a user category, the greater the likelihood that the user category is fraudulent. The shorter the time interval between the publication of each suspected credit fraud text within a user category, the more concentrated the frequency of suspected credit fraud text publication within the user category, and the greater the likelihood that the user category is fraudulent.
[0124] Therefore, in one embodiment, the likelihood of fraud in a user category is positively correlated with the number of users in the user category, the number of fixed triples in the user category, and the degree of credit fraud in each suspected credit fraud text in the user category. The likelihood of fraud in a user category is negatively correlated with the time interval between the publication of each suspected credit fraud text in the user category.
[0125] In one embodiment, the product of the number of users in a user category, the number of fixed triples in the user category, and the mean credit fraud degree of each suspected credit fraud text in the user category is calculated. The likelihood of fraud in the user category is determined based on the ratio of this product to the mean time interval between the publication of each suspected credit fraud text in the user category. The formula is as follows:
[0126]
[0127] in, Indicates the possibility that the user category to which the w-th user belongs is fraudulent. Indicates the number of users in the user category to which the w-th user belongs. represents the number of fixed triplets in the user category to which the w-th user belongs. represents the mean of the credit fraud degree of each suspected credit fraud text in the user category to which the w-th user belongs. Represents the mean time interval between the release of each suspected credit fraud text in this user category.
[0128] Because fraudulent activity typically involves a single core operating node (e.g., a fraudulent user) linking numerous accounts and using fixed templates to generate corresponding fraudulent text, and legitimate users (e.g., family users) may also have multiple accounts linked, traditional methods struggle to distinguish between them. Therefore, based on differences in user credit behavior, if a user posts credit texts with similar structures at similar times as multiple other users (i.e., a high likelihood of a fixed template), the posted texts are highly consistent with credit fraud, and the multiple users who posted the credit texts are highly connected, then the likelihood of fraudulent activity is high.
[0129] Therefore, in the above embodiment, the possibility that each triple in each suspected credit fraud text in the user category belongs to the fixed template of the user category is determined. According to the possibility that each triple in the user category corresponds to the fixed template of the user category, the fixed triple of the user category can be accurately determined. Then, according to the number of users in the user category, the number of fixed triples of the user category, the degree of credit fraud of each suspected credit fraud text in the user category, and the time interval between the release of each suspected credit fraud text in the user category, the possibility of fraud in the user category can be accurately determined.
[0130] In one embodiment, determining the possibility that each triple in each suspected credit fraud text in a user category belongs to a fixed template of the user category includes: for each triple in each suspected credit fraud text in the user category, determining the possibility that the triple belongs to the fixed template of the user category based on the frequency of occurrence of the triple in each suspected credit fraud text in the user category, the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text in the user category and the degree of discreteness of the number of characters, and the difference in the proportion of single sentences of the triple in each suspected credit fraud text in the user category.
[0131] The spacing position is the position between the elements of the triples in a single sentence of the suspected credit fraud text. For example, a single sentence may be [triplet element 1, spacing position 1, triplet element 2, spacing position 2, triplet element 3].
[0132] In one embodiment, the dispersion degree may be a standard deviation or a variance, etc. The difference may be a range.
[0133] In one embodiment, the proportion of triples in each sentence in each suspected credit fraud text of the user category can be determined based on the ratio between the total number of words in all elements in the triple and the total number of words in the sentence in which the triple is located. The formula is as follows:
[0134]
[0135] in, It represents the percentage of the j-th triple in a single sentence in each suspected credit fraud text of the user category. Represents the total number of words in all elements of the j-th triple. Indicates the total number of words in the sentence where the j-th triple is located.
[0136] In one embodiment, the triples in all suspected credit fraud texts may be traversed in descending order of frequency of occurrence, and the likelihood of each triple belonging to a fixed template of a user category may be determined.
[0137] It can be understood that the smaller the difference in the proportion of triples in each sentence in each suspected credit fraud text of the user category, the more similar the proportion of triples in the corresponding sentences, and the greater the possibility that the triples belong to the fixed template of the user category. The smaller the dispersion of the number of characters between the phrases at the corresponding interval positions of the triples in each suspected credit fraud text of the user category, and the greater the similarity between the phrases at the interval positions, the smaller the fluctuation in the number of characters at the interval position, and the closer the semantics of the phrases at the interval position, the greater the possibility that the triples belong to the fixed template of the user category. The greater the frequency of occurrence of the triples in each suspected credit fraud text of the user category, the greater the possibility that the triples belong to the fixed template of the user category.
[0138] Therefore, in one embodiment, the likelihood that a triple belongs to a fixed template for a user category is positively correlated with the frequency of the triple's appearance in each suspected credit fraud text for the user category. The likelihood that a triple belongs to a fixed template for a user category is positively correlated with the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text for the user category. The likelihood that a triple belongs to a fixed template for a user category is negatively correlated with the difference in the proportion of single sentences of the triple in each suspected credit fraud text for the user category. The likelihood that a triple belongs to a fixed template for a user category is negatively correlated with the degree of dispersion of the number of characters at the corresponding interval positions of the triple in each suspected credit fraud text for the user category.
[0139] In one embodiment, a first product is calculated between the frequency of occurrence of a triple in each suspected credit fraud text of the user category and the mean similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text of the user category. A second product is calculated between the range of the percentage of single sentences of the triple in each suspected credit fraud text of the user category and the standard deviation of the number of characters at the corresponding interval positions of the triple in each suspected credit fraud text of the user category. Based on the ratio of the first product to the second product, the probability that the triple belongs to the fixed template of the user category is determined. The formula is as follows:
[0140]
[0141] in, represents the likelihood that the jth triplet belongs to the fixed template of the user category. represents the frequency of occurrence of the j-th triple in each suspected credit fraud text of the user category. represents the mean similarity between the phrases at the corresponding interval positions of the j-th triple in each suspected credit fraud text of the user category. It represents the extreme difference of the proportion of single sentences of the j-th triple in each suspected credit fraud text of the user category. represents the standard deviation of the number of characters at the interval position corresponding to the j-th triple in each suspected credit fraud text of the user category.
[0142] In the above embodiment, for each triple in each suspected credit fraud text in the user category, the possibility that the triple belongs to the fixed template of the user category can be accurately determined based on the frequency of occurrence of the triple in each suspected credit fraud text in the user category, the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text in the user category and the discrete degree of the number of characters, and the difference in the proportion of single sentences of the triple in each suspected credit fraud text in the user category.
[0143] In one embodiment, the degree to which a user complies with credit fraud is determined based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs, including: determining the ratio between the user's own historical bad repayment record times and the historical bad repayment record times of the user corresponding to each credit text to be tested, to obtain the user's own bad credit behavior rate; determining the degree to which the user complies with credit fraud based on the user's own bad credit behavior rate and the possibility of fraud in the user category to which the user belongs.
[0144] In one embodiment, the degree to which a user is eligible for credit fraud is positively correlated with the user's own bad credit behavior rate. The degree to which a user is eligible for credit fraud is positively correlated with the possibility of fraud in the user category to which the user belongs.
[0145] In one embodiment, the degree to which a user qualifies for credit fraud can be determined based on the product of the user's bad credit behavior rate and the fraud probability of the user category to which the user belongs. The formula is as follows:
[0146]
[0147] in, Indicates the degree to which the wth user meets the credit fraud requirement. Indicates the number of bad repayment records of the w-th user. Represents the average number of historical bad repayment records of users corresponding to all credit documents to be tested. It indicates the possibility of fraud in the user category to which the wth user belongs.
[0148] In the above embodiment, since the greater the number of bad repayment records a user has, the greater the likelihood of fraud in the user category to which the user belongs, and the greater the degree to which the user qualifies for credit fraud, the user's bad credit behavior rate is obtained by determining the ratio of the number of bad repayment records a user has to the number of bad repayment records of the user corresponding to each credit document to be tested. Based on the user's bad credit behavior rate and the likelihood of fraud in the user category to which the user belongs, the degree to which the user qualifies for credit fraud can be accurately determined. Through the above embodiment, the degree to which the user corresponding to each suspected credit fraud document qualifies for credit fraud can be determined.
[0149] In one embodiment, see Figure 3 After determining the extent to which the user meets the criteria for credit fraud based on the user's bad credit behavior record and the possibility of fraud in the user category to which the user belongs, the method further includes the following steps:
[0150] Step 301 : Determine a priority value for warning the user based on the degree to which the user meets the credit fraud criteria, the number of suspected credit fraud texts posted by the user, and the time interval between the user's last posting of a suspected credit fraud text and the present time.
[0151] In one embodiment, the priority value for alerting a user is positively correlated with the degree to which the user meets the criteria for credit fraud. The priority value for alerting a user is positively correlated with the number of texts suspected of credit fraud posted by the user. The priority value for alerting a user is negatively correlated with the time since the user last posted a text suspected of credit fraud.
[0152] In one embodiment, the priority value for alerting the user can be obtained by multiplying the degree to which the user meets the credit fraud criteria by the number of suspected credit fraud texts posted by the user, and then dividing the product by the time since the user last posted a suspected credit fraud text. The formula is as follows:
[0153]
[0154] in, Indicates the priority value for alerting the w-th user. Indicates the degree to which the wth user meets the credit fraud requirement. Indicates the number of suspected credit fraud texts posted by the user. The time since the user last posted a text suspected of credit fraud.
[0155] Step 302: alert the user of the suspected credit fraud text based on the priority value.
[0156] In one embodiment, the priority values of each user may be normalized, and the user's suspected credit fraud text may be warned based on the normalized priority value.
[0157] In one embodiment, a target user whose normalized priority value is greater than a preset priority value threshold may be identified, and a warning may be issued to the target user regarding the text suspected of credit fraud. The preset priority value threshold may be set to 0.5.
[0158] In the above embodiment, the priority value for warning the user can be accurately determined based on the degree to which the user meets the requirements of credit fraud, the number of suspected credit fraud texts posted by the user, and the length of time since the user last posted a suspected credit fraud text, thereby accurately warning the user of suspected credit fraud texts.
[0159] In one embodiment, alerting users of suspected credit fraud texts based on priority values includes: determining target users whose priority values are greater than a preset priority value threshold; and alerting each target user of suspected credit fraud texts in descending order of priority values.
[0160] In one embodiment, the warning information of each target user and the suspected credit fraud text of the target user may be output to the terminal used by the staff in descending order of priority value.
[0161] In the above embodiment, target users whose priority values are greater than a preset priority value threshold are determined, and alerts are issued to the suspected credit fraud texts of each target user in descending order of priority values. This allows for more orderly alerts and improves the accuracy of alerts.
[0162] like Figure 4 As shown, the present invention provides an overall flow chart of a credit text data processing method based on NLP, which includes the following steps: first, obtaining the user's multi-source credit text data (i.e., the credit text to be tested) and preprocessing it, the preprocessing including cleaning, standardization, word segmentation and triple extraction; then, the segmented phrases are reorganized, and a credit fraud intention vocabulary set (i.e., a set of credit fraud intention words) is established based on the degree of association between the reorganized words and high-frequency fraud intention words; then, based on the consistency and contextual connection between the multi-source credit texts, the text data suspected of credit fraud (i.e., suspected credit fraud text) is preliminarily screened; then, based on the preliminary screening results and combined with the connection between the user's credit behaviors, the possibility of the user committing credit fraud is analyzed; based on the possibility of credit fraud, the alarm priority corresponding to the user (i.e., the priority value for alerting the user) is determined; finally, the credit fraud text is alarmed and the corresponding user information is output.
[0163] See Figure 5The present invention provides a credit text data processing system based on NLP, the system includes a memory and a processor; the memory is used to store executable program code; the processor is used to call and run the executable program code from the memory to implement the following steps: segmenting multiple credit texts to be tested to obtain segmentation results; reorganizing each phrase in each single sentence in the segmentation results to obtain multiple reorganized words; establishing a set of credit fraud intention words based on the association of each reorganized word with high-frequency fraud intention words, historical normal credit texts, and context; screening suspected credit fraud texts based on the consistency and contextual connection between each credit text to be tested containing credit fraud intention words; determining the possibility of fraud in each user category based on the connection between each suspected credit fraud text corresponding to each user in the user category; determining the degree to which the user meets the credit fraud requirement based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs.
[0164] In one embodiment, the processor further implements the following steps: for each recombined word, the possibility that the recombined word is a credit fraud intention word is determined based on the similarity between the recombined word and each high-frequency fraud intention word, the frequency of appearance in historical normal credit texts, the similarity with each high-frequency phrase in historical normal credit texts, and the similarity with the remaining words in the credit text to be tested except the recombined word; the credit fraud intention word in each recombined word is determined based on the possibility that each recombined word belongs to a credit fraud intention word; and a set of credit fraud intention words is established based on the high-frequency fraud intention words and the credit fraud intention words in each recombined word.
[0165] In one embodiment, the processor further implements the following steps: screening target credit texts to be tested that contain credit fraud intention words from each credit text to be tested; determining the degree of fuzziness of each credit fraud intention word appearing in each target credit text to be tested; determining the degree of credit fraud of the target credit text to be tested for each target credit text to be tested based on the growth of the frequency of credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text, the degree of fuzziness of each credit fraud intention word appearing in the target credit text to be tested, the number of target paragraphs containing credit fraud intention words in the target credit text to be tested, and the frequency of the triples in the target credit text to be tested in the remaining target credit texts to be tested of the corresponding user; screening suspected credit fraud texts from each target credit text to be tested based on the degree of credit fraud corresponding to each target credit text.
[0166] In one embodiment, the processor further implements the following steps: for each credit fraud intention word appearing in the target credit text to be tested, the degree of ambiguity of the credit fraud intention word is determined based on the difference in similarity between the credit fraud intention word and normal words in each paragraph in which it is located, and the similarity between the word vectors corresponding to the credit fraud intention word in each paragraph.
[0167] In one embodiment, the processor further implements the following steps: determining the possibility that each triple in each suspected credit fraud text in the user category belongs to the fixed template of the user category; determining the fixed triples of the user category based on the possibility that each triple in the user category corresponds to the fixed template of the user category; determining the possibility that fraud exists in the user category based on the number of users in the user category, the number of fixed triples of the user category, the degree of credit fraud of each suspected credit fraud text in the user category, and the time interval between the release of each suspected credit fraud text in the user category.
[0168] In one embodiment, the processor further implements the following steps: for each triple in each suspected credit fraud text in the user category, determine the possibility that the triple belongs to the fixed template of the user category based on the frequency of occurrence of the triple in each suspected credit fraud text in the user category, the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text in the user category and the discrete degree of the number of characters, and the difference in the proportion of single sentences of the triple in each suspected credit fraud text in the user category.
[0169] In one embodiment, the processor further implements the following steps: determining the ratio between the number of historical bad repayment records of the user himself and the number of historical bad repayment records of the user corresponding to each credit text to be tested, and obtaining the user's own bad credit behavior rate; determining the degree to which the user meets the credit fraud requirement based on the user's own bad credit behavior rate and the possibility of fraud in the user category to which the user belongs.
[0170] In one embodiment, the processor further implements the following steps: determining a priority value for alerting the user based on the degree to which the user meets the criteria for credit fraud, the number of suspected credit fraud texts posted by the user, and the length of time since the user last posted a suspected credit fraud text; and alerting the user to the suspected credit fraud text based on the priority value.
[0171] In one embodiment, the processor further implements the following steps: determining target users whose priority values are greater than a preset priority value threshold; and alerting each target user of suspected credit fraud texts in descending order of priority values.
[0172] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0173] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A credit text data processing method based on NLP, characterized in that: The method comprises: Perform word segmentation on multiple credit texts to be tested to obtain word segmentation results; Recombining each phrase in each single sentence in the word segmentation result to obtain multiple recombined words; Establishing a set of credit fraud intent words based on the association between each of the recombined words and high-frequency fraud intent words, historical normal credit texts, and context; Screening suspected credit fraud texts based on the consistency and contextual connection between the credit texts to be tested that contain the credit fraud intention words; For each user category, determining the possibility of fraud in the user category based on the relationship between the suspected credit fraud texts corresponding to the users in the user category; Determine the extent to which the user meets the criteria for credit fraud based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs; The screening of suspected credit fraud texts based on the consistency and contextual connection between the credit texts to be tested containing the credit fraud intention words includes: screening target credit texts to be tested containing the credit fraud intention words from the credit texts to be tested; determining the fuzziness level of each of the credit fraud intention words appearing in each of the target credit documents to be tested; For each target credit text to be tested, determine the credit fraud degree of the target credit text to be tested based on the growth of the frequency of the credit fraud intention words appearing in the target credit text to be tested published by the user corresponding to the target credit text to be tested, the fuzziness of each credit fraud intention word appearing in the target credit text to be tested, the number of target paragraphs containing the credit fraud intention words in the target credit text to be tested, and the frequency of the triples in the target credit text to be tested in the remaining target credit texts to be tested of the corresponding user; Based on the credit fraud degrees corresponding to the target credit texts to be tested, suspected credit fraud texts are screened from the target credit texts to be tested.
2. The NLP-based credit text data processing method according to claim 1, characterized in that: The method of establishing a set of credit fraud intent words based on the association between each of the recombined words and high-frequency fraud intent words, historical normal credit texts, and context includes: For each of the recombined words, determine the likelihood that the recombined word is a credit fraud intent word based on its similarity to various high-frequency fraud intent words, its frequency of appearance in historical normal credit texts, its similarity to various high-frequency phrases in the historical normal credit texts, and its similarity to the remaining words in the credit text to be tested, excluding the recombined word. Determining the credit fraud intention words in each of the recombined words based on the likelihood that each of the recombined words corresponds to the credit fraud intention word; A set of credit fraud intention words is established based on the high-frequency fraud intention words and the credit fraud intention words in each of the reorganized words.
3. The NLP-based credit text data processing method according to claim 1, characterized in that: Determining the fuzziness level of each credit fraud intention word appearing in each target credit text to be tested includes: For each credit fraud intention word appearing in the target credit text to be tested, the fuzziness of the credit fraud intention word is determined based on the difference in similarity between the credit fraud intention word and normal words in each paragraph, and the similarity between the word vectors corresponding to the credit fraud intention word in each paragraph.
4. The NLP-based credit text data processing method according to claim 1, characterized in that: Determining the possibility of fraud in the user category based on the connection between the suspected credit fraud texts corresponding to the users in the user category includes: determining the likelihood that each triple in each of the suspected credit fraud texts in the user category belongs to a fixed template of the user category; determining a fixed triplet of the user category according to the likelihood that each triplet in the user category corresponds to a fixed template of the user category; The possibility of fraud in the user category is determined based on the number of users in the user category, the number of fixed triples in the user category, the degree of credit fraud of each suspected credit fraud text in the user category, and the time interval between the release of each suspected credit fraud text in the user category.
5. The NLP-based credit text data processing method according to claim 4, characterized in that: Determining the possibility that each triple in each suspected credit fraud text in the user category belongs to the fixed template of the user category includes: For each triple in each suspected credit fraud text in the user category, determine the possibility that the triple belongs to the fixed template of the user category based on the frequency of occurrence of the triple in each suspected credit fraud text in the user category, the similarity between the phrases at the corresponding interval positions of the triple in each suspected credit fraud text in the user category and the discrete degree of the number of characters, and the difference in the proportion of single sentences of the triple in each suspected credit fraud text in the user category.
6. The NLP-based credit text data processing method according to claim 1, characterized in that: Determining the extent to which a user meets the criteria for credit fraud based on the user's own bad credit behavior record and the possibility of fraud in the user category to which the user belongs includes: Determine the ratio between the number of bad repayment records in the history of the user and the number of bad repayment records in the history of the user corresponding to each of the credit documents to be tested, and obtain the bad credit behavior rate of the user; The degree to which the user meets the credit fraud criteria is determined based on the user's own bad credit behavior rate and the possibility of fraud in the user category to which the user belongs.
7. The NLP-based credit text data processing method according to claim 1, characterized in that: After determining the extent to which the user meets the criteria for credit fraud based on the user's bad credit behavior record and the possibility of fraud in the user category to which the user belongs, the method further includes: Determining a priority value for alerting the user based on the degree to which the user meets the criteria for credit fraud, the number of suspected credit fraud texts posted by the user, and the time since the user last posted the suspected credit fraud text; According to the priority value, the user is warned of the suspected credit fraud text.
8. The NLP-based credit text data processing method according to claim 7, characterized in that: The step of alerting the user of the suspected credit fraud text according to the priority value includes: Determine the target user whose priority value is greater than a preset priority value threshold; The suspected credit fraud text of each target user is warned in descending order of the priority value.
9. A credit text data processing system based on NLP, characterized in that: The system includes a memory and a processor; the memory is used to store executable program code; the processor is used to call and run the executable program code from the memory to implement the NLP-based credit text data processing method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Individual behavior modeling and fraud detection method for low-frequency transactions
CN111242744A
Terminal fraud phone recognition method based on call text word vector
CN111669757A