Abnormal address identification method and device, electronic equipment and storage medium

By obtaining and analyzing the terms and part of speech of the address, and generating feature vectors to identify abnormal information, the problem of the inability to understand address text in the prior art is solved, the accuracy of abnormal address recognition is improved, and intelligent anti-fraud is realized.

CN120223331APending Publication Date: 2025-06-27JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311798481.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When identifying abnormal addresses, the prior art cannot truly understand the address text, resulting in poor recognition effects and inability to achieve intelligent anti-fraud.

Method used

By obtaining the part of speech of the candidate address and the words, a feature vector is generated, and an exception information is determined, thereby identifying the exception address.

Benefits of technology

It improves the accuracy of abnormal address recognition and achieves the purpose of intelligent anti-fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223331A_ABST
    Figure CN120223331A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal address recognition method and device, electronic equipment and a storage medium, and the abnormal address recognition method comprises the steps: obtaining a word of each candidate address in address data and the part of speech of the word; generating a feature vector of the candidate address based on the word of the candidate address and the part of speech of the word; determining abnormal information of the candidate address according to the feature vector of the candidate address; and identifying the abnormal address in the address data according to the abnormal information of the candidate address, thereby solving the technical problems of poor abnormal address identification effect and incapability of intelligent fraud prevention in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for identifying abnormal addresses. Background Art

[0002] A large-scale Internet user group has not only created the Internet ecosystem but also given rise to black industrial chains surrounding mainstream Internet products and services, gathering a large number of black production groups. To promptly identify fraud behaviors in online black production and avoid financial losses, risk control and anti-fraud technology is an essential part of e-commerce and financial companies. With the rapid development and popularization of the Internet, address information has become a relatively common means of collection, and it is easy for black production to use this information for fraud.

[0003] Currently, the commonly used fraud address identification methods mainly include statistical analysis-based methods, supervised learning based on intelligent algorithms, and unsupervised learning based on anomaly detection; only the classification patterns of abnormal addresses can be learned through machine learning model training or summary by professionals, and the identification effect is poor, unable to truly achieve the purpose of intelligent anti-fraud. Summary of the Invention

[0004] This application aims to at least partly solve one of the technical problems in the related art.

[0005] To this end, the first objective of this application is to propose a method for identifying abnormal addresses to accurately identify abnormal addresses and achieve the purpose of intelligent anti-fraud.

[0006] The second objective of this application is to propose a device for identifying abnormal addresses.

[0007] The third objective of this application is to propose an electronic device.

[0008] The fourth objective of this application is to propose a computer-readable storage medium.

[0009] The fifth objective of this application is to propose a computer program product.

[0010] To achieve the above objectives, the first aspect embodiment of this application proposes a method for identifying abnormal addresses, including:

[0011] Obtain the words of each candidate address in the address data and the part of speech of the words;

[0012] Generate a feature vector of the candidate address based on the words of the candidate address and the part of speech of the words;

[0013] Determine the abnormal information of the candidate address according to the feature vector of the candidate address;

[0014] Identify the abnormal addresses in the address data according to the abnormal information of the candidate addresses.

[0015] To achieve the above object, an embodiment of the second aspect of the present application provides a device for identifying abnormal addresses, including:

[0016] A first acquisition module, configured to acquire the words of each candidate address in the address data and the part of speech of the words;

[0017] A second acquisition module, configured to generate a feature vector of the candidate address based on the words of the candidate address and the part of speech of the words;

[0018] A third acquisition module, configured to determine the abnormal information of the candidate address according to the feature vector of the candidate address;

[0019] An identification module, configured to identify the abnormal addresses in the address data according to the abnormal information of the candidate addresses.

[0020] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0021] The memory stores computer-executable instructions;

[0022] The processor executes the computer-executable instructions stored in the memory to implement the method described in the embodiment of the first aspect.

[0023] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the first aspect.

[0024] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the embodiment of the first aspect.

[0025] The method, device, electronic device and storage medium for identifying abnormal addresses provided by the present application obtain the words of the candidate addresses and the part of speech of the words, determine the feature vector of the candidate address, thereby determine the abnormal information according to the feature vector, and identify the abnormal addresses based on the abnormal information, solving the problem that the traditional identification method cannot truly understand the address text, improving the identification effect, and achieving the purpose of intelligent anti-fraud.

[0026] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0028] Figure 1 A flowchart of a method for identifying an abnormal address provided in an embodiment of the present application;

[0029] Figure 2 A schematic diagram of a process for obtaining a feature vector provided in an embodiment of the present application;

[0030] Figure 3 A schematic diagram of a process for obtaining abnormal information provided by an embodiment of the present application;

[0031] Figure 4 A schematic diagram of a process for identifying abnormal addresses provided in an embodiment of the present application;

[0032] Figure 5 A schematic diagram of a flow chart of a model distillation process for identifying abnormal addresses provided in an embodiment of the present application;

[0033] Figure 6 A flowchart of another abnormal address identification method provided in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of the structure of an abnormal address identification device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0036] The large-scale Internet user group has generated huge Internet business needs and interactive traffic. In addition to the Internet ecosystem built by major Internet manufacturers through websites, applications, applets, and online and offline services, there are also a series of black industry chains derived from mainstream Internet products and services, which have gathered large-scale black industry groups. For example, "薅羊毛" is a typical black industry behavior, using false and unstable identity information to participate in marketing, discounts, and full-reduction activities to make profits, and cannot bring actual active users or order transactions to the platform. The Internet black industry is far more than just the wool party. Scalpers, brushing orders, cashing out, and junk registrations are all black industry behaviors, causing huge economic losses to Internet companies.

[0037] The first step in anti-fraud against black production behavior is to start from the detailed data of all interaction actions between users and the system. In terms of user data, compared with the commonly used user account information, contact phone numbers and other information in anti-fraud, the application of address information in the risk control field is relatively superficial, but the information contained in the address is very valuable; it is easy for black production to use this address information for fraud, for example, address fraud is carried out in the following ways:

[0038] False address: On some e-commerce or food delivery platforms, large-amount incentive red envelopes will be issued for new users or the first order of users, which will attract a large number of black production of wool pullers to carry out arbitrage. For example, fraudulent users will collude with merchants in advance, place orders in this store with multiple new accounts, and fill in a false address that does not exist at all for the delivery address. They may even directly note "no delivery required" or other words at the end of the order note or address as a secret signal to achieve the purpose of defrauding platform subsidies.

[0039] Vague address: The rise of consumer loans has triggered a wave of cash-out by black production, usually by purchasing easily realizable goods for resale and cash-out. For example, black production places an order for mobile devices, and the merchant does not ship directly, but takes advantage of the interest-free installment benefits during the platform promotion period to cash out; currently, e-commerce platforms will restrict orders with concentrated addresses. Therefore, in order to bypass the risk control rules, criminals may use vague addresses for transactions, such as only writing the address to a certain square or a certain community, rather than specifying the building number, so as to achieve address fraud.

[0040] Special address characters: In order to counter the risk control rules of the platform, traditional Chinese characters, misspelled words, pinyin or special characters are interspersed in the address to split keywords, etc., so as to complete black production behavior.

[0041] Currently, the identification schemes for fraudulent addresses mainly include methods based on statistical analysis, supervised learning based on intelligent algorithms, and unsupervised learning based on anomaly detection; among them, the method based on statistical analysis mainly relies on the experience accumulation of professionals, configures rules and intercepts based on known abnormal address features, but manual summary is time-consuming and laborious and has limited effects, and it is relatively easy to be bypassed; supervised learning based on intelligent algorithms can utilize the big data performance of all addresses, and with the help of deep learning models, automatically extract and summarize the key hidden features of fraudulent addresses. It requires a certain amount of representative address samples, and the identification quality depends on the annotation quality, which requires a large amount of manpower and time costs; unsupervised learning based on anomaly detection, although it can cover more types of abnormal addresses, has a lower identification accuracy; therefore, there is an urgent need for a method that can accurately identify abnormal addresses, based on the specific text understanding of address data, to achieve true intelligent anti-fraud.

[0042] Next, the method, device, electronic device and storage medium for identifying abnormal addresses according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0043] Figure 1 This is a schematic flowchart of a method for identifying abnormal addresses provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0044] S101, obtain the words and the part-of-speech of the words in each candidate address in the address data.

[0045] In some implementations, the address data may be addresses in the payment service of an e-commerce platform or a financial platform; optionally, the addresses can be analyzed on a daily basis, that is to say, the candidate addresses generated by the e-commerce platform or the financial platform within one day may be included in the address data.

[0046] In some implementations, the address data can be preliminarily screened, and candidate addresses such as null values, irregular lengths, or white lists are removed from the address data to reduce the computational amount of the candidate addresses in the address data.

[0047] It can be understood that a candidate address refers to the address data, which can be composed of texts such as provinces, cities, districts, counties, townships, or communities. Therefore, the candidate address will include multiple words that conform to semantic logic, and the words in the candidate address are obtained and analyzed.

[0048] Optionally, the candidate address can be segmented, the words in the candidate address are obtained, and named entity recognition is performed on the words in the candidate address to obtain the part-of-speech of the words.

[0049] Optionally, existing means such as named entity recognition technology and jieba segmentation can be used to implement the segmentation of the candidate address. In this embodiment, named entity recognition technology can be used to analyze the candidate address. Compared with other segmentation means, it has a better segmentation effect on complete named entities.

[0050] Exemplarily, for the candidate address "5th Floor, D Building, C Street, B District, A City", after segmenting this candidate address, five words that conform to semantic logic, namely "A City", "B District", "C Street", "D Building", and "5th Floor", will be obtained.

[0051] Further, after determining the words in the candidate address, the words in the candidate address can be part-of-speech tagged to facilitate the analysis of richer text information in the candidate address.

[0052] Optionally, the named entity recognition method can be used to judge the part-of-speech of the words. That is to say, when the named entity recognition method segments the words in the candidate address, it can additionally judge the part-of-speech of a single word.

[0053] Optionally, the part of speech can be a noun, verb, adverb, adjective, locative word, numeral classifier, auxiliary word, conjunction, punctuation mark, special symbol, etc. Optionally, the part of speech can be further divided in more detail. For example, a noun can be specifically labeled as a personal name, place name, animal name, organization name, food name, etc., and punctuation marks and special symbols can be specifically labeled as exclamation marks, commas, parentheses, caesuras, and full stops, etc.; so as to accurately identify subsequent abnormal addresses based on richer word and part-of-speech information in the candidate address.

[0054] S102, generate a feature vector of the candidate address based on the words and parts of speech of the candidate address.

[0055] Optionally, a word vector can be generated from the words of the candidate address respectively, and a part-of-speech vector can be generated from the parts of speech of the words. The feature vector of the candidate address is generated from the word vector and the part-of-speech vector.

[0056] In some implementations, the feature vector can include a word vector and a part-of-speech vector, that is, both the word vector and the part-of-speech vector are feature vectors of the candidate address; in some implementations, the word vector and the part-of-speech vector can also be concatenated to obtain the feature vector of the candidate address.

[0057] Furthermore, in order to facilitate the analysis of the feature vector, the feature vector can be digitally transformed. For example, different words and parts of speech are numerically labeled to obtain the numbers corresponding to each word and part of speech, and the corresponding words or parts of speech are replaced with the numbers, so as to obtain a word vector and a part-of-speech vector with elements being data, that is, a feature vector with elements being data.

[0058] S103, determine the abnormal information of the candidate address according to the feature vector of the candidate address.

[0059] In some implementations, the feature vector of the candidate address includes the text information of the candidate address. Therefore, a preliminary judgment can be made on the candidate address based on the feature vector of the candidate address to determine whether there is abnormal information for this candidate address.

[0060] Optionally, feature extraction can be performed on the feature vector, and abnormal analysis can be performed according to the extracted features. For example, the feature vector is input into a pre-trained abnormal recognition model, and the abnormal information indicating whether there is an abnormality in the feature vector is output.

[0061] S104, identify the abnormal addresses in the address data according to the abnormal information of the candidate addresses.

[0062] In some implementations, the abnormal information of the candidate address can include a judgment on whether this candidate address may be abnormal. That is to say, according to the abnormal information of the candidate address, the abnormal addresses that exist can be determined from all the candidate addresses in the address data.

[0063] In some implementations, the exception information for a candidate address with an exception can also be used as a suspected exception address, and the suspected exception address can be re-identified to improve the accuracy of exception address identification. For example, the suspected exception address can be input into an exception identification model for a second judgment to determine the exception address in the address data.

[0064] In this embodiment, by obtaining the words of the candidate address and the part-of-speech of the words, the feature vector of the candidate address is determined. The feature vector contains rich text information of the candidate address. It is relatively accurate to determine the exception information of the candidate address according to the feature vector, and then identify the exception address based on the accurate exception information. Identification is performed based on the rich text information of the candidate address itself, improving the accuracy of exception address identification and achieving the purpose of intelligent anti-fraud.

[0065] Based on the above embodiments, the feature vector includes a first vector and a second vector, and the process of obtaining the feature vector of the candidate address is described. As Figure 2 shown, the method includes the following steps:

[0066] S201, obtain the words of each candidate address in the address data and the part-of-speech of the words.

[0067] In the embodiments of the present application, the implementation method of step S201 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0068] S202, perform vector encoding on the words of the candidate address to generate the first vector of the candidate address.

[0069] Optionally, the words in the candidate address can be numbered, and the same words have the same number. Then, the numbers of all the words that have appeared in the address data can be determined, so that all the words in the candidate address can be converted into corresponding numbers, obtaining the number sequence corresponding to the words. The number sequence corresponding to the words is used as the first vector of the candidate address, and each element in the first vector is the number of the corresponding word in the candidate address.

[0070] S203, perform vector encoding on the part-of-speech of the words to generate the second vector of the candidate address.

[0071] Optionally, the part-of-speech of the words can be numbered, and the same part-of-speech has the same number. Then, the numbers of all the part-of-speech that have appeared in the address data can be determined, so that all the part-of-speech of the words in the candidate address can be converted into corresponding numbers, obtaining the number sequence corresponding to the part-of-speech. The number sequence corresponding to the part-of-speech is used as the second vector of the candidate address, and each element in the second vector is the number of the corresponding part-of-speech of the word in the candidate address.

[0072] After determining the first vector and the second vector of each candidate address, the feature vector corresponding to the candidate address is obtained, where the feature vector of the candidate address includes the first vector of the candidate address and the second vector of the candidate address.

[0073] In this embodiment, the first vector and the second vector are respectively generated based on the words of the candidate address and the part-of-speech of the words. The first vector and the second vector are used as the feature vector of the candidate address, which fully represents the text information of the candidate address. The abnormal information of the candidate address is identified and obtained based on the feature vector, so as to fully understand the address text and the accuracy is higher.

[0074] Based on the above embodiment, the process of obtaining the abnormal information of the candidate address is described. As Figure 3 shown, the method includes the following steps:

[0075] S301, obtain the probability score of the candidate address according to the feature vector of the candidate address.

[0076] Optionally, the first vector of the candidate address can be input into a pre-trained first anomaly recognition model to obtain the first probability score of the candidate address; the second vector of the candidate address can be input into a pre-trained second anomaly recognition model to obtain the second probability score of the candidate address.

[0077] In some implementations, the first anomaly recognition model can be an algorithm model for conventional unsupervised anomaly detection, such as the Hidden Markov Model (HMM). The HMM model is a type of Markov chain, which can be used to describe a Markov process with hidden unknown parameters, determine the hidden parameters of the process from the observable parameters, and then use the hidden parameters for further analysis, such as pattern recognition. The HMM model can be applied to fields such as speech recognition, behavior recognition, text recognition, and fault diagnosis. During the training process of the HMM model using address samples, the HMM model will give the probability score corresponding to each address sample according to the occurrence frequency of each word and the transition probability of the order between words. The address sample includes the first vectors of multiple sample data.

[0078] In some implementations, the second anomaly recognition model can also be an HMM model. The HMM model is trained using the part-of-speech corresponding to each word in the address sample. During the training process of the HMM model, it will give the probability score corresponding to the candidate address in terms of part-of-speech according to the occurrence frequency of each part-of-speech and the transition probability of the order between parts-of-speech. The address sample includes the second vectors of multiple samples.

[0079] Further, input the first vector of the candidate address into the trained HMM model related to words, and output the first probability score corresponding to the candidate address; input the second vector of the candidate address into the trained HMM model related to part of speech, and output the second probability score corresponding to the candidate address.

[0080] It can be understood that both the first probability score and the second probability score are used to represent the normality degree of the candidate address. The smaller the first probability score and the second probability score are, the lower the probability of the candidate address appearing in the overall address data, that is, the more abnormal the candidate address is; correspondingly, the larger the first probability score and the second probability score are, the more likely the candidate address is a normal address.

[0081] In some implementations, the first probability score and the second probability score can be aggregated to obtain a more accurate probability score corresponding to the candidate address. Optionally, the average value of the first probability score and the second probability score can be calculated, and the average value is used as the probability score of the candidate address; or, the larger value of the first probability score and the second probability score can be used as the probability score of the candidate address; alternatively, the first probability score and the second probability score can be weighted and summed to obtain the probability score of the candidate address.

[0082] S302. Determine the abnormal information of the candidate address according to the probability score of the candidate address.

[0083] It can be understood that the probability score of the candidate address reflects the normal situation of the candidate address. The smaller the probability score is, the more likely the candidate address is an abnormal address.

[0084] Optionally, the candidate addresses can be sorted in ascending order according to the probability score to obtain the first sequence; starting from the first candidate address in the first sequence, intercept the first sequence according to a preset ratio to obtain the first address set.

[0085] In some implementations, the abnormal information of the candidate address can include the category of whether the candidate address is a normal address or a possibly abnormal address. It can be understood that the probability scores corresponding to the candidate addresses in the first address set are smaller, that is, the candidate addresses in the first address set may be abnormal addresses. Therefore, the category of each candidate address in the first address combination is a suspected abnormal address.

[0086] Correspondingly, the candidate addresses in the first sequence except the first address set are used as the second address set, and the category of each candidate address in the second address set is a normal address.

[0087] After determining the categories corresponding to all candidate addresses in the address data, determine the abnormal information of each candidate address in the address data, and the abnormal information can be used to indicate whether the candidate address is a suspected abnormal address or a normal address.

[0088] In this embodiment, the first vector and the second vector in the feature vector are processed by the first anomaly recognition model and the second anomaly recognition model to obtain the corresponding first probability score and second probability score, thereby determining the probability score reflecting the normality degree of the candidate address. The efficiency and accuracy of processing the candidate address by the two anomaly recognition models are relatively high. Based on the probability score of the candidate address, the anomaly information is obtained, and the anomaly information of the more reasonable candidate address is obtained.

[0089] Based on the above embodiment, the process of identifying abnormal addresses in the address data is described. As Figure 4 shown, the method includes the following steps:

[0090] S401. Based on the anomaly information, determine all suspected abnormal addresses in the address data.

[0091] Since the anomaly information can be used to indicate that the candidate address is a suspected abnormal address or a normal address; therefore, all suspected abnormal addresses in the address data can be determined based on the anomaly information corresponding to each candidate address, and the suspected abnormal addresses are identified again for anomaly detection to improve the accuracy of address anomaly detection.

[0092] S402. Input the suspected abnormal addresses into a pre-trained third anomaly recognition model to obtain the abnormal addresses.

[0093] In some implementations, the third anomaly recognition model can be a large language model. The large language model has a more powerful text understanding ability. Therefore, based on the text understanding ability of the large language model, the suspected abnormal addresses can be identified and judged again to ensure a more accurate recognition result.

[0094] In some implementations, an open-source large language model can be obtained. Since the large language model has not been covered and trained in address anomaly classification, the large language model needs to be fine-tuned so that the large language model can better apply to the scenario of address anomaly classification.

[0095] Optionally, random sampling and annotation can be performed from the historically accumulated abnormal address samples, and normal address samples similar in number to the abnormal address samples are obtained to form a sample set of the large language model, and the sample set is divided into a training set and a validation set according to a ratio.

[0096] The large language model is fine-tuned using the labeled training set. Each address sample is input into the large language model to obtain the corresponding output result. The difference between the output result and the actual result labeled by the address sample is used as the loss function, and the large language model is continuously optimized and adjusted according to the loss function until the value of the loss function no longer decreases, and the training of the large language model is ended to obtain the fine-tuned large language model, that is, the pre-trained third anomaly recognition model.

[0097] Further, input the suspected abnormal address into the pre-trained third abnormal recognition model to output a judgment result, which indicates whether the suspected abnormal address is a definite abnormal address; input all the suspected abnormal addresses into the large language model to output the abnormal addresses in the address data.

[0098] In this embodiment, the suspected abnormal addresses are determined according to the abnormal information, and the suspected abnormal addresses are re-identified to reduce the time-consuming of abnormal recognition and improve the recognition accuracy. Using the large language model as the third abnormal recognition model to analyze the suspected abnormal addresses solves the problem that traditional recognition methods cannot truly understand the address text. By understanding the address text, the real abnormal addresses are determined from the suspected abnormal addresses, achieving better abnormal recognition.

[0099] On the basis of the above embodiment, after the abnormal addresses are recognized, model distillation can be performed for real-time abnormal address recognition. As Figure 5 shown, the method includes the following steps:

[0100] S501, record the abnormal addresses in the sample set.

[0101] In some implementations, all the abnormal addresses output by the third abnormal recognition model are recorded in the sample set, and the abnormal addresses in the sample set are the abnormal addresses determined based on the judgment rules of the large language model.

[0102] It can be understood that the judgment of the abnormal addresses based on the third abnormal recognition model is carried out in an offline environment, that is, all the address data generated by the platform the previous day is obtained, and each candidate address in the address data is judged to obtain the abnormal addresses. If abnormal address recognition needs to be carried out in a real-time environment, a lightweight model can be obtained based on the abnormal addresses in the sample set and deployed in the online environment.

[0103] S502, in response to the number of abnormal addresses in the sample set being greater than or equal to the preset quantity, perform model distillation on the fourth abnormal recognition model based on the sample set to obtain the pre-trained fourth abnormal recognition model.

[0104] In some implementations, the preset quantity can be 20,000, 30,000 or larger data to ensure that there is enough data in the sample set.

[0105] In some implementations, after the number of abnormal addresses in the sample set is greater than or equal to the preset quantity, normal addresses similar to the number of abnormal addresses can be obtained and recorded in the sample set to facilitate model distillation of the fourth abnormal recognition model based on the sample set.

[0106] Optionally, the normal address can be the normal address output by the third anomaly recognition model, or can be the normal address sampled from the existing address database.

[0107] It can be understood that model distillation is a model compression technology, aiming to transfer the knowledge of a complex and large model (teacher model) to another smaller and simpler model (student model); the key idea is to use the output of the teacher model as the target and let the student model learn how to approximate the prediction results of the teacher model.

[0108] Optionally, the fourth anomaly recognition model can be a Long Short-Term Memory (LSTM) model, which has excellent performance in short text classification tasks and has a small number of parameters, and can achieve fast recognition and reasoning required by the online risk control environment. Based on the sample set composed of abnormal addresses, the LSTM model is distilled and trained to obtain a pre-trained fourth anomaly recognition model. The fourth anomaly recognition model can learn the recognition pattern of abnormal addresses under the guidance of the classification results of the third anomaly recognition model.

[0109] S503, obtain the real-time address to be processed.

[0110] In some implementations, the pre-trained fourth anomaly recognition model, that is, the LSTM model, can be deployed in the online lockdown environment. When a new address to be processed is generated in real time, the address to be processed is obtained and analyzed.

[0111] S504, input the address to be processed into the pre-trained fourth anomaly recognition model to generate indication information on whether the address to be processed is an abnormal address.

[0112] Optionally, the address to be processed can be input into the pre-trained fourth anomaly recognition model to output indication information on whether the address to be processed is an abnormal address.

[0113] In some implementations, if it is recognized that the address to be processed is an abnormal address, the payment request or purchase request and other orders corresponding to the current abnormal address can be marked or rejected to achieve the purpose of intelligent anti-fraud.

[0114] In this embodiment, after obtaining the abnormal address through the third anomaly recognition model, the abnormal addresses are accumulated and distilled for learning, so as to obtain the lightweight model, the fourth anomaly model, so that the fourth anomaly model can be deployed in the online risk control environment. After obtaining the address to be processed in real time, the address to be processed is recognized based on the fourth anomaly recognition model to determine whether it is an abnormal address, so as to determine whether to reject the order, support more application environments, and achieve accurate intelligent anti-fraud.

[0115] Figure 6It is a schematic flowchart of another method for identifying abnormal addresses provided by an embodiment of this application. As Figure 6 shown, the method includes the following steps:

[0116] S601, Obtain the words and the part-of-speech of the words in each candidate address in the address data.

[0117] In the embodiments of this application, the implementation method of step S601 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0118] S602, Perform vector encoding on the words of the candidate address to generate a first vector of the candidate address.

[0119] In the embodiments of this application, the implementation method of step S602 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0120] S603, Perform vector encoding on the part-of-speech of the words to generate a second vector of the candidate address.

[0121] In the embodiments of this application, the implementation method of step S603 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0122] S604, Obtain the probability score of the candidate address according to the feature vector of the candidate address.

[0123] In the embodiments of this application, the implementation method of step S604 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0124] S605, Determine the abnormal information of the candidate address according to the probability score of the candidate address.

[0125] In the embodiments of this application, the implementation method of step S605 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0126] S606, Based on the abnormal information, determine all suspected abnormal addresses in the address data.

[0127] In the embodiments of this application, the implementation method of step S606 can be implemented in any one of the embodiments of this disclosure, and no limitation is made here and will not be elaborated further.

[0128] S607, Input the suspected abnormal address into a pre-trained third abnormal recognition model to obtain the abnormal address.

[0129] In the embodiments of the present application, the implementation method of step S607 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0130] S608, record the abnormal address in the sample set.

[0131] In the embodiments of the present application, the implementation method of step S608 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0132] S609, in response to the number of abnormal addresses in the sample set being greater than or equal to a preset number, perform model distillation on the fourth abnormal recognition model based on the sample set to obtain a pre-trained fourth abnormal recognition model.

[0133] In the embodiments of the present application, the implementation method of step S609 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0134] S610, obtain the real-time address to be processed.

[0135] In the embodiments of the present application, the implementation method of step S610 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0136] S611, input the address to be processed into the pre-trained fourth abnormal recognition model to generate indication information on whether the address to be processed is an abnormal address.

[0137] In the embodiments of the present application, the implementation method of step S611 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0138] In this embodiment, by obtaining the words and the part-of-speech of the words of the candidate address, a first vector and a second vector are respectively generated, and the first vector and the second vector are used as the feature vectors of the candidate address. The feature vectors contain rich text information of the candidate address. The first vector and the second vector in the feature vectors are processed by the first anomaly recognition model and the second anomaly recognition model to obtain the corresponding first probability score and second probability score, so as to determine the probability score reflecting the normality degree of the candidate address. It is more reasonable to obtain the anomaly information based on the probability score of the candidate address. The suspected abnormal address is determined according to the anomaly information, and the suspected abnormal address is recognized again, reducing the time-consuming of anomaly recognition and improving the accuracy of recognition, solving the problem that the traditional recognition method cannot truly understand the address text, and realizing better anomaly recognition; after determining the abnormal address, model distillation can also be performed to obtain a fourth anomaly recognition model that can be deployed in the online environment, and the address text is recognized and judged in real time, achieving the purpose of intelligent anti-fraud.

[0139] To implement the above embodiment, the present application also proposes an identification device for abnormal addresses.

[0140] Figure 7 The structural schematic diagram of an identification device for abnormal addresses provided by an embodiment of the present application is as Figure 7 shown. The identification device for abnormal addresses includes:

[0141] A first acquisition module 701, configured to acquire the words and the part-of-speech of the words of each candidate address in the address data;

[0142] A second acquisition module 702, configured to generate a feature vector of the candidate address based on the words and the part-of-speech of the candidate address;

[0143] A third acquisition module 703, configured to determine the anomaly information of the candidate address according to the feature vector of the candidate address;

[0144] An identification module 704, configured to identify the abnormal address in the address data according to the anomaly information of the candidate address.

[0145] Further, in a possible implementation manner of the embodiment of the present application, the third acquisition module 703 includes:

[0146] Obtain the probability score of the candidate address according to the feature vector of the candidate address, where the probability score represents the normality degree of the candidate address;

[0147] Determine the anomaly information of the candidate address according to the probability score of the candidate address.

[0148] Further, in a possible implementation manner of the embodiment of the present application, the first acquisition module 701 includes:

[0149] Segment the candidate address to obtain the words in the candidate address, and perform named entity recognition on the words in the candidate address to obtain the part-of-speech of the words.

[0150] Further, in a possible implementation manner of the embodiment of the present application, the second acquisition module 702 includes:

[0151] Perform vector encoding on the words of the candidate address to generate a first vector of the candidate address;

[0152] Perform vector encoding on the part-of-speech of the words to generate a second vector of the candidate address;

[0153] Wherein the feature vector of the candidate address includes the first vector of the candidate address and the second vector of the candidate address.

[0154] Further, in a possible implementation manner of the embodiment of the present application, the third acquisition module 703 includes:

[0155] Input the first vector of the candidate address into a pre-trained first anomaly recognition model to obtain a first probability score of the candidate address;

[0156] Input the second vector of the candidate address into a pre-trained second anomaly recognition model to obtain a second probability score of the candidate address;

[0157] Determine the probability score of the candidate address according to the first probability score and the second probability score.

[0158] Further, in a possible implementation manner of the embodiment of the present application, the anomaly information includes the category of the candidate address, and the third acquisition module 703 includes:

[0159] Arrange the candidate addresses in ascending order according to the probability scores to obtain a first sequence;

[0160] Starting from the first candidate address in the first sequence, intercept the first sequence according to a preset ratio to obtain a first address set;

[0161] Wherein, the candidate addresses in the first sequence other than the first address set are the second address set, the category of each candidate address in the first address set is a suspected abnormal address, and the category of each candidate address in the second address set is a normal address.

[0162] Further, in a possible implementation manner of the embodiment of the present application, the recognition module 704 includes:

[0163] Based on the anomaly information, determine all the suspected abnormal addresses in the address data;

[0164] Input the suspected abnormal addresses into a pre-trained third anomaly recognition model to obtain abnormal addresses.

[0165] Further, in a possible implementation manner of the embodiment of the present application, after the recognition module 704, it includes:

[0166] Record the abnormal address in the sample set;

[0167] In response to the number of abnormal addresses in the sample set being greater than or equal to a preset number, perform model distillation on the fourth abnormal recognition model based on the sample set to obtain a pre-trained fourth abnormal recognition model.

[0168] Further, in a possible implementation manner of the embodiment of the present application, after the recognition module 704, it further includes:

[0169] Obtain the real-time address to be processed;

[0170] Input the address to be processed into the pre-trained fourth abnormal recognition model to generate indication information on whether the address to be processed is an abnormal address.

[0171] It should be noted that the foregoing explanation of the embodiment of the method for identifying abnormal addresses is also applicable to the device for identifying abnormal addresses in this embodiment, and will not be elaborated here.

[0172] In this embodiment, by obtaining the words of the candidate address and the part-of-speech of the words, the feature vector of the candidate address is determined. The feature vector contains rich text information of the candidate address. It is relatively accurate to determine the abnormal information of the candidate address according to the feature vector, and then identify the abnormal address according to the accurate abnormal information. Based on the rich text information of the candidate address itself for identification, the accuracy of identifying the abnormal address is improved, achieving the purpose of intelligent anti-fraud.

[0173] To implement the above embodiment, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiment.

[0174] To implement the above embodiment, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiment.

[0175] To implement the above embodiment, the present application also proposes a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method provided in the foregoing embodiment.

[0176] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0177] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to safeguard and secure access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0178] This application is expected to provide an implementation plan for users to selectively block the use or access of personal information data. That is, this disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0179] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0180] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0181] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0182] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained, for example, electronically by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0183] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0184] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0185] In addition, in each of the embodiments of the present application, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0186] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for identifying an abnormal address, characterized in that, The method includes: Obtaining the words of each candidate address in the address data and the part-of-speech of the words; Generating a feature vector of the candidate address based on the words of the candidate address and the part-of-speech of the words; Determining the abnormal information of the candidate address according to the feature vector of the candidate address; Identifying the abnormal addresses in the address data according to the abnormal information of the candidate address.

2. The method according to claim 1, wherein The determining the abnormal information of the candidate address according to the feature vector of the candidate address includes: Obtaining a probability score of the candidate address according to the feature vector of the candidate address, where the probability score represents the normal degree of the candidate address; Determining the abnormal information of the candidate address according to the probability score of the candidate address.

3. The method according to claim 1 or 2, characterized in that, The obtaining the words of each candidate address in the address data and the part-of-speech of the words includes: Performing word segmentation on the candidate address to obtain the words in the candidate address, and performing named entity recognition on the words in the candidate address to obtain the part-of-speech of the words.

4. The method according to claim 3, characterized in that, The generating a feature vector of the candidate address based on the words of the candidate address and the part-of-speech of the words includes: Performing vector encoding on the words of the candidate address to generate a first vector of the candidate address; Performing vector encoding on the part-of-speech of the words to generate a second vector of the candidate address; Wherein the feature vector of the candidate address includes the first vector of the candidate address and the second vector of the candidate address.

5. The method according to claim 4, wherein The obtaining a probability score of the candidate address according to the feature vector of the candidate address includes: Inputting the first vector of the candidate address into a pre-trained first abnormal recognition model to obtain a first probability score of the candidate address; Inputting the second vector of the candidate address into a pre-trained second abnormal recognition model to obtain a second probability score of the candidate address; Determining the probability score of the candidate address according to the first probability score and the second probability score.

6. The method according to claim 5, wherein The abnormal information includes the category of the candidate address. The determining the abnormal information of the candidate address according to the probability score of the candidate address includes: Performing ascending sorting on the candidate addresses according to the probability score to obtain a first sequence; Starting from the first candidate address in the first sequence, intercepting the first sequence according to a preset ratio to obtain a first address set; Wherein, the candidate addresses in the first sequence except the first address set are a second address set, the category of each candidate address in the first address set is a suspected abnormal address, and the category of each candidate address in the second address set is a normal address.

7. The method according to any one of claims 1-6, characterized in that, The identifying the abnormal addresses in the address data according to the abnormal information of the candidate address includes: Determining all the suspected abnormal addresses in the address data based on the abnormal information; Inputting the suspected abnormal addresses into a pre-trained third abnormal recognition model to obtain the abnormal addresses.

8. The method according to claim 7, characterized in that, After the identifying the abnormal addresses in the address data, it includes: Recording the abnormal addresses in a sample set; In response to the number of abnormal addresses in the sample set being greater than or equal to a preset number, model distillation is performed on the fourth abnormal recognition model based on the sample set to obtain a pre-trained fourth abnormal recognition model.

9. The method according to claim 8, wherein After obtaining the pre-trained fourth abnormal recognition model, it further includes: Obtaining real-time addresses to be processed; Inputting the addresses to be processed into the pre-trained fourth abnormal recognition model to generate indication information on whether the addresses to be processed are abnormal addresses.

10. An abnormal address recognition device, characterized in that, It includes: A first acquisition module for acquiring the words of each candidate address in the address data and the part of speech of the words; A second acquisition module for generating a feature vector of the candidate address based on the words of the candidate address and the part of speech of the words; A third acquisition module for determining the abnormal information of the candidate address according to the feature vector of the candidate address; An identification module for identifying the abnormal addresses in the address data according to the abnormal information of the candidate address.

11. An electronic device, characterized in that, It includes: A processor and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-9.

13. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-9.