Case series and parallel method, device and equipment based on transmission relationship reasoning and storage medium
Through a method based on transitive relationship reasoning, the clue characteristics of fraud cases are automatically analyzed, and related case sets are identified and merged, which solves the problem of low efficiency of traditional manual analysis and realizes the intelligent merging and rapid detection of fraud cases.
Patent Information
- Application Number
- CN202411735012.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Traditional manual analysis of fraud cases is inefficient and difficult to quickly handle complex and diverse online fraud cases. Existing technologies make it difficult to effectively merge and identify the correlations between cases.
A method based on transitive relationship reasoning is used to extract communication, network behavior, financial transactions and application clue features from multiple fraud cases. Through intersection analysis and transitive relationship configuration, related case sets are automatically identified and merged to achieve intelligent case processing.
It has improved the efficiency of solving fraud-related cases, enhanced the accuracy of investigations, optimized resource allocation, promoted cross-regional collaboration, shortened the investigation cycle and improved the accuracy of case mergers.
Smart Images

Figure CN119398985B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a case stringing and parallelizing method and device based on transitive relationship reasoning, a computer device and a storage medium. BACKGROUND
[0002] With the rapid development of information technology, network criminal behaviors have the characteristics of increasing complexity, concealment and cross-regional nature, such as network fraud, single brushing and pig killing, and other fraud-related cases are emerging in an endless stream, which poses a serious challenge to public safety. Traditional analysis of fraud-related cases mainly relies on manual review, information comparison and assistance from some data analysis tools. However, in the current complex network environment, there are various fraud-related means, and the amount of fraud-related case data is large. Artificial analysis of fraud-related cases is inefficient in the face of big data, which is not conducive to the rapid investigation of fraud-related cases. SUMMARY
[0003] Therefore, it is necessary to provide a case stringing and parallelizing method and device based on transitive relationship reasoning, a computer device and a storage medium to solve the above technical problems, which can analyze and effectively merge new fraud-related cases, realize intelligent merging of fraud-related cases, and improve the investigation efficiency of fraud-related cases.
[0004] A case stringing and parallelizing method based on transitive relationship reasoning, comprising: obtaining a plurality of fraud-related cases from a plurality of network-related cases; extracting a plurality of clue features of each fraud-related case, the plurality of clue features including any one of a communication clue, a fraud-related network behavior clue, a fund transaction clue and a fraud-related application program clue; performing intersection analysis on the plurality of clue features of each fraud-related case to determine a first fraud-related case set and a second fraud-related case set having clue feature correlations, wherein the fraud-related cases in the first fraud-related case set have an intersection of first clue features, the fraud-related cases in the second fraud-related case set have an intersection of second clue features, and the first clue features and the second clue features are different; if the first fraud-related case set and the second fraud-related case set have one or more same fraud-related cases, it is determined that the first fraud-related case set and the second fraud-related case set have a relationship of belonging to the same fraud-related nature case; if it is identified that the first fraud-related case set and the second fraud-related case set belong to the same fraud-related nature case, a transitive relationship of the first fraud-related case set and the second fraud-related case set is configured; and the fraud-related cases in the first fraud-related case set and the fraud-related cases in the second fraud-related case set are processed according to the transitive relationship of the first fraud-related case set and the second fraud-related case set.
[0005] In one of the embodiments, the communication clues include any one or more of a mobile phone number and a communication contact in the mobile phone number, an instant messaging number and a communication contact in the instant messaging number, an email number and an email contact in the email number; and / or, the fund transaction clues include a fund transaction account number and a transfer account number; and / or, the cybercrime network behavior clues include one or more of a cybercrime website domain name, a cybercrime website IP address, and a third-party service ID.
[0006] In one of the embodiments, the cybercrime application clues include information of the cybercrime application and associated information of the cybercrime application, the information of the cybercrime application includes a package name and packaging information of the cybercrime application, MD5 values of cybercrime application files, a signature of the cybercrime application, and MD5 values of cybercrime application icons, the associated information of the cybercrime application includes a distribution website of the cybercrime application, a distribution download address of the cybercrime application, an application server IP of the cybercrime application, an application server domain name of the cybercrime application, a packaging service system of the cybercrime application, a customer service system of the cybercrime application, and a push platform of the cybercrime application.
[0007] In one of the embodiments, the intersection analysis is performed on the multiple clues features of each cybercrime case to determine a first cybercrime case set in which the clues features are associated, including: performing intersection analysis on the information of the cybercrime application and the associated information of the cybercrime application of each cybercrime case; if the package name and packaging information of the cybercrime application of any two cybercrime cases are associated, determining whether any one or more of the MD5 values of the cybercrime application files, the signature of the cybercrime application, and the MD5 values of the cybercrime application icons are associated, if any one or more of the MD5 values of the cybercrime application files, the signature of the cybercrime application, and the MD5 values of the cybercrime application icons are associated, determining that the any two cybercrime cases are cybercrime cases in the first cybercrime case set, wherein the first clue feature is the cybercrime application clue; if the package name and packaging information of the cybercrime application of any two cybercrime cases are not associated, determining whether any one or more of the distribution website of the cybercrime application, the distribution download address of the cybercrime application, the application server IP of the cybercrime application, the application server domain name of the cybercrime application, the packaging service system of the cybercrime application, the customer service system of the cybercrime application, and the push platform of the cybercrime application are associated, if any one or more of the distribution website of the cybercrime application, the distribution download address of the cybercrime application, the application server IP of the cybercrime application, the application server domain name of the cybercrime application, the packaging service system of the cybercrime application, the customer service system of the cybercrime application, and the push platform of the cybercrime application are associated, determining that the any two cybercrime cases are cybercrime cases in the first cybercrime case set, wherein the first clue feature is the cybercrime application clue.
[0008] In one of the embodiments, the multiple clue features of each fraud case are subjected to intersection analysis to determine a second fraud case set in which the clue features are associated, including: subjecting the multiple clue features of each fraud case to intersection analysis to determine multiple fraud cases in which the second clue features are associated, and constructing the second fraud case set according to the multiple fraud cases in which the second clue features are associated, the second clue features being any one or more of the communication clue, the fraud network behavior clue, and the fund transaction clue.
[0009] In one of the embodiments, after the step of processing the fraud cases in the first fraud case set and the fraud cases in the second fraud case set in series and in parallel according to the transmission relationship between the first fraud case set and the second fraud case set, the method further includes: obtaining the collection unit information of each target fraud case obtained through the series and parallel processing, and dividing the multiple target fraud cases according to the attribution places according to the collection unit information of each target fraud case to obtain the target fraud case set of each attribution place.
[0010] In one of the embodiments, the first clue feature is one, and the second clue feature is one; or, the first clue feature is multiple, the second clue feature is one, and the second clue feature is different from any first clue feature; or, the first clue feature is multiple, the second clue feature is multiple, and any second clue feature is different from any first clue feature; and the method further includes: configuring the retrieval keywords of the first fraud case set according to the first clue feature, and configuring the retrieval keywords of the second fraud case set according to the second clue feature; wherein, when the first clue feature is one, the retrieval keywords of the first fraud case set is one, when the first clue feature is multiple, the retrieval keywords of the first fraud case set is multiple and the first fraud case set is retrieved by inputting multiple retrieval keywords; when the second clue feature is one, the retrieval keywords of the second fraud case set is one, and when the second clue feature is multiple, the retrieval keywords of the second fraud case set is multiple and the second fraud case set is retrieved by inputting multiple retrieval keywords.
[0011] A case string and parallel device based on transmission relationship reasoning, comprising: an acquisition module configured to acquire multiple fraud cases from multiple network-related cases; an extraction module configured to extract multiple clue features of each fraud case, the multiple clue features including any of communication clues, fraud network behavior clues, fund transaction clues, and fraud application program clues; an intersection analysis module configured to perform intersection analysis on the multiple clue features of each fraud case to determine a first fraud case set and a second fraud case set in which clue features are associated, wherein the first fraud case set has an intersection of first clue features between fraud cases, the second fraud case set has an intersection of second clue features between fraud cases, and the first clue features and the second clue features are different; a determination module configured to determine that the first fraud case set and the second fraud case set have a relationship of belonging to the same fraud nature case if the first fraud case set and the second fraud case set have one or more same fraud cases; a configuration module configured to configure transmission relationships of the first fraud case set and the second fraud case set if it is identified that the first fraud case set and the second fraud case set belong to the same fraud nature case; and a string and parallel processing module configured to perform string and parallel case processing on fraud cases in the first fraud case set and fraud cases in the second fraud case set according to the transmission relationships of the first fraud case set and the second fraud case set.
[0012] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method of any of the above embodiments when executing the computer program.
[0013] A computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the method of any of the above embodiments.
[0014] The case stringing and parallel processing method, device, computer equipment and storage medium based on the transmission relationship reasoning can automatically analyze the multiple clue features of the fraud-related cases, determine the first fraud-related case set and the second fraud-related case set with associated clue features, and preliminarily screen out the fraud-related cases with the same clue features. Then, it is determined whether the first fraud-related case set and the second fraud-related case set contain the same fraud-related cases. If yes, it is inferred that the first fraud-related case set and the second fraud-related case set have the relationship of belonging to the same fraud-related nature case. Then, the transmission relationship of the first fraud-related case set and the second fraud-related case set is configured. Finally, the fraud-related cases in the first fraud-related case set and the second fraud-related case set are processed by stringing and parallel processing based on the transmission relationship of the first fraud-related case set and the second fraud-related case set. The whole process does not need manual review and analysis, can deeply analyze and effectively combine the new fraud-related cases, realizes the intelligent combination of the fraud-related cases, and improves the investigation efficiency of the fraud-related cases.
[0015] Therefore, the multiple clue features of the fraud-related cases can be automatically analyzed, the first fraud-related case set and the second fraud-related case set with associated clue features can be determined, and the fraud-related cases with the same clue features can be preliminarily screened out. Then, it is determined whether the first fraud-related case set and the second fraud-related case set contain the same fraud-related cases. If yes, it is inferred that the first fraud-related case set and the second fraud-related case set have the relationship of belonging to the same fraud-related nature case. Then, the transmission relationship of the first fraud-related case set and the second fraud-related case set is configured. Finally, the fraud-related cases in the first fraud-related case set and the second fraud-related case set are processed by stringing and parallel processing based on the transmission relationship of the first fraud-related case set and the second fraud-related case set. The whole process does not need manual review and analysis, can deeply analyze and effectively combine the new fraud-related cases, realizes the intelligent combination of the fraud-related cases, and improves the investigation efficiency of the fraud-related cases. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The application environment diagram of the case stringing and parallel processing method based on the transmission relationship reasoning in one embodiment;
[0017] Figure 2 The flowchart of the case stringing and parallel processing method based on the transmission relationship reasoning in one embodiment;
[0018] Figure 3 The system configuration module diagram of the case stringing and parallel processing method based on the transmission relationship reasoning in one embodiment;
[0019] Figure 4 The processing diagram of the case stringing and parallel processing method based on the transmission relationship reasoning in one example;
[0020] Figure 5 A structural block diagram of a case string and parallel device based on a transmission relationship reasoning in an embodiment;
[0021] Figure 6 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0022] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0023] The present application provides a case string and parallel method based on a transmission relationship reasoning, which is applied to an application environment as shown in Figure 1 As shown in Figure 1 , the server 102 obtains multiple fraud cases from the databases of various platforms, extracts multiple clue features of each fraud case, the multiple clue features including any multiple of communication clues, fraud network behavior clues, fund transaction clues and fraud application program clues, performs intersection analysis on the multiple clue features of each fraud case, determines multiple target fraud cases with clue feature association, and performs string and parallel case processing on the multiple target fraud cases. The server 102 can be implemented by an independent server or a server cluster composed of multiple servers.
[0024] Among the various embodiments of the present application, APP, application and application program are the same meaning, all expressing application program.
[0025] In an embodiment, as shown in Figure 2 , a case string and parallel method based on a transmission relationship reasoning is provided, which is taken as an example to illustrate the server 102 in Figure 1 , including the following steps:
[0026] S202, obtaining multiple fraud cases from multiple network-related cases.
[0027] In the present embodiment, multiple network-related cases are collected, and each network-related case can come from a case collection platform in different regions. That is, the multiple fraud cases can come from the fraud case management platform of the collection unit in different administrative regions, or from the same fraud case management platform. Specifically, multiple fraud cases can be searched and obtained based on fraud case investigation related information, to provide intelligent assistance for fraud case investigation and analysis.
[0028] S204, extract multiple clue features of each fraud case, the multiple clue features including any one or more of a communication clue, a fraud network behavior clue, a fund transaction clue, and a fraud application program clue.
[0029] In this embodiment, the multiple clue features of each fraud case include any one or more of the communication clue, the fraud network behavior clue, the fund transaction clue, and the fraud application program clue. The communication clue is used to identify the communication relationship between the fraud personnel in the fraud case, the fund transaction clue is used to identify the fund transaction relationship between the fraud personnel in the fraud case, the fraud network behavior clue is used to identify the network fraud operation process adopted by the fraud personnel in the fraud case, and the fraud application program clue is used to identify the network fraud tool adopted by the fraud personnel in the fraud case. These clue features can reflect the correlation relationship of multiple fraud cases, and can assist the rapid case investigation of the case investigators when investigating a large number of fraud cases in a big data environment.
[0030] In one example, the communication clue includes any one or more of a mobile phone number and a communication contact in the mobile phone number, an instant messaging number and a communication contact in the instant messaging number, an email number and an email contact in the email number; and / or, the fund transaction clue includes a fund transaction account number and a transfer account number; and / or, the fraud network behavior clue includes one or more of a fraud website domain name, a fraud website IP address, and a third-party service ID.
[0031] Specifically, the communication clue: the fraud case is merged through the number identifier data such as the mobile phone number, bank card, ID card, email, WeChat, WeChat group, QQ, QQ group, etc. in the fraud case investigation data of the fraud case. And, the fraud case is merged through the information such as the WeChat group member, QQ group member, etc. in the fraud case investigation data of the fraud case. The fraud network behavior clue: the fraud case is merged through the fraud website and the associated information such as the website domain name, website IP, third-party service ID, etc. in the fraud case. And, the fraud case is merged through the networking data service across regions.
[0032] In one example, the fraud application program clue includes information of the fraud application program and associated information of the fraud application program, the information of the fraud application program including the package name and packaging information of the fraud application program, the MD5 value of the fraud application program file, the signature of the fraud application program, and the MD5 value of the fraud application program icon, the associated information of the fraud application program including the distribution website of the fraud application program, the distribution download address of the fraud application program, the application server IP of the fraud application program, the application server domain name of the fraud application program, the packaging service system of the fraud application program, the customer service system of the fraud application program, and the push platform of the fraud application program.
[0033] In this example, the fraud cases are merged by the information of the fraud application and the associated information of the fraud application. For example, the information of the fraud application includes APP package name, APP file MD5 value, APP signature information, and MD5 value of APP icon. The associated information of the fraud application includes APP distribution website, APP distribution download address, APP packaging information, APP application server IP, APP application server domain name, APP packaging service, APP customer service system, and APP push platform.
[0034] In S206, the intersection analysis is performed on the multiple clue features of each fraud case to determine a first fraud case set and a second fraud case set in which the clue features are associated. The first clue feature is intersected between the fraud cases in the first fraud case set, and the second clue feature is intersected between the fraud cases in the second fraud case set. The first clue feature and the second clue feature are different.
[0035] In this embodiment, the intersection analysis is performed on the clue features of all the fraud cases. If there is an intersection of the clue features, the fraud cases are attributed to the corresponding fraud case set. If there is an intersection of the fraud network behavior clues between the multiple fraud cases, the multiple fraud cases are included in one of the fraud case sets. If there is an intersection of the fund transaction clues between the multiple fraud cases, the multiple fraud cases are included in another of the fraud case sets. If there is an intersection of the communication clues between the multiple fraud cases, the multiple fraud cases are included in another of the fraud case sets. If there is an intersection of the fraud application clues between the multiple fraud cases, the multiple fraud cases are included in another of the fraud case sets. Therefore, the multiple fraud cases are divided into multiple fraud case sets according to the clue features. The first fraud case set and the second fraud case set can be any two of the fraud case sets.
[0036] For example, there is an intersection of the communication clues between the fraud case A and the fraud case B. The fraud case A and the fraud case B are included in the first fraud case set, and the first clue feature is the communication clue. There is an intersection of the fraud application clues between the fraud case B and the fraud case C. The fraud case B and the fraud case C are included in the second fraud case set, and the second clue feature is the fraud application clue.
[0037] In one example, the intersection analysis of the multiple clue features of each fraud case determines the first set of fraud cases and the second set of fraud cases associated with the clue features, including: performing intersection analysis on the multiple clue features of each fraud case to determine the target clue features that exist in the intersection of the fraud cases, identifying the number of fraud cases containing the target clue features, and if the number is greater than a set number threshold, determining the first set of fraud cases and the second set of fraud cases associated with the target clue features.
[0038] In this example, after the intersection of the clue features between the fraud cases, further judgment is needed on the intersection clue features. When the number of times the intersection clue features are collected is greater than a set number, the intersection clue features are considered as the associated features between the two fraud cases. The number of times the intersection clue features are collected is determined by the number of fraud cases containing the clue features. For example, as shown in the following table:
[0039]
[0040]
[0041]
[0042]
[0043] In one embodiment, the intersection analysis of the multiple clue features of each fraud case determines the first set of fraud cases associated with the clue features, including: performing intersection analysis on the information of the fraud application of each fraud case and the associated information of the fraud application; if the package name and the packaging information of the fraud application of any two fraud cases are associated, determining whether any one of the MD5 value of the fraud application file, the signature of the fraud application, and the MD5 value of the fraud application icon of the two fraud cases is associated, and if any one or more are associated, determining that the two fraud cases are both fraud cases in the first set of fraud cases, wherein the first clue feature is the fraud application clue; if the package name and the packaging information of the fraud application of any two fraud cases are not associated, determining whether any one of the distribution website of the fraud application, the distribution download address of the fraud application, the application server IP of the fraud application, the application server domain name of the fraud application, the packaging service system of the fraud application, the customer service system of the fraud application, and the push platform of the fraud application is associated, and if any one or more are associated, determining that the two fraud cases are both fraud cases in the first set of fraud cases, wherein the first clue feature is the fraud application clue.
[0044] In the embodiment, when the package names and packaging information of the fraud applications of the two fraud cases are associated, it is further determined whether any of the MD5 values of the corresponding fraud application files, the signatures of the fraud applications, and the MD5 values of the fraud application icons are associated. If yes, it is indicated that the same fraud application is used in the two fraud cases, and the fraud cases are treated as fraud cases in the same fraud case set. Therefore, through the multiple determination mode, the fraud cases using the same fraud application can be accurately found and processed.
[0045] In addition, when the front-end fraud application is developed by the fraud user, the information of the fraud application may be changed by some means to avoid investigation. Therefore, in the embodiment, if it is detected that the package names and packaging information of the fraud applications of the two fraud cases are not associated, it is further determined whether any one or more of the following background information of the fraud application is associated: the distribution website of the fraud application, the distribution download address of the fraud application, the application server IP of the fraud application, the application server domain name of the fraud application, the packaging service system of the fraud application, the customer service system of the fraud application, and the push platform of the fraud application. If yes, it is indicated that the same fraud application is used by the fraud user, and the fraud cases are associated. The corresponding fraud cases are treated as fraud cases in the same fraud case set and are processed. Therefore, the accuracy of screening fraud cases during the same case investigation can be improved.
[0046] In one embodiment, the intersection analysis of the plurality of clue characteristics of each fraud case is performed to determine the second fraud case set having the associated clue characteristics, including: performing the intersection analysis of the plurality of clue characteristics of each fraud case to determine a plurality of fraud cases having the second clue characteristics, and constructing the second fraud case set according to the plurality of fraud cases having the second clue characteristics, the second clue characteristics being any one or more of the communication clue, the fraud network behavior clue, and the fund transaction clue.
[0047] In the embodiment, the second clue characteristics of the fraud cases in the second fraud case set can be any one or more of the communication clue, the fraud network behavior clue, and the fund transaction clue. That is, when the first clue characteristic is the fraud application clue, the second clue characteristic is any one or more of the other non-fraud application clues.
[0048] In the embodiment, the second clue characteristics of the fraud cases in the second fraud case set can be any one or more of the communication clue, the fraud network behavior clue, and the fund transaction clue. That is, when the first clue characteristic is the fraud application clue, the second clue characteristic is any one or more of the other non-fraud application clues.
[0049] In this embodiment, the fraud cases in the first fraud case set all have the first clue feature, the fraud cases in the second fraud case set all have the second clue feature, and if there is a same fraud case in the two fraud case sets, it means that the fraud case has both the first clue feature and the second clue feature. Therefore, it can be determined that the fraud cases in the two fraud case sets belong to the same fraud nature. For example, fraud case A and fraud case B have the same communication clue, fraud case B and fraud case C have the same fund transaction clue, but fraud case A and fraud case C do not have the same clue feature. By having the same clue feature with the same fraud case, i.e. fraud case B, it can be considered that fraud case A and fraud case C have a relationship of belonging to the same fraud nature case.
[0050] As described above, the first fraud case set and the second fraud case set are any two fraud case sets in a plurality of fraud case sets. By having one or more same fraud cases, it can be determined that the first fraud case set and the second fraud case set have a relationship of belonging to the same fraud nature case. Further, if there is one or more same fraud cases in the third fraud case set and the second fraud case set or the first fraud case set, the third fraud case set and the first fraud case set and the second fraud case set all have a relationship of belonging to the same fraud nature case. By analogy, a plurality of fraud case sets having a relationship of belonging to the same fraud nature case can be screened out.
[0051] S210, if it is identified that the first fraud case set and the second fraud case set have a relationship of belonging to the same fraud nature case, the transmission relationship of the first fraud case set and the second fraud case set is configured.
[0052] After it is determined in the above step S208 that the first fraud case set and the second fraud case set have a relationship of belonging to the same fraud nature case, in this embodiment, the transmission relationship of the first fraud case set and the second fraud case set is configured. The transmission relationship indicates that there is a same fraud case in the two case sets, and the same fraud attribute can be transmitted between the fraud cases when the case string processing is performed. By analogy, if there are a plurality of fraud case sets such as the third fraud case set and the fourth fraud case set having a relationship of belonging to the same fraud nature case with the first fraud case set and the second fraud case set, the transmission relationship between these fraud case sets is also configured.
[0053] S212, according to the transmission relationship of the first fraud case set and the second fraud case set, the fraud cases in the first fraud case set and the fraud cases in the second fraud case set are processed in a string and parallel manner.
[0054] In this embodiment, the series-parallel processing refers to a processing mode combining series and parallel cases. The series refers to putting a series of different cases with a connection together for investigation by analyzing the criminal means, traces, and physical evidence. The parallel refers to putting two cases with a connection together for investigation by analyzing the criminal means, traces, and physical evidence. If there is a transmission relationship between any two sets of fraud cases, the fraud cases in the two sets of fraud cases are processed in series and parallel. Similarly, if there is a transmission relationship between multiple sets of fraud cases, the fraud cases in the multiple sets of fraud cases are processed in series and parallel.
[0055] Specifically, the multiple fraud cases can be processed in series and parallel by taking the case numbers of the fraud cases as the associated connection. Alternatively, the multiple fraud cases can be processed in series and parallel by taking the clues features with an intersection of the fraud cases as the associated connection.
[0056] In one embodiment, after the step of processing the fraud cases in the first set of fraud cases and the fraud cases in the second set of fraud cases in series and parallel according to the transmission relationship between the first set of fraud cases and the second set of fraud cases, the method further includes: obtaining collection unit information of each target fraud case obtained by the series-parallel processing; and dividing the multiple target fraud cases according to the collection unit information of each target fraud case and according to the attribution to obtain a set of target fraud cases in each attribution.
[0057] In this embodiment, the collection unit information is used to indicate a unit that collects and manages the target fraud cases. For example, the collection unit information indicates that the collection unit of the target fraud case is a public security bureau or a fraud case investigation agency. Further, the set of target fraud cases is divided according to the attribution. In one implementation, the attribution can be divided according to an administrative region, for example, the attribution can be divided into provinces, cities, districts, and counties. Therefore, the investigation and analysis of the fraud cases can be gradually refined in the province, the city, the district, and the county, and the distribution of the criminal network in the region, the activity hotspot, and the gang structure are revealed, thereby providing accurate information for local governance.
[0058] In one embodiment, the first clue feature is one, and the second clue feature is one; or, the first clue feature is multiple, the second clue feature is one, and the second clue feature is different from any first clue feature; or, the first clue feature is multiple, the second clue feature is multiple, and any second clue feature is different from any first clue feature; after the step of processing the fraud cases in the first fraud case set and the fraud cases in the second fraud case set according to the transmission relationship between the first fraud case set and the second fraud case set, the method further comprises: configuring a search keyword for the first fraud case set according to the first clue feature, and configuring a search keyword for the second fraud case set according to the second clue feature; wherein, when the first clue feature is one, the search keyword for the first fraud case set is one; when the first clue feature is multiple, the search keyword for the first fraud case set is multiple, and the first fraud case set is searched by inputting multiple search keywords; when the second clue feature is one, the search keyword for the second fraud case set is one; when the second clue feature is multiple, the search keyword for the second fraud case set is multiple, and the second fraud case set is searched by inputting multiple search keywords.
[0059] Specifically, the first clue feature and the second clue feature are both one, or the first clue feature is multiple and the second clue feature is one, or the first clue feature and the second clue feature are both multiple. It is only required that there is no feature overlap between the first clue feature and the second clue feature. In addition, each fraud case set is configured with a corresponding search keyword for querying and retrieving the fraud cases in the fraud case set. When a fraud case set is associated by one clue feature, one search keyword is determined by the clue feature. When a fraud case set is associated by multiple clue features, multiple search keywords are determined by the respective clue features, and the second fraud case set is searched by inputting multiple search keywords. Therefore, the fraud cases in the corresponding fraud case set can be accurately retrieved.
[0060] The above case stringing and merging method based on transmission relationship reasoning can automatically analyze multiple clue features of fraud cases, determine the first fraud case set and the second fraud case set associated by the clue features, and preliminarily screen out fraud cases with the same clue feature. Further, it is judged whether the first fraud case set and the second fraud case set contain the same fraud cases, and if so, it is reasoned that the first fraud case set and the second fraud case set have the relationship of being the same fraud nature cases. Further, the transmission relationship of the two is configured. Finally, in the stringing and merging process of fraud cases, the fraud cases in the two case sets are processed according to the transmission relationship. The whole process does not need manual review and analysis, can deeply analyze and effectively merge new fraud cases, realizes intelligent merging of fraud cases, and improves the efficiency of fraud case investigation.
[0061] In summary, the present application can achieve the following effects:
[0062] 1. Improve case solving efficiency: Through the transmission relationship reasoning analysis of the case elements, the internal relationship between seemingly independent cases can be quickly identified, and the related cases can be combined for processing, helping investigators to identify series of cases that may be committed by the same criminal gang or individual, shorten the investigation period, and speed up the case cracking speed.
[0063] 2. Enhance the accuracy of investigation: Use big data and model algorithm to deeply mine the criminal pattern, compare multiple cases through transmission relationship, such as fraud tools, case time, case location, etc., and case series and merger analysis can reveal the internal relationship between cases.
[0064] 3. Optimize resource allocation: Integrate multi-source data, avoid repeated investigation, concentrate on key points, realize the optimal allocation of resources, improve efficiency and cost-effectiveness.
[0065] 4. Promote cross-regional cooperation: String and parallel analysis across regional boundaries, promote cross-regional information sharing, improve collaborative investigation capabilities, and deal with cross-regional criminal activities.
[0066] For a case stringing and parallel method based on transmission relationship reasoning of the present application, a specific example is provided as follows:
[0067] First, set up a data integration module: design a highly compatible data integration framework, extract and schedule through NiFi, gather suspect account, APP, website information, fraud server information, third-party service ID, communication account, and other case element characteristic data. Specifically, use NiFi+Kafka combination to complete the extraction, distribution and data buffering of suspect number clues, fraud APP element clues, group member clues, relationship network and other data collected based on fraud cases. Subscribe to Kafka topic data using SparkStreaming, maintain Kafka data consumption offset to MySql database, use real-time data warehouse modeling layering technology and columnar storage engine technology, combine natural language processing, machine learning and other technologies to automatically identify and extract key feature elements, and build a feature element library.
[0068] Second, as shown in Figure 3 , the data integration module: design a highly compatible data integration framework, extract and schedule through NiFi, gather suspect account, APP, website information, fraud server information, third-party service ID, communication account, and other case element characteristic data.
[0069] Serial-parallel model and analysis algorithm: A variety of complex models including number serial-parallel, APP serial-parallel, networking serial-parallel, relationship serial-parallel are designed, and cases are serial-parallel, through serial-parallel cases, analyze the crime method, these models use real-time data warehouse modeling layer and column storage engine technology, automatically identify, extract key association clues in massive data, build serial-parallel clue library, use multiple case clue cross verification, get more "APP, number, bank card" and other suspect clues and characteristics. Use the shortest path algorithm of graph and the density-based community discovery algorithm to batch serial-parallel cases, identify potential indirect contacts, dig out hideouts and gangs, expand the achievements, and speed up the investigation of cases.
[0070] Transitive relationship reasoning: that is, the two cases of case A and case B are serial-parallel, the two cases of case B and case C are serial-parallel, and then the transitive relationship reasoning is used to combine case A, case B, and case C into a serial case.
[0071] Visual relationship graph construction: Construct a case entity relationship graph, taking cases, virtual account numbers (translated real people), crime tools, and fund accounts as nodes, and direct connections between entities (such as IP, domain name, fund transfer, server information, etc.) as edges, forming a multi-layer relationship network, which facilitates investigators to quickly understand the structure of the criminal network and guide in-depth investigation.
[0072] For example, as shown in Figure 4 , through the path transmission of the suspect's account, APP, website, application server information, virtual account number, mobile phone number, social group number, and payment account, more fraud cases are serial-parallel and combined. That is, as shown in Figure 4 , the first fraud case set obtains fraud case A and fraud case B, and the two cases of fraud case A and fraud case B are serial-parallel. The second fraud case set obtains fraud case B and fraud case C, and the two cases of fraud case B and fraud case C are serial-parallel, and then the transitive relationship reasoning is used to combine fraud case A, fraud case B, and fraud case C in the first fraud case set and the second fraud case set into a serial case.
[0073] In summary, the above-mentioned case serial-parallel method based on transitive relationship reasoning uses intelligent algorithms to quickly process and analyze massive data, constructs suspect account, APP, website, application server information, virtual account number, mobile phone number, social group number, and payment account as case serial-parallel and combined network graph, intuitively displays the structure of the criminal network, reveals hidden relationship links, makes complex case clues clear, and helps investigators quickly understand the overall situation of the case. The technical features of the present application are as follows:
[0074] 1. Enhancing case correlation identification: By comparing multiple fraud-related cases, cross-referenced clues such as fraud tools, time of involvement, and location of occurrence can reveal the internal connections between fraud-related cases. Through transitive relation reasoning, a series of cases, a piece of evidence, and a group of suspects can be identified, helping investigators to recognize a series of cases that may be committed by the same criminal gang or individual.
[0075] 2. Optimizing resource allocation: By handling fraud-related cases in series and parallel, repeated investigations on seemingly independent but actually related cases can be avoided, and resources can be concentrated on one or a few core cases for in-depth investigation, improving investigation efficiency and resource utilization.
[0076] 3. Promoting information sharing: The handling of cases in series and parallel promotes information exchange and sharing between different regions and departments, helping to build a cross-regional and cross-police collaboration mechanism to form a joint force to combat crime.
[0077] 4. Deepening understanding of criminal patterns: Analyzing multiple cases in series helps to reveal the action patterns, psychological characteristics and motives of criminals, providing a basis for developing more effective prevention and suppression strategies.
[0078] 5. Improving evidence collection and utilization capabilities: In the combined cases, some information that may be considered isolated or secondary in individual cases may become key clues after comprehensive analysis, helping to discover more evidence and strengthen the completeness and persuasiveness of the evidence chain.
[0079] 6. Speeding up case investigation: Integrating information from multiple cases can quickly identify suspects or criminal networks, shorten the case investigation period, respond to social concerns in a timely manner, and protect public safety.
[0080] It should be understood that although each step in the flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified in this document, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the drawing can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or sub-steps or stages of other steps.
[0081] The application also provides a case series and parallel device based on transitive relation reasoning. As shown in FIG. 6, the device includes a case information collection module 601, a case information preprocessing module 602, a case information storage module 603, a case information retrieval module 604, a case information analysis module 605, a case information output module 606, and a case information management module 607. Figure 5As shown, a case string and parallel device based on transmission relationship reasoning includes an acquisition module 502, an extraction module 504, an intersection analysis module 506, a determination module 508, and a configuration module 510, a string and parallel processing module 512. Among them, the acquisition module 502 is used to acquire multiple fraud cases from multiple network-related cases; the extraction module 504 is used to extract multiple clue features of each fraud case, and the multiple clue features include any one or more of the communication clues, the fraud network behavior clues, the fund transaction clues, and the fraud application program clues; the intersection analysis module 506 is used to analyze the intersection of the multiple clue features of each fraud case, and determine the first fraud case set and the second fraud case set that exist clue feature association, wherein the first fraud case set exists the intersection of the first clue feature between the fraud cases, the second fraud case set exists the intersection of the second clue feature between the fraud cases, and the first clue feature and the second clue feature are different; the determination module 508 is used to determine that the first fraud case set and the second fraud case set exist the relationship of the same fraud nature case if the first fraud case set and the second fraud case set exist one or more same fraud cases; the configuration module 510 is used to configure the transmission relationship of the first fraud case set and the second fraud case set if it is identified that the first fraud case set and the second fraud case set belong to the same fraud nature case relationship; the string and parallel processing module 512 is used to process the fraud cases in the first fraud case set and the fraud cases in the second fraud case set according to the transmission relationship of the first fraud case set and the second fraud case set.
[0082] In one of the embodiments, the communication clues include any one or more of the mobile phone number and the communication contact in the mobile phone number, the instant messaging number and the communication contact in the instant messaging number, the email number and the email contact in the email number; and / or, the fund transaction clues include the fund transaction account number and the transfer account number; and / or, the fraud network behavior clues include one or more of the fraud website domain name, the fraud website IP address, and the third party service ID.
[0083] In one of the embodiments, the fraud application program clues include the information of the fraud application program and the associated information of the fraud application program, the information of the fraud application program includes the package name and the packaging information of the fraud application program, the MD5 value of the fraud application program file, the signature of the fraud application program, and the MD5 value of the fraud application program icon, the associated information of the fraud application program includes the distribution website of the fraud application program, the distribution download address of the fraud application program, the application server IP of the fraud application program, the application server domain name of the fraud application program, the packaging service system of the fraud application program, the customer service system of the fraud application program, and the push platform of the fraud application program.
[0084] In one of the embodiments, the multiple clue features of each fraud case are subjected to intersection analysis, and a first fraud case set with associated clue features is determined, including: performing intersection analysis on the information of the fraud application of each fraud case and the associated information of the fraud application; if the package name and packaging information of the fraud application of any two fraud cases are associated, determining whether any one of the MD5 value of the fraud application file, the signature of the fraud application and the MD5 value of the fraud application icon of the fraud application of any two fraud cases is associated, if any one or more are associated, determining that any two fraud cases are fraud cases in the first fraud case set, wherein the first clue feature is the fraud application clue; if the package name and packaging information of the fraud application of any two fraud cases are not associated, determining whether any one of the distribution website of the fraud application, the distribution download address of the fraud application, the application server IP of the fraud application, the application server domain name of the fraud application, the packaging service system of the fraud application, the customer service system of the fraud application and the push platform of the fraud application of any two fraud cases is associated, if any one or more are associated, determining that any two fraud cases are fraud cases in the first fraud case set, wherein the first clue feature is the fraud application clue.
[0085] In one of the embodiments, the multiple clue features of each fraud case are subjected to intersection analysis, and a second fraud case set with associated clue features is determined, including: performing intersection analysis on the multiple clue features of each fraud case, determining multiple fraud cases with associated second clue features, and constructing a second fraud case set according to the multiple fraud cases with associated second clue features, wherein the second clue feature is any one or more of the communication clue, the fraud network behavior clue and the fund transaction clue.
[0086] In one of the embodiments, after the fraud cases in the first fraud case set and the fraud cases in the second fraud case set are subjected to stringing and parallel processing according to the transmission relationship, the method further includes: obtaining the collection unit information of each target fraud case obtained by the stringing and parallel processing; and dividing the multiple target fraud cases according to the attribution, and obtaining the target fraud case set of each attribution according to the collection unit information of each target fraud case.
[0087] In one of the embodiments, the first clue feature is one, the second clue feature is one; or, the first clue feature is multiple, the second clue feature is one and different from any first clue feature; or, the first clue feature is multiple, the second clue feature is multiple, and any second clue feature is different from any first clue feature.
[0088] The case series and parallel device based on the transmission relationship reasoning further comprises a configuration module configured to configure a retrieval keyword of the first fraud case set according to the first clue feature, and configure a retrieval keyword of the second fraud case set according to the second clue feature; wherein when the first clue feature is one, the retrieval keyword of the first fraud case set is one, when the first clue feature is multiple, the retrieval keyword of the first fraud case set is multiple and the first fraud case set is retrieved by inputting multiple retrieval keywords; when the second clue feature is one, the retrieval keyword of the second fraud case set is one, when the second clue feature is multiple, the retrieval keyword of the second fraud case set is multiple and the second fraud case set is retrieved by inputting multiple retrieval keywords.
[0089] The specific limitation of the case series and parallel device based on the transmission relationship reasoning can refer to the limitation of the case series and parallel method based on the transmission relationship reasoning, which will not be repeated here. Each module in the case series and parallel device based on the transmission relationship reasoning can be realized by software, hardware and their combination in whole or in part. Each module can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to call and execute the operation corresponding to each module by the processor.
[0090] In one embodiment, a computer device is provided, which can be Figure 1 The internal structure diagram of the server shown can be as shown in Figure 6 The computer device comprises a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is used to provide computing and control capability. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the clue features of multiple fraud cases. The network interface of the computer device is used to communicate with the external system platform through network connection. The computer program is executed by the processor to implement a case series and parallel method based on the transmission relationship reasoning.
[0091] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implements the following steps when executing the computer program: obtaining a plurality of fraud cases from a plurality of network-related cases; extracting a plurality of clue features of each fraud case, the plurality of clue features comprising any one or more of a communication clue, a fraud network behavior clue, a fund transaction clue, and a fraud application program clue; performing intersection analysis on the plurality of clue features of each fraud case to determine a first fraud case set and a second fraud case set in which there is a clue feature association, wherein there is an intersection of a first clue feature between fraud cases in the first fraud case set, there is an intersection of a second clue feature between fraud cases in the second fraud case set, and the first clue feature and the second clue feature are different; if the first fraud case set and the second fraud case set have one or more same fraud cases, determining that the first fraud case set and the second fraud case set have a relationship of belonging to the same fraud nature case; if it is identified that the first fraud case set and the second fraud case set belong to the same fraud nature case relationship, configuring a transitive relationship of the first fraud case set and the second fraud case set; and performing stringing and parallel processing on fraud cases in the first fraud case set and fraud cases in the second fraud case set according to the transitive relationship of the first fraud case set and the second fraud case set.
[0092] In one embodiment, the communication clue comprises any one or more of a mobile phone number and a communication contact in the mobile phone number, an instant messaging number and a communication contact in the instant messaging number, an email number and an email contact in the email number; and / or, the fund transaction clue comprises a fund transaction account number and a transfer account number; and / or, the fraud network behavior clue comprises one or more of a fraud website domain name, a fraud website IP address, and a third-party service ID.
[0093] In one embodiment, the fraud application program clue comprises information of a fraud application program and associated information of the fraud application program, the information of the fraud application program comprises a package name and packaging information of the fraud application program, an MD5 value of a fraud application program file, a signature of the fraud application program, and an MD5 value of a fraud application program icon, and the associated information of the fraud application program comprises a distribution website of the fraud application program, a distribution download address of the fraud application program, an application server IP of the fraud application program, an application server domain name of the fraud application program, a packaging service system of the fraud application program, a customer service system of the fraud application program, and a push platform of the fraud application program.
[0094] In one of the embodiments, when the processor executes the computer program to implement the step of performing the intersection analysis on the multiple clue features of each fraud case to determine the first fraud case set in which the clue features are associated, the following steps are implemented: performing the intersection analysis on the information of the fraud application of each fraud case and the associated information of the fraud application; if the package name and the packaging information of the fraud application of any two fraud cases are associated, determining whether any one of the MD5 value of the fraud application file, the signature of the fraud application and the MD5 value of the icon of the fraud application of the two fraud cases is associated, and if any one or more of them are associated, determining that the two fraud cases are both the fraud cases in the first fraud case set, wherein the first clue feature is the fraud application clue; if the package name and the packaging information of the fraud application of any two fraud cases are not associated, determining whether any one of the distribution website of the fraud application, the distribution download address of the fraud application, the application server IP of the fraud application, the application server domain name of the fraud application, the packaging service system of the fraud application, the customer service system of the fraud application and the push platform of the fraud application of the two fraud cases is associated, and if any one or more of them are associated, determining that the two fraud cases are both the fraud cases in the first fraud case set, wherein the first clue feature is the fraud application clue.
[0095] In one of the embodiments, when the processor executes the computer program to implement the step of performing the intersection analysis on the multiple clue features of each fraud case to determine the second fraud case set in which the clue features are associated, the following steps are implemented: performing the intersection analysis on the multiple clue features of each fraud case to determine the multiple fraud cases in which the second clue features are associated, and constructing the second fraud case set according to the multiple fraud cases in which the second clue features are associated, wherein the second clue feature is any one or more of the communication clue, the fraud network behavior clue and the fund transaction clue.
[0096] In one of the embodiments, when the processor executes the computer program, the following steps are implemented: obtaining the collection unit information of each target fraud case obtained through the string and parallel case processing; and dividing the multiple target fraud cases according to the attribution place according to the collection unit information of each target fraud case to obtain the target fraud case set of each attribution place.
[0097] In one of the embodiments, the first clue feature is one, and the second clue feature is one; or, the first clue feature is multiple, the second clue feature is one, and the second clue feature is different from any of the first clue features; or, the first clue feature is multiple, the second clue feature is multiple, and any of the second clue features is different from any of the first clue features; when the computer program is executed by the processor, the following steps are implemented: configuring the search keywords of the first set of fraud cases according to the first clue feature, and configuring the search keywords of the second set of fraud cases according to the second clue feature; wherein, when the first clue feature is one, the search keywords of the first set of fraud cases are one; when the first clue feature is multiple, the search keywords of the first set of fraud cases are multiple, and the first set of fraud cases is searched by inputting multiple search keywords; when the second clue feature is one, the search keywords of the second set of fraud cases are one; when the second clue feature is multiple, the search keywords of the second set of fraud cases are multiple, and the second set of fraud cases is searched by inputting multiple search keywords.
[0098] In one of the embodiments, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: obtaining multiple pieces of fraud cases from multiple pieces of network-related cases; extracting multiple clue features of each fraud case, the multiple clue features including any of the following: communication clues, fraud network behavior clues, fund transaction clues, and fraud application program clues; performing intersection analysis on the multiple clue features of each fraud case to determine a first set of fraud cases and a second set of fraud cases with associated clue features, wherein the first set of fraud cases has an intersection of the first clue features between the fraud cases, the second set of fraud cases has an intersection of the second clue features between the fraud cases, and the first clue features and the second clue features are different; if the first set of fraud cases and the second set of fraud cases have one or more same fraud cases, it is determined that the first set of fraud cases and the second set of fraud cases have a relationship of belonging to the same fraud nature case; if it is identified that the first set of fraud cases and the second set of fraud cases belong to the same fraud nature case, a transmission relationship of the first set of fraud cases and the second set of fraud cases is configured; and the fraud cases in the first set of fraud cases and the fraud cases in the second set of fraud cases are handled in series and in parallel according to the transmission relationship of the first set of fraud cases and the second set of fraud cases.
[0099] In one of the embodiments, the communication clues include any of the following: a mobile phone number and communication contacts in the mobile phone number, an instant messaging number and communication contacts in the instant messaging number, an email number and email contacts in the email number; and / or, the fund transaction clues include fund transaction account numbers and transfer account numbers; and / or, the fraud network behavior clues include one or more of the following: fraud website domain names, fraud website IP addresses, and third-party service IDs.
[0100] In one of the embodiments, the fraud application clue application includes information of the fraud application and associated information of the fraud application, the information of the fraud application includes package name and packaging information of the fraud application, MD5 value of the fraud application file, signature of the fraud application and MD5 value of the fraud application icon, the associated information of the fraud application includes distribution website of the fraud application, distribution download address of the fraud application, application server IP of the fraud application, application server domain name of the fraud application, packaging service system of the fraud application, customer service system of the fraud application and push platform of the fraud application.
[0101] In one of the embodiments, when the computer program is executed by the processor to realize the step of performing intersection analysis on the multiple clue characteristics of each fraud case to determine the first fraud case set with the clue characteristic association, the following steps are realized: performing intersection analysis on the information of the fraud application and the associated information of the fraud application of each fraud case; if the package name and the packaging information of the fraud application of any two fraud cases are associated, determining whether any one or more of the MD5 value of the fraud application file, the signature of the fraud application and the MD5 value of the fraud application icon are associated, if any one or more are associated, determining that the any two fraud cases are the fraud cases in the first fraud case set, wherein the first clue characteristic is the fraud application clue; if the package name and the packaging information of the fraud application of any two fraud cases are not associated, determining whether any one or more of the distribution website of the fraud application, the distribution download address of the fraud application, the application server IP of the fraud application, the application server domain name of the fraud application, the packaging service system of the fraud application, the customer service system of the fraud application and the push platform of the fraud application are associated, if any one or more are associated, determining that the any two fraud cases are the fraud cases in the first fraud case set, wherein the first clue characteristic is the fraud application clue.
[0102] In one of the embodiments, when the computer program is executed by the processor to realize the step of performing intersection analysis on the multiple clue characteristics of each fraud case to determine the second fraud case set with the clue characteristic association, the following steps are realized: performing intersection analysis on the multiple clue characteristics of each fraud case to determine the multiple fraud cases with the second clue characteristic association, constructing the second fraud case set according to the multiple fraud cases with the second clue characteristic association, and the second clue characteristic is any one or more of the communication clue, the fraud network behavior clue and the fund transaction clue.
[0103] In one of the embodiments, the computer program, when executed by the processor, implements the following steps: obtaining collection unit information of each target fraud case obtained through stringing and parallel processing; and dividing the plurality of target fraud cases according to the collection unit information of each target fraud case and according to the place of origin, to obtain a set of target fraud cases of each place of origin.
[0104] In one of the embodiments, the first clue feature is one, the second clue feature is one; or, the first clue feature is multiple, the second clue feature is one and the second clue feature is different from any first clue feature; or, the first clue feature is multiple, the second clue feature is multiple, any second clue feature is different from any first clue feature; the computer program, when executed by the processor, implements the following steps: configuring a retrieval keyword of the first fraud case set according to the first clue feature, and configuring a retrieval keyword of the second fraud case set according to the second clue feature; wherein, when the first clue feature is one, the retrieval keyword of the first fraud case set is one, when the first clue feature is multiple, the retrieval keyword of the first fraud case set is multiple and the first fraud case set is retrieved by inputting multiple retrieval keywords; when the second clue feature is one, the retrieval keyword of the second fraud case set is one, when the second clue feature is multiple, the retrieval keyword of the second fraud case set is multiple and the second fraud case set is retrieved by inputting multiple retrieval keywords.
[0105] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0106] Any combination of the technical features in the above embodiments can be made. For the sake of brevity, not all possible combinations are described, however, it should be understood that the scope of the specification includes all possible combinations.
[0107] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, some modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.
Claims
1. A case concatenation and merging method based on transitive relationship reasoning, characterized in that: The method comprises: Obtain multiple fraud cases from multiple Internet-related cases; Extracting multiple clue features of each fraud case, wherein the multiple clue features include any multiple of communication clues, fraudulent network behavior clues, financial transaction clues, and fraudulent application clues; Performing an intersection analysis on multiple clue features of each fraud case to determine a first fraud case set and a second fraud case set having associated clue features, wherein the fraud cases in the first fraud case set have an intersection of first clue features, and the fraud cases in the second fraud case set have an intersection of second clue features, and the first clue features and the second clue features are different; If the first fraud case set and the second fraud case set contain one or more identical fraud cases, then it is determined that the first fraud case set and the second fraud case set are cases of the same fraud nature; If it is identified that the first fraud case set and the second fraud case set are cases of the same fraud nature, a transfer relationship is configured between the first fraud case set and the second fraud case set. The transfer relationship is used to indicate that the same fraud case exists in the two fraud case sets and to transfer the same fraud attribute between the fraud cases during case concatenation and merging. combining the fraud cases in the first fraud case set and the fraud cases in the second fraud case set according to the transfer relationship between the first fraud case set and the second fraud case set; The intersection analysis of multiple clue features of each fraud case is performed to determine the first fraud case set with associated clue features, including: Conduct intersection analysis on the information of fraudulent applications involved in each fraud case and the related information of fraudulent applications; If the package names and packaging information of the fraudulent applications of any two fraud cases are correlated, then determine whether any one of the MD5 values of the fraudulent application files, the signatures of the fraudulent applications, and the MD5 values of the fraudulent application icons of the any two fraud cases is correlated; if any one or more correlations exist, then determine that the any two fraud cases are both fraudulent cases in the first fraud case set, and the first clue feature is the fraudulent application clue; If there is no correlation between the package names and packaging information of the fraudulent applications in any two fraud cases, then determine whether there is a correlation between any one of the distribution URLs of the fraudulent applications, the distribution download addresses of the fraudulent applications, the application server IPs of the fraudulent applications, the application server domain names of the fraudulent applications, the packaging service systems of the fraudulent applications, the customer service systems of the fraudulent applications, and the push platforms of the fraudulent applications in any two fraud cases. If there is any one or more correlations, then determine that the any two fraud cases are both fraudulent cases in the first fraud case concentration, and the first clue feature is the fraudulent application clue.
2. The method according to claim 1, characterized in that The communication clues include any one or more of a mobile phone number and the communication contacts in the mobile phone number, an instant messaging number and the communication contacts in the instant messaging number, an email address and the email contacts in the email address number; and / or, The fund transaction clues include fund transaction account numbers and transfer account numbers; and / or, The fraudulent network behavior clues include one or more of the fraudulent website domain name, the fraudulent website IP address, and the third-party service ID.
3. The method according to claim 1, characterized in that The intersection analysis of multiple clue features of each fraud case is performed to determine a second set of fraud cases with associated clue features, including: An intersection analysis is performed on multiple clue features of each fraud-related case to determine multiple fraud-related cases associated with a second clue feature. The second fraud-related case set is constructed based on the multiple fraud-related cases associated with the second clue feature, where the second clue feature is any one or more of communication clues, fraud-related network behavior clues, and financial transaction clues.
4. The method according to claim 1, wherein After the step of combining the fraud cases in the first fraud case set and the fraud cases in the second fraud case set according to the transfer relationship between the first fraud case set and the second fraud case set, the method further includes: Obtain the information of the collecting units of each target fraud case obtained through case consolidation and combination processing; According to the information of the collection unit of each target fraud case, the multiple target fraud cases are divided according to the location of attribution to obtain a target fraud case set for each location.
5. The method according to claim 1, wherein There is one first clue feature and one second clue feature; or, there are multiple first clue features and one second clue feature, and the second clue feature is different from any of the first clue features; or, there are multiple first clue features and multiple second clue features, and any of the second clue features is different from any of the first clue features; The method further includes: configuring search keywords for the first set of fraud-related cases based on the first clue feature, and configuring search keywords for the second set of fraud-related cases based on the second clue feature; When the first clue feature is one, the search keyword for the first fraud case set is one; when the first clue feature is multiple, the search keyword for the first fraud case set is multiple and the first fraud case set is searched by inputting multiple search keywords; When the second clue feature is one, the search keyword for the second set of fraud cases is one; when the second clue feature is multiple, the search keywords for the second set of fraud cases are multiple and the second set of fraud cases is retrieved by inputting multiple search keywords.
6. A case concatenation and paralleling device based on transitive relationship reasoning, characterized in that: The device comprises: An acquisition module is used to obtain multiple fraud-related cases from multiple Internet-related cases; An extraction module, configured to extract multiple clue features of each fraud case, wherein the multiple clue features include any one of communication clues, fraudulent network behavior clues, financial transaction clues, and fraudulent application clues; an intersection analysis module, configured to perform intersection analysis on multiple clue features of each fraud case to determine a first fraud case set and a second fraud case set having associated clue features, wherein the fraud cases in the first fraud case set have an intersection of first clue features, and the fraud cases in the second fraud case set have an intersection of second clue features, and the first clue features and the second clue features are different; a determination module, configured to determine that the first fraud case set and the second fraud case set are cases of the same fraud nature if the first fraud case set and the second fraud case set have one or more identical fraud cases; a configuration module configured to, upon identifying that the first fraud case set and the second fraud case set are cases of the same fraud nature, configure a transfer relationship between the first fraud case set and the second fraud case set, wherein the transfer relationship is used to indicate that the same fraud case exists in the two fraud case sets and to transfer the same fraud attribute between the fraud cases during case concatenation and merging; a serial and parallel processing module, configured to serially and parallelize the fraud cases in the first fraud case set and the fraud cases in the second fraud case set according to the transfer relationship between the first fraud case set and the second fraud case set; The intersection analysis of multiple clue features of each fraud case is performed to determine the first fraud case set with associated clue features, including: Conduct intersection analysis on the information of fraudulent applications involved in each fraud case and the related information of fraudulent applications; If the package names and packaging information of the fraudulent applications of any two fraud cases are correlated, then determine whether any one of the MD5 values of the fraudulent application files, the signatures of the fraudulent applications, and the MD5 values of the fraudulent application icons of the any two fraud cases is correlated; if any one or more correlations exist, then determine that the any two fraud cases are both fraudulent cases in the first fraud case set, and the first clue feature is the fraudulent application clue; If there is no correlation between the package names and packaging information of the fraudulent applications in any two fraud cases, then determine whether there is a correlation between any one of the distribution URLs of the fraudulent applications, the distribution download addresses of the fraudulent applications, the application server IPs of the fraudulent applications, the application server domain names of the fraudulent applications, the packaging service systems of the fraudulent applications, the customer service systems of the fraudulent applications, and the push platforms of the fraudulent applications in any two fraud cases. If there is any one or more correlations, then determine that the any two fraud cases are both fraudulent cases in the first fraud case concentration, and the first clue feature is the fraudulent application clue.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Video aided analytical method for criminal investigation and case detection and system
CN101655860A
Serial-parallel case correlation analysis method and device, equipment and storage medium
CN111753872A