Risk detection method, device, apparatus and storage medium
By extracting semantics and constructing storylines from user communication data, and combining multimodal large models and user profiles, the limitations of traditional telecom fraud detection are overcome, enabling efficient identification and response to complex fraud strategies.
Patent Information
- Application Number
- CN202311658115.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-12-05
AI Technical Summary
Existing technologies are unable to effectively combat complex and ever-changing telecom fraud strategies. Traditional prevention methods, such as keyword filtering and single-information analysis, have limitations, cannot identify complex semantic information, and detection solutions are delayed and have high maintenance costs.
By extracting semantics from user communication data, constructing user communication storylines, and combining multimodal big data models and user profiles, risk detection is performed to identify fraud risks.
It has achieved accurate detection of complex and ever-changing fraudulent activities, reduced false positives, and improved detection efficiency and accuracy.
Smart Images

Figure CN120128345B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a risk detection method, apparatus, device, and storage medium. Background Technology
[0002] In recent years, with the widespread adoption of mobile communication networks, telecommunications fraud incidents have continued to rise, seriously threatening users' lives and property safety. Traditional prevention methods, such as keyword filtering, sample labeling, and single-information analysis, have several limitations, including the inability to identify complex semantic information, long training cycles, delayed detection schemes, high maintenance costs, and data isolation issues, making them difficult to effectively counter the complex and ever-changing strategies of telecommunications fraud. Summary of the Invention
[0003] The main objective of this invention is to provide a risk detection method, apparatus, device, and storage medium, aiming to solve the technical problem that existing technologies cannot effectively cope with complex and ever-changing telecommunications fraud strategies.
[0004] To achieve the above objectives, the present invention provides a risk detection method, the method comprising the following steps:
[0005] Semantic extraction is performed on the collected user communication data to obtain communication semantic data;
[0006] A user communication storyline is constructed based on the communication semantic data and the user communication data. The user communication storyline is a data stream constructed based on at least one communication data corresponding to the same user and the same semantic data.
[0007] Risk detection is performed based on the communication semantic data and the user's communication storyline to determine whether there is a risk of fraud.
[0008] Optionally, the user communication data is multimodal data;
[0009] The step of extracting semantic data from the collected user communication data to obtain communication semantic data includes:
[0010] Semantic extraction is performed on the collected user communication data using a multimodal large model to obtain communication semantic data. The multimodal large model is a pre-trained model that performs semantic extraction on various types of data.
[0011] Optionally, the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk includes:
[0012] Obtain the target user identifier corresponding to the user communication storyline;
[0013] Find the target user profile corresponding to the target user identifier;
[0014] The existence of fraud risk is determined based on the target user profile, the user communication storyline, and the communication semantic data.
[0015] Optionally, the step of determining whether there is a risk of fraud based on the target user profile, the user communication storyline, and the communication semantic data includes:
[0016] The target user profile is updated based on the user communication storyline and the communication semantic data to obtain an updated user profile;
[0017] The target user profile is compared with the updated user profile to determine the profile difference.
[0018] If the difference in the portrait is greater than a preset difference threshold, it is determined that there is a risk of fraud.
[0019] Optionally, the step of comparing the target user profile with the updated user profile to determine the profile difference includes:
[0020] The user communication storyline is examined to determine if any risky behavior exists;
[0021] If risky behavior is found, the target user profile is compared with the updated user profile to determine the degree of profile difference.
[0022] Optionally, the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk includes:
[0023] The user communication storyline is marked based on the communication semantic data to obtain a semantically marked storyline;
[0024] Obtain the target user corresponding to the user communication storyline;
[0025] Search the storyline repository for semantically tagged storylines for users other than the target user to obtain comparison storylines;
[0026] Construct broadcast event information based on the semantically labeled storyline and the contrasting storyline;
[0027] The presence of fraud risk is determined based on the broadcast event information.
[0028] Optionally, the step of constructing broadcast event information based on the semantically marked storyline and the contrasting storyline includes:
[0029] Semantic statistics are performed on the semantically marked storyline and the contrasting storyline to obtain the intersection semantic information, wherein the intersection semantic information consists of the same or similar semantics possessed by at least two different semantically marked storylines;
[0030] The number of associated users corresponding to each intersection semantic information is determined by statistical analysis of the semantically labeled storylines.
[0031] Broadcast event information is constructed based on the intersection semantic information and the number of associated users.
[0032] Optionally, the step of determining whether there is a risk of fraud based on the broadcast event information includes:
[0033] Extract the number of associated users from the broadcast event information;
[0034] If the number of associated users is greater than or equal to a preset risk threshold, then a fraud risk is determined to exist.
[0035] Optionally, after the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk, the method further includes:
[0036] When there is a risk of fraud, determine the type of risk.
[0037] Search the response strategy library for the corresponding risk type;
[0038] Risk warnings will be issued based on the aforementioned response strategies.
[0039] Optionally, the step of issuing a risk warning based on the response strategy includes:
[0040] Obtain risk alert permission, which indicates the alert methods currently allowed to be used;
[0041] The response strategies are filtered based on the risk alert permissions to determine the target response strategy;
[0042] Risk warnings will be issued based on the aforementioned target response strategies.
[0043] Optionally, the step of constructing a user communication storyline based on the communication semantic data and the user communication data includes:
[0044] The communication semantic data is clustered to obtain at least one cluster.
[0045] The user communication data corresponding to each cluster is sorted in ascending order of the corresponding communication time to obtain the user communication storyline.
[0046] Furthermore, to achieve the above objectives, the present invention also proposes a risk detection device, which includes the following modules:
[0047] The extraction module is used to extract semantics from the collected user communication data to obtain communication semantic data.
[0048] The construction module is used to construct a user communication storyline based on the communication semantic data and the user communication data. The user communication storyline is a data stream constructed based on at least one communication data corresponding to the same semantic data of the same user.
[0049] The detection module is used to perform risk detection based on the communication semantic data and the user communication storyline to determine whether there is a risk of fraud.
[0050] Optionally, the user communication data is multimodal data;
[0051] The extraction module is also used to perform semantic extraction on the collected user communication data through a multimodal large model to obtain communication semantic data. The multimodal large model is a pre-trained model that performs semantic extraction on multiple different types of data.
[0052] Optionally, the detection module is further configured to obtain the target user identifier corresponding to the user communication storyline; find the target user profile corresponding to the target user identifier; and determine whether there is a fraud risk based on the target user profile, the user communication storyline, and the communication semantic data.
[0053] Optionally, the detection module is further configured to update the target user profile based on the user communication storyline and the communication semantic data to obtain an updated user profile; compare the target user profile with the updated user profile to determine the profile difference degree; if the profile difference degree is greater than a preset difference threshold, it is determined that there is a fraud risk.
[0054] Optionally, the detection module is further configured to detect the user communication storyline to determine whether there is any risky behavior; if there is risky behavior, the target user profile is compared with the updated user profile to determine the profile difference.
[0055] Optionally, the detection module is further configured to: mark the user communication storyline according to the communication semantic data to obtain a semantically marked storyline; obtain the target user corresponding to the user communication storyline; search for semantically marked storylines corresponding to other users besides the target user in the storyline repository to obtain a comparison storyline; construct broadcast event information based on the semantically marked storyline and the comparison storyline; and determine whether there is a fraud risk based on the broadcast event information.
[0056] Optionally, the detection module is further configured to perform semantic statistics on the semantically marked storyline and the comparison storyline to obtain intersection semantic information, wherein the intersection semantic information consists of the same or similar semantics possessed by at least two different semantically marked storylines; perform user attribution statistics on the semantically marked storylines corresponding to the intersection semantic information to determine the number of associated users corresponding to each intersection semantic information; and construct broadcast event information based on the intersection semantic information and the number of associated users.
[0057] In addition, to achieve the above objectives, the present invention also proposes a risk detection device, which includes: a processor, a memory, and a risk detection program stored in the memory and executable on the processor. When the risk detection program is executed by the processor, it implements the steps of the risk detection method described above.
[0058] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium storing a risk detection program, which, when executed, implements the steps of the risk detection method described above.
[0059] This invention extracts semantic data from collected user communication data to obtain communication semantic data. Based on this semantic data and user communication data, it constructs user communication storylines, which are data streams built from at least one communication data point corresponding to the same semantic meaning for the same user. Risk detection is then performed based on the semantic data and user communication storylines to determine the presence of fraud risk. Because it does not rely on single keyword filtering or single information analysis, but rather on constructing user communication storylines based on corresponding communication semantic information, and then performing overall behavioral analysis and information statistics based on these storylines, it provides a more accurate analysis of the user's behavior before and after the event, thus ensuring that even complex and ever-changing fraudulent behaviors can be detected. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the structure of an electronic device in the hardware operating environment involved in the embodiments of the present invention;
[0061] Figure 2 This is a flowchart illustrating the first embodiment of the risk detection method of the present invention;
[0062] Figure 3 This is a flowchart illustrating the second embodiment of the risk detection method of the present invention;
[0063] Figure 4 This is a flowchart illustrating the second embodiment of the risk detection method of the present invention;
[0064] Figure 5 This is a schematic diagram of a risk detection and processing flow according to an embodiment of the present invention;
[0065] Figure 6 This is a structural block diagram of the first embodiment of the risk detection device of the present invention.
[0066] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0067] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0068] Reference Figure 1 , Figure 1 This is a schematic diagram of the risk detection device structure for the hardware operating environment involved in the embodiments of the present invention.
[0069] like Figure 1 As shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0070] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0071] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a risk detection program.
[0072] exist Figure 1In the electronic device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present invention can be set in the risk detection device. The electronic device calls the risk detection program stored in the memory 1005 through the processor 1001 and executes the risk detection method provided in the embodiment of the present invention.
[0073] This invention provides a risk detection method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of a risk detection method according to the present invention.
[0074] In this embodiment, the risk detection method includes the following steps:
[0075] Step S10: Extract semantic information from the collected user communication data to obtain communication semantic data.
[0076] It should be noted that the execution subject of this embodiment can be the risk detection device, which can be a personal computer, server or other electronic device, or other device that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the risk detection device is used as an example to illustrate the risk detection method of the present invention.
[0077] It should be noted that user communication data can be data generated during user communications, such as SMS, voice messages, MMS, or chat logs from specific software applications. Risk detection devices can collect user communication data after obtaining the user's authorization.
[0078] In practical use, since there may be many users involved when collecting user communication data, the amount of user communication data collected may be large. Processing all of it directly may be slow and the performance of the device may not be able to support it. Therefore, after collecting user communication data, it can be stored (such as in a specific database or storage space). Then the risk detection device can read the stored user communication data in batches or in quantitative quantities and process it sequentially.
[0079] In specific implementation, semantic extraction is performed on the collected user communication data to obtain communication semantic data. This can be done by using semantic extraction algorithms (such as TextRank algorithm, KeyBert algorithm, or similar algorithms) or pre-trained semantic extraction models (such as deep learning models or neural network models) to extract semantics from the collected user communication data.
[0080] The communication semantic data corresponds one-to-one with the user communication data. The communication semantic data may include semantic content and classification vectors used to represent the content association categories (such as classification vectors for daily necessities, luxury goods, daily electronic products, high-end electronic products, etc.). Of course, other data may also be included if necessary, but this embodiment does not limit this.
[0081] In practical applications, to ensure the data can be targeted at complex and ever-changing fraud strategies, the collected user communication data can be multimodal data, meaning that the collected user communication data simultaneously includes various types of data such as text (e.g., SMS), images (e.g., MMS or screenshots of software chat interfaces), and voice (e.g., call recordings). In this case, to ensure proper semantic extraction of the user communication data, step S10 in this embodiment may include:
[0082] Semantic extraction is performed on the collected user communication data using a multimodal large model to obtain communication semantic data.
[0083] It should be noted that a multimodal large model can be a model pre-trained on a model training set constructed from multimodal data, capable of semantic extraction from multiple different types of data (i.e., different data types). Specifically, a multimodal large model can be a model that has been partially pre-trained using transfer learning and then transferred to the semantic extraction domain for further training, thus saving model training time.
[0084] In practical applications, a multimodal large model can also be a large model composed of multiple different sub-models. A sub-model can be trained on a model training set constructed from one type of data and can perform semantic extraction on one type of user communication data. In this case, the multimodal large model can be used to extract semantics from the collected user communication data to obtain communication semantic data. Alternatively, the collected user communication data can be input into the multimodal large model, which will then distribute it to the corresponding sub-model for semantic extraction processing according to the data type of the user communication data, and generate communication semantic data in a unified format.
[0085] Step S20: Construct a user communication storyline based on the communication semantic data and the user communication data.
[0086] It should be noted that a user communication storyline can be a data stream constructed based on at least one communication data corresponding to the same user and the same semantics. For example, a user communication storyline can be obtained by constructing a data stream based on all the communication data involved by user A in purchasing daily necessities B.
[0087] In this context, the same user may correspond to multiple different semantics, meaning that the same user may correspond to multiple user communication storylines.
[0088] In a specific implementation, in order to quickly construct the user communication storyline, step S20 of this embodiment may include:
[0089] The communication semantic data is clustered to obtain at least one cluster.
[0090] The user communication data corresponding to each cluster is sorted in ascending order of the corresponding communication time to obtain the user communication storyline.
[0091] It should be noted that clustering communication semantic data to obtain at least one cluster can be achieved by grouping communication semantic data with the same classification vector and similar or identical semantic content into the same cluster, thereby obtaining at least one cluster.
[0092] It is understandable that after obtaining the clusters, the user communication data corresponding to the same cluster is all the communication data of the same user for the same semantic (or similar semantic) purpose. At this time, a data stream can be constructed based on it. In order to ensure accurate analysis, the user communication data corresponding to the clusters can be sorted in ascending order of the corresponding communication time to clarify the chronological order of each user's communication data, thereby constructing the user communication storyline.
[0093] In practice, the collected user communication data may include user communication data from multiple different users. In this case, the user communication data can be classified first to obtain the user communication data corresponding to each user. Then, each user's data can be processed separately to obtain at least one user communication storyline corresponding to each user.
[0094] Step S30: Perform risk detection based on the communication semantic data and the user communication storyline to determine whether there is a risk of fraud.
[0095] In practical use, risk detection based on communication semantic data and user communication storylines can determine whether there is a risk of fraud. This can be achieved by conducting overall behavior detection based on communication semantic data and user communication storylines to determine whether the user has been misled or to detect whether a malicious team is conducting a group fraud, thereby determining whether there is a risk of fraud.
[0096] In a specific implementation, to ensure the security of users' assets as much as possible, after step S30 in this embodiment, the following may also be included:
[0097] When there is a risk of fraud, determine the type of risk.
[0098] Search the response strategy library for the corresponding risk type;
[0099] Risk warnings will be issued based on the aforementioned response strategies.
[0100] It should be noted that, depending on the type of fraud, risk types can be categorized into several different types, such as telecommunications fraud, information manipulation, and information simulation. The risk type can be determined based on the user's behavior within the user's communication storyline where fraud risk is identified, as well as the type of user communication data used to construct the user's communication storyline.
[0101] For example, if the user's communication data is text-based, and the user's behavior is to receive a message and then provide a response, ultimately sending personal account information, then the risk type can be determined as information inducement.
[0102] It should be noted that the response strategy library can be a database that stores different response strategies. In the response strategy library, each risk type corresponds to at least one response strategy.
[0103] In practical use, response strategies can include procedures for alerting users to fraud risks, such as message notifications, call alerts, service delays, and service interruptions. Providing risk alerts based on the response strategy can involve executing the procedures documented within that strategy to issue a risk warning.
[0104] In practical applications, since there may be multiple different response strategies for different types of risks, risk warnings based on response strategies can be implemented by executing the corresponding processes of each response strategy to issue an alert; alternatively, the response strategy with the highest execution priority can be selected as the target strategy, and the processing process recorded in the target strategy can be executed to issue a risk warning. The execution priority of each response strategy can be preset by the administrator of the risk detection equipment.
[0105] Furthermore, since the execution procedures may differ depending on the company or department to which the risk detection equipment belongs, in order to ensure reasonable warnings, the steps for issuing risk warnings based on the response strategy described in this embodiment may include:
[0106] Obtain risk alert permissions;
[0107] The response strategies are filtered based on the risk alert permissions to determine the target response strategy;
[0108] Risk warnings will be issued based on the aforementioned target response strategies.
[0109] It should be noted that risk alert permissions are used to characterize the permitted alert methods. For example, if the department to which the risk detection device belongs is an official department, then the permitted alert methods include various means such as telephone, voice, and SMS, and the risk alert permissions can be "phone, voice, message". However, if the department or enterprise to which the risk detection device belongs belongs is an asset management department, then the permitted alert methods include means such as business blocking and SMS notifications, and the risk alert permissions can be "block, message".
[0110] In practical use, the response strategies can be screened based on risk warning permissions. The target response strategy can be determined by screening response strategies based on risk warning permissions, removing the part of the response strategy whose processing flow does not comply with the risk warning permissions, and taking the remaining response strategies as the target response strategies.
[0111] Since there may still be multiple target response strategies, the same approach as when there are multiple strategies can also be adopted, which will not be elaborated here.
[0112] This embodiment extracts semantic data from collected user communication data to obtain communication semantic data. Based on this semantic data and user communication data, a user communication storyline is constructed. The user communication storyline is a data stream built from at least one communication data point corresponding to the same semantic meaning for the same user. Risk detection is then performed based on the communication semantic data and the user communication storyline to determine the presence of fraud risk. Because it is not based on single keyword filtering or single information analysis, but rather on constructing a user communication storyline based on corresponding communication semantic information, and then performing overall behavioral analysis and information statistics based on the user communication storyline, a more accurate analysis of the user's behavior before and after the connection is established. This ensures that even complex and ever-changing fraudulent behaviors can be detected.
[0113] refer to Figure 3 , Figure 3 This is a flowchart illustrating a second embodiment of a risk detection method according to the present invention.
[0114] Based on the first embodiment described above, step S30 of the risk detection method in this embodiment includes:
[0115] Step S301: Obtain the target user identifier corresponding to the user communication storyline.
[0116] It should be noted that obtaining the target user identifier corresponding to the user's communication storyline can mean obtaining the user identifier of the user corresponding to the user's communication storyline and using the obtained user identifier as the target user identifier. The user identifier can be unique identifier data used to identify a user, such as a mobile phone number.
[0117] Step S302: Locate the target user profile corresponding to the target user identifier.
[0118] It should be noted that finding the target user profile corresponding to the target user identifier can be done by searching for the corresponding user profile in a user profile database and using the found user profile as the target user profile. The target user profile is used to represent information such as the target user's habits and personal circumstances.
[0119] Step S303: Determine whether there is a risk of fraud based on the target user profile, the user communication storyline, and the communication semantic data.
[0120] In practical applications, determining the existence of fraud risk based on target user profiles, user communication stories, and communication semantic data can be achieved by marking user communication stories using communication semantic data, extracting behavioral features from the marked user communication stories to obtain user behavioral characteristics, comparing these characteristics with the target user profile, and determining whether the user's behavior in the communication story conforms to the user habits and personal circumstances recorded in the target user profile. If they do, it is determined that there is no fraud risk; otherwise, it is determined that there is a fraud risk.
[0121] In actual use, there may be multiple user communication storylines. Based on this, each user communication storyline can be processed separately. For example, user A may have four user communication storylines, numbered 1 to 4. In this case, steps S301 to S303 can be executed once for each of the four user communication storylines, numbered 1 to 4.
[0122] In practical implementation, since the deduction of whether it conforms to behavioral habits involves certain subjective factors, the probability of misjudgment may be relatively high. In order to minimize misjudgment, step S303 in this embodiment may include:
[0123] The target user profile is updated based on the user communication storyline and the communication semantic data to obtain an updated user profile;
[0124] The target user profile is compared with the updated user profile to determine the profile difference.
[0125] If the difference in the portrait is greater than a preset difference threshold, it is determined that there is a risk of fraud.
[0126] It should be noted that the profile difference score can be a quantitative score used to characterize the difference between two user profiles. The higher the profile difference score, the greater the difference between the two user profiles.
[0127] In practical use, the target user profile is updated based on the user communication storyline and communication semantic data. The updated user profile can be obtained by marking the user communication storyline with communication semantic data, extracting behavioral features from the marked user communication storyline, updating the user habits and personal information recorded in the target user profile based on the extracted behavioral features, and using the updated user profile as the updated user profile.
[0128] In practical applications, user profiles can be represented by multiple different categories of labels. The difference between the target user profile and the updated user profile is determined by comparing the two profiles. This can be achieved by calculating the label differences for each category in the target user profile and the updated user profile using a preset difference algorithm, and then averaging the weighted average of these differences to obtain the overall profile difference. The preset difference algorithm can be pre-set by the administrator of the risk detection equipment, such as using Euclidean distance or cosine similarity algorithms.
[0129] Understandably, if the difference in the user profile exceeds the preset difference threshold, it means that the user profile updated after the user's communication storyline is significantly different from the user's previous user profile. In this case, the user may have been misled, leading to abnormal behavior. Therefore, it can be determined that there is a risk of fraud.
[0130] In practical implementation, since comparing user profiles requires a large amount of computing resources, in order to reduce unnecessary resource consumption, the step of comparing the target user profile with the updated user profile to determine the profile difference degree, as described in this embodiment, may include:
[0131] The user communication storyline is examined to determine if any risky behavior exists;
[0132] If risky behavior is found, the target user profile is compared with the updated user profile to determine the degree of profile difference.
[0133] It should be noted that risky behavior can be any behavior that may affect a user's assets, such as spending, transferring money, or sending red envelopes; of course, behaviors that may indirectly affect a user's assets, such as sending personal information or related account information, can also be identified as risky behavior.
[0134] It is understandable that if a user exhibits risky behavior, it means that the user's assets may be affected. Therefore, if risky behavior is detected when the user's communication storyline is being monitored, the target user profile can be compared with the updated user profile to determine the profile difference. Based on the profile difference, it can be further determined whether the user has been misled, i.e., whether there is a risk of fraud.
[0135] If the user does not engage in any risky behavior, it means that the user's assets will not be affected. Therefore, subsequent steps can be discontinued to reduce unnecessary performance loss and ensure that the device's computing resources can be used for users whose assets may be affected.
[0136] This embodiment obtains the target user identifier corresponding to the user's communication storyline; finds the target user profile corresponding to the target user identifier; and determines whether there is a fraud risk based on the target user profile, the user's communication storyline, and the communication semantic data. Because it combines the user profile with the behavior extracted from the user's communication storyline to analyze the rationality of the behavior, it ensures that when a user is misled, leading to an unreasonable change in behavior, it can be detected.
[0137] refer to Figure 4 , Figure 4 This is a flowchart illustrating a third embodiment of a risk detection method according to the present invention.
[0138] Based on the first embodiment described above, step S30 of the risk detection method in this embodiment includes:
[0139] Step S301': Mark the user communication storyline according to the communication semantic data to obtain a semantically marked storyline.
[0140] It should be noted that, based on the semantic data of communication, the user communication storyline is marked to obtain the semantically marked storyline. This can be done by marking the content and type of each piece of user communication data that constructs the user communication storyline according to the correspondence between the semantic data of communication and the user communication data, and then using the marked user communication storyline as the semantically marked storyline.
[0141] Step S302': Obtain the target user corresponding to the user communication storyline.
[0142] It should be noted that the target user corresponding to the user's communication storyline can be the user corresponding to the user's communication storyline.
[0143] Step S303': Search the storyline repository for semantically tagged storylines for users other than the target user to obtain comparison storylines.
[0144] It should be noted that the storyline repository can be a database used to store constructed semantically marked storylines. Finding semantically marked storylines for users other than the target user in the storyline repository to obtain comparison storylines can be done by reading the semantically marked storylines for the remaining users besides the target user from the storyline repository.
[0145] In practical use, since the collected user communication data may be large and include communication data from multiple different users at different times and of different types, it may be processed in batches. The semantic tag storylines generated during the processing will be stored in the storyline repository for later horizontal comparison (i.e., comparing the semantic tag storylines of different users). Therefore, after obtaining the semantic tag storyline in the current processing flow, the semantic tag storylines corresponding to other users besides the target user can be searched in the storyline repository, and the found semantic tag storylines can be used as comparison storylines.
[0146] In practice, since horizontal comparison requires comparing a large amount of data, the overall execution will consume a lot of computing resources. Furthermore, horizontal comparison has high requirements for the comprehensiveness of the data. Therefore, step S303' of this embodiment can be executed after the last batch of semantic tag storylines is obtained in the current risk detection cycle.
[0147] For example, in the current risk detection cycle, the amount of user communication data is large. The user communication data is divided into 5 batches for detection according to different users. The semantic tag storylines generated by the first 1-4 batches can be stored in the storyline repository. When the 5th batch is processed and the semantic tag storylines generated by the 5th batch are obtained, step S303' of this embodiment can be executed to obtain all the semantic tag storylines generated in the current risk detection cycle.
[0148] Step S304': Construct broadcast event information based on the semantically marked storyline and the contrasting storyline.
[0149] It should be noted that constructing broadcast event information based on semantically marked storylines and contrasting storylines can be achieved by analyzing the semantically marked storylines and contrasting storylines to obtain the same or similar information received by different users, thereby generating broadcast event information.
[0150] In a specific implementation, in order to quickly construct broadcast event information, step S304' in this embodiment may include:
[0151] Semantic statistics are performed on the semantically marked storyline and the contrasting storyline to obtain the intersection semantic information;
[0152] The number of associated users corresponding to each intersection semantic information is determined by statistical analysis of the semantically labeled storylines.
[0153] Broadcast event information is constructed based on the intersection semantic information and the number of associated users.
[0154] It should be noted that the intersection semantic information can be the same or similar semantics possessed by at least two different semantic tag storylines, and the number of associated users can be the number of users whose corresponding semantic tag storylines include the intersection semantic information.
[0155] In practical applications, semantic statistics are performed on semantically labeled storylines and contrasting storylines to obtain intersecting semantic information. This can be achieved by statistically analyzing the frequency of each semantic information occurrence in both storylines, and identifying semantic information that appears more than or equal to a preset number in different storylines as intersecting semantic information. The preset number can be pre-set by the administrator of the risk detection equipment; for example, setting the preset number to 2.
[0156] In practical applications, the number of associated users corresponding to the semantic tag storylines of the intersection semantic information can be determined by adding the user identifiers corresponding to the semantic tag storylines of the intersection semantic information to a set, then deduplicating the set, counting the number of user identifiers in the deduplicated set, and using this number as the number of associated users corresponding to the intersection semantic information.
[0157] For example: Suppose the intersection semantic information is A, and there are 60 corresponding semantic tag storylines. After adding the user tags corresponding to these 60 semantic tag storylines to the set and removing duplicates, the number of remaining user tags is 30. At this time, the number of associated users corresponding to the intersection semantic information is 30, and the generated broadcast event information is "Ms: A, times: 30".
[0158] In order to facilitate subsequent responses, when determining the intersection semantic information, the sender of each piece of information can also be obtained. Semantic information with the same sender and appearing more than or equal to a preset number in different storylines is taken as the intersection semantic information.
[0159] Step S305': Determine whether there is a risk of fraud based on the broadcast event information.
[0160] In practical use, broadcast event information can determine how many different users received the same or similar information. By combining this information with an analysis of fraud strategies, it can be determined whether there is a risk of fraud.
[0161] In specific implementation, to minimize false alarms, step S305' in this embodiment may include:
[0162] Extract the number of associated users from the broadcast event information;
[0163] If the number of associated users is greater than or equal to a preset risk threshold, then a fraud risk is determined to exist.
[0164] It should be noted that the preset risk threshold can be set by the administrator of the risk detection device after analyzing the fraud strategy. For example, if the administrator of the risk detection device analyzes the fraud strategy and determines that if more than 30 different users receive information with the same or similar meaning, there is a risk of fraud, then the preset risk threshold can be set to 30.
[0165] Understandably, if the number of associated users is greater than or equal to the preset risk threshold, it means that a large number of different users have received information with the same or similar semantics as the intersection of semantic information contained in the broadcast event information. At this time, it may be that a fraud team is carrying out broadcasting or casting a wide net behavior. Therefore, it can be determined that there is a fraud risk.
[0166] If the number of associated users is less than the preset risk threshold, it means that although multiple users have received information with the same or similar semantics as the intersection of semantic information contained in the broadcast event information, the number of users involved is small. In this case, it may be that some users are sending mass notifications or other notification behaviors, and it can be determined that there is no risk of fraud.
[0167] To facilitate understanding, we will now combine... Figure 5 This explanation is provided, but it does not limit the scope of this solution. Figure 5 This is a schematic diagram of the risk detection and processing flow in this embodiment.
[0168] like Figure 5 As shown, the system can seamlessly connect to various communication networks and acquire text, voice, and image data in real time through real-time raw data acquisition technology from communication networks. Figure 5 The system processes various types of communication information (such as SMS, MMS, and phone calls). Then, it uses a multimodal large-scale model (also known as a deep semantic understanding model) for semantic extraction, extracting key features (i.e., classification vectors representing content association categories). Simultaneously, it clusters user communication information based on corresponding semantic relevance to form storylines (i.e., the aforementioned user communication storylines). These storylines are then input into a decision tree, where they are combined with user profiles and statistical rules for horizontal comparison (i.e., constructing broadcast event information and determining the presence of fraud risk based on broadcast event information) and / or vertical in-depth analysis (i.e., determining the presence of fraud risk based on user profiles). This determines whether fraud risk exists. Finally, based on the risk assessment results (i.e., whether fraud risk exists), and combined with a response strategy library (i.e.... Figure 5 The knowledge graph in the database, also known as the knowledge graph strategy base, generates corresponding processing decisions (i.e., ...). Figure 5 (The processing decisions shown include release, reminder, warning, delay, and blocking).
[0169] This embodiment obtains semantically tagged storylines by marking user communication storylines based on the communication semantic data; it then identifies the target user corresponding to each user communication storyline; it searches the storyline repository for semantically tagged storylines corresponding to other users besides the target user to obtain comparison storylines; it constructs broadcast event information based on the semantically tagged storylines and the comparison storylines; and it determines whether there is a risk of fraud based on the broadcast event information. Because it involves a horizontal comparison between the currently constructed semantically tagged storylines and the semantically tagged storylines of other users, and uses broadcast event information to represent situations where the same or similar information is received by different users, it ensures that the broadcast or wide-net fraudulent activities of fraud teams can be identified.
[0170] Furthermore, embodiments of the present invention also propose a storage medium storing a risk detection program, which, when executed by a processor, implements the steps of the risk detection method described above.
[0171] Reference Figure 6 , Figure 6 This is a structural block diagram of the first embodiment of the risk detection device of the present invention.
[0172] like Figure 6 As shown, the risk detection device proposed in this embodiment of the invention includes:
[0173] Extraction module 10 is used to perform semantic extraction on the collected user communication data to obtain communication semantic data;
[0174] Construction module 20 is used to construct a user communication storyline based on the communication semantic data and the user communication data, wherein the user communication storyline is a data stream constructed based on at least one communication data corresponding to the same semantic data of the same user;
[0175] The detection module 30 is used to perform risk detection based on the communication semantic data and the user communication storyline to determine whether there is a risk of fraud.
[0176] This embodiment extracts semantic data from collected user communication data to obtain communication semantic data. Based on this semantic data and user communication data, a user communication storyline is constructed. The user communication storyline is a data stream built from at least one communication data point corresponding to the same semantic meaning for the same user. Risk detection is then performed based on the communication semantic data and the user communication storyline to determine the presence of fraud risk. Because it is not based on single keyword filtering or single information analysis, but rather on constructing a user communication storyline based on corresponding communication semantic information, and then performing overall behavioral analysis and information statistics based on the user communication storyline, a more accurate analysis of the user's behavior before and after the connection is established. This ensures that even complex and ever-changing fraudulent behaviors can be detected.
[0177] Furthermore, the user communication data is multimodal data;
[0178] The extraction module 10 is also used to perform semantic extraction on the collected user communication data through a multimodal large model to obtain communication semantic data. The multimodal large model is a pre-trained model that performs semantic extraction on multiple different types of data.
[0179] Furthermore, the detection module 30 is also used to obtain the target user identifier corresponding to the user communication storyline; find the target user profile corresponding to the target user identifier; and determine whether there is a fraud risk based on the target user profile, the user communication storyline, and the communication semantic data.
[0180] Furthermore, the detection module 30 is also used to update the target user profile based on the user communication storyline and the communication semantic data to obtain an updated user profile; compare the target user profile with the updated user profile to determine the profile difference degree; if the profile difference degree is greater than a preset difference threshold, it is determined that there is a fraud risk.
[0181] Furthermore, the detection module 30 is also used to detect the user communication storyline to determine whether there is any risky behavior; if there is risky behavior, the target user profile is compared with the updated user profile to determine the profile difference.
[0182] Furthermore, the detection module 30 is also used to mark the user communication storyline according to the communication semantic data to obtain a semantically marked storyline; obtain the target user corresponding to the user communication storyline; search for semantically marked storylines corresponding to other users besides the target user in the storyline repository to obtain a comparison storyline; construct broadcast event information based on the semantically marked storyline and the comparison storyline; and determine whether there is a fraud risk based on the broadcast event information.
[0183] Furthermore, the detection module 30 is also used to perform semantic statistics on the semantically marked storyline and the comparison storyline to obtain intersection semantic information, wherein the intersection semantic information is the same or similar semantics possessed by at least two different semantically marked storylines; to perform user attribution statistics on the semantically marked storylines corresponding to the intersection semantic information to determine the number of associated users corresponding to each intersection semantic information; and to construct broadcast event information based on the intersection semantic information and the number of associated users.
[0184] Furthermore, the detection module 30 is also used to extract the number of associated users from the broadcast event information; if the number of associated users is greater than or equal to a preset risk threshold, it is determined that there is a fraud risk.
[0185] Furthermore, the detection module 30 is also used to obtain the risk type when there is a risk of fraud; search for the corresponding response strategy in the response strategy library; and issue a risk warning based on the response strategy.
[0186] Furthermore, the detection module 30 is also used to acquire risk warning permissions, which represent the warning methods currently allowed to be used for issuing warnings; to filter the response strategies according to the risk warning permissions and determine the target response strategy; and to issue risk warnings according to the target response strategy.
[0187] Furthermore, the construction module 20 is also used to cluster the communication semantic data to obtain at least one cluster; and sort the user communication data corresponding to each cluster in ascending order of the corresponding communication time to obtain the user communication storyline.
[0188] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0189] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0190] In addition, for technical details not described in detail in this embodiment, please refer to the risk detection method provided in any embodiment of the present invention, which will not be repeated here.
[0191] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0192] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0194] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
[0195] This invention discloses A1, a risk detection method, which includes the following steps:
[0196] Semantic extraction is performed on the collected user communication data to obtain communication semantic data;
[0197] A user communication storyline is constructed based on the communication semantic data and the user communication data. The user communication storyline is a data stream constructed based on at least one communication data corresponding to the same user and the same semantic data.
[0198] Risk detection is performed based on the communication semantic data and the user's communication storyline to determine whether there is a risk of fraud.
[0199] A2. In the risk detection method described in A1, the user communication data is multimodal data;
[0200] The step of extracting semantic data from the collected user communication data to obtain communication semantic data includes:
[0201] Semantic extraction is performed on the collected user communication data using a multimodal large model to obtain communication semantic data. The multimodal large model is a pre-trained model that performs semantic extraction on various types of data.
[0202] A3. The risk detection method as described in A1, wherein the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk includes:
[0203] Obtain the target user identifier corresponding to the user communication storyline;
[0204] Find the target user profile corresponding to the target user identifier;
[0205] The existence of fraud risk is determined based on the target user profile, the user communication storyline, and the communication semantic data.
[0206] A4. The risk detection method as described in A3, wherein the step of determining whether there is a fraud risk based on the target user profile, the user communication storyline, and the communication semantic data includes:
[0207] The target user profile is updated based on the user communication storyline and the communication semantic data to obtain an updated user profile;
[0208] The target user profile is compared with the updated user profile to determine the profile difference.
[0209] If the difference in the portrait is greater than a preset difference threshold, it is determined that there is a risk of fraud.
[0210] A5. The risk detection method as described in A4, wherein the step of comparing the target user profile with the updated user profile to determine the profile difference includes:
[0211] The user communication storyline is examined to determine if any risky behavior exists;
[0212] If risky behavior is found, the target user profile is compared with the updated user profile to determine the degree of profile difference.
[0213] A6. The risk detection method as described in A1, wherein the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk includes:
[0214] The user communication storyline is marked based on the communication semantic data to obtain a semantically marked storyline;
[0215] Obtain the target user corresponding to the user communication storyline;
[0216] Search the storyline repository for semantically tagged storylines for users other than the target user to obtain comparison storylines;
[0217] Construct broadcast event information based on the semantically labeled storyline and the contrasting storyline;
[0218] The presence of fraud risk is determined based on the broadcast event information.
[0219] A7. The risk detection method as described in A6, wherein the step of constructing broadcast event information based on the semantically marked storyline and the contrasting storyline includes:
[0220] Semantic statistics are performed on the semantically marked storyline and the contrasting storyline to obtain the intersection semantic information, wherein the intersection semantic information consists of the same or similar semantics possessed by at least two different semantically marked storylines;
[0221] The number of associated users corresponding to each intersection semantic information is determined by statistical analysis of the semantically labeled storylines.
[0222] Broadcast event information is constructed based on the intersection semantic information and the number of associated users.
[0223] A8. The risk detection method as described in A6, wherein the step of determining whether there is a risk of fraud based on the broadcast event information includes:
[0224] Extract the number of associated users from the broadcast event information;
[0225] If the number of associated users is greater than or equal to a preset risk threshold, then a fraud risk is determined to exist.
[0226] A9. The risk detection method as described in A1, after the step of performing risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk, further includes:
[0227] When there is a risk of fraud, determine the type of risk.
[0228] Search the response strategy library for the corresponding risk type;
[0229] Risk warnings will be issued based on the aforementioned response strategies.
[0230] A10. The risk detection method as described in A9, wherein the step of issuing a risk warning based on the response strategy includes:
[0231] Obtain risk alert permission, which indicates the alert methods currently allowed to be used;
[0232] The response strategies are filtered based on the risk alert permissions to determine the target response strategy;
[0233] Risk warnings will be issued based on the aforementioned target response strategies.
[0234] A11. The risk detection method as described in any one of A1-A10, wherein the step of constructing a user communication storyline based on the communication semantic data and the user communication data includes:
[0235] The communication semantic data is clustered to obtain at least one cluster.
[0236] The user communication data corresponding to each cluster is sorted in ascending order of the corresponding communication time to obtain the user communication storyline.
[0237] The present invention also discloses B12, a risk detection device, the risk detection device comprising the following modules:
[0238] The extraction module is used to extract semantics from the collected user communication data to obtain communication semantic data.
[0239] The construction module is used to construct a user communication storyline based on the communication semantic data and the user communication data. The user communication storyline is a data stream constructed based on at least one communication data corresponding to the same semantic data of the same user.
[0240] The detection module is used to perform risk detection based on the communication semantic data and the user communication storyline to determine whether there is a risk of fraud.
[0241] B13. The risk detection device as described in B12, wherein the user communication data is multimodal data;
[0242] The extraction module is also used to perform semantic extraction on the collected user communication data through a multimodal large model to obtain communication semantic data. The multimodal large model is a pre-trained model that performs semantic extraction on multiple different types of data.
[0243] B14. The risk detection device as described in B12, wherein the detection module is further configured to obtain the target user identifier corresponding to the user communication storyline; find the target user profile corresponding to the target user identifier; and determine whether there is a fraud risk based on the target user profile, the user communication storyline, and the communication semantic data.
[0244] B15. The risk detection device as described in B14, wherein the detection module is further configured to update the target user profile based on the user communication storyline and the communication semantic data to obtain an updated user profile; compare the target user profile with the updated user profile to determine the profile difference degree; if the profile difference degree is greater than a preset difference threshold, it is determined that there is a fraud risk.
[0245] B16. The risk detection device as described in B14, wherein the detection module is further configured to detect the user communication storyline to determine whether there is any risky behavior; if there is any risky behavior, the target user profile is compared with the updated user profile to determine the profile difference.
[0246] B17. The risk detection device as described in B12, wherein the detection module is further configured to: mark the user communication storyline according to the communication semantic data to obtain a semantically marked storyline; obtain the target user corresponding to the user communication storyline; search in the storyline repository for semantically marked storylines corresponding to other users besides the target user to obtain a comparison storyline; construct broadcast event information based on the semantically marked storyline and the comparison storyline; and determine whether there is a fraud risk based on the broadcast event information.
[0247] B18. The risk detection device as described in B17, wherein the detection module is further configured to perform semantic statistics on the semantically marked storyline and the comparison storyline to obtain intersection semantic information, wherein the intersection semantic information consists of the same or similar semantics possessed by at least two different semantically marked storylines; perform user attribution statistics on the semantically marked storylines corresponding to the intersection semantic information to determine the number of associated users corresponding to each intersection semantic information; and construct broadcast event information based on the intersection semantic information and the number of associated users.
[0248] The present invention also discloses C19, a risk detection device, the risk detection device comprising: a processor, a memory, and a risk detection program stored in the memory and executable on the processor, wherein the risk detection program, when executed by the processor, implements the steps of the risk detection method as described above.
[0249] The present invention also discloses D20, a computer-readable storage medium storing a risk detection program, wherein the risk detection program, when executed, implements the steps of the risk detection method described above.
Claims
1. A risk detection method, characterized by, The risk detection method comprises the following steps: performing semantic extraction on the collected user communication data to obtain communication semantic data, the communication semantic data corresponding one-to-one to the user communication data, the communication semantic data comprising semantic content and a classification vector for representing a content association category; constructing a user communication storyline based on the communication semantic data and the user communication data, the user communication storyline being a data flow constructed according to at least one communication data corresponding to the same user and the same semantic, and one user corresponding to at least one user communication storyline; performing risk detection according to the communication semantic data and the user communication storyline to determine whether there is a fraud risk, the risk detection being based on user communication storyline to find a corresponding user portrait to confirm whether there is a fraud risk, and / or based on user communication storyline to find a comparison storyline to construct broadcast event information to determine whether there is a fraud risk; the step of constructing a user communication storyline based on the communication semantic data and the user communication data comprises: grouping communication semantic data with the same classification vector and similar or identical semantic content into the same cluster to obtain at least one clustering cluster; sorting the user communication data corresponding to each clustering cluster in ascending order of corresponding communication time to obtain a user communication storyline.
2. The risk detection method of claim 1, wherein, The user communication data is multi-modal data. The step of performing semantic extraction on the collected user communication data to obtain communication semantic data comprises: performing semantic extraction on the collected user communication data by a multi-modal large model to obtain communication semantic data, the multi-modal large model being a pre-trained model for performing semantic extraction on multiple different types of data.
3. The risk detection method of claim 1, wherein, The step of performing risk detection according to the communication semantic data and the user communication storyline to determine whether there is a fraud risk comprises: obtaining a target user identifier corresponding to the user communication storyline; finding a target user portrait corresponding to the target user identifier; determining whether there is a fraud risk according to the target user portrait, the user communication storyline and the communication semantic data.
4. The risk detection method of claim 3, wherein, The step of determining whether there is a fraud risk according to the target user portrait, the user communication storyline and the communication semantic data comprises: updating the target user portrait according to the user communication storyline and the communication semantic data to obtain an updated user portrait; comparing the target user portrait with the updated user portrait to determine a portrait difference degree; if the portrait difference degree is greater than a preset difference threshold, it is determined that there is a fraud risk.
5. The risk detection method of claim 4, wherein, The step of comparing the target user portrait with the updated user portrait to determine a portrait difference degree comprises: detecting the user communication storyline to determine whether there is a risk behavior; if there is a risk behavior, comparing the target user portrait with the updated user portrait to determine a portrait difference degree.
6. The risk detection method of claim 1, wherein, The step of performing risk detection according to the communication semantic data and the user communication storyline to determine whether there is a fraud risk comprises: labeling the user communication storyline according to the communication semantic data to obtain a semantic labeled storyline; Obtain the target user corresponding to the user communication storyline; Search the storyline repository for semantically tagged storylines for users other than the target user to obtain comparison storylines; Construct broadcast event information based on the semantically labeled storyline and the contrasting storyline; The presence of fraud risk is determined based on the broadcast event information.
7. The risk detection method of claim 6, wherein, The step of constructing broadcast event information based on the semantically marked storyline and the contrasting storyline includes: Semantic statistics are performed on the semantically marked storyline and the contrasting storyline to obtain the intersection semantic information, wherein the intersection semantic information consists of the same or similar semantics possessed by at least two different semantically marked storylines; The number of associated users corresponding to each intersection semantic information is determined by statistical analysis of the semantically labeled storylines. Broadcast event information is constructed based on the intersection semantic information and the number of associated users.
8. A risk detection apparatus, characterized by, The risk detection device includes the following modules: The extraction module is used to perform semantic extraction on the collected user communication data to obtain communication semantic data. The communication semantic data corresponds one-to-one with the user communication data. The communication semantic data includes semantic content and a classification vector used to represent the content association category. The construction module is used to construct a user communication storyline based on the communication semantic data and the user communication data. The user communication storyline is a data stream constructed based on at least one communication data corresponding to the same semantic data of the same user. One user corresponds to at least one user communication storyline. The detection module is used to perform risk detection based on the communication semantic data and the user communication storyline to determine whether there is a fraud risk. The risk detection is to find the corresponding user profile based on the user communication storyline to confirm whether there is a fraud risk, and / or to find and compare the storylines based on the user communication storyline to construct broadcast event information to determine whether there is a fraud risk. The construction module is also used to cluster communication semantic data with the same classification vector and similar or identical semantic content into the same cluster to obtain at least one cluster. The user communication data corresponding to each cluster is sorted in ascending order of the corresponding communication time to obtain the user communication storyline.
9. A risk detection apparatus, characterized by, The risk detection device includes: a processor, a memory, and a risk detection program stored in the memory and executable on the processor. When the risk detection program is executed by the processor, it implements the steps of the risk detection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a risk detection program, which, when executed, implements the steps of the risk detection method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method and system for real-time detection of communication fraud base on suspicious behavior recognition
CN107222865A
The invention relates to a transaction security control method and a system based on subject portrait
CN109509093A
Massive fraud short message detection method and device, server and storage medium
CN111083705A