Fraud warning method, device, electronic device and storage medium

By segmenting and classifying conversation data, matching it with preset speech templates, and identifying the fraud process, the problems of timeliness and comprehensiveness of online fraud warning effects in existing technologies are solved, and efficient and reliable fraud warning is achieved.

CN117271723BActive Publication Date: 2025-09-26ANHUI IFLYREC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311173507.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-11
Publication Date
2025-09-26
Estimated Expiration
2043-09-11

AI Technical Summary

Technical Problem

In the existing online fraud warning methods, the warning effect is less timely and comprehensive, and it is difficult to accurately and quickly identify fraudulent behavior, especially when criminals use changing phone numbers or virtual phones, the existing methods are difficult to effectively prevent.

Method used

By segmenting the conversation data to be warned, at least one round of conversation data is obtained, and each round of conversation data is classified. Based on the arrangement order and speech type of the conversation data in the conversation data, it is matched with the preset speech template, and a preset speech template is constructed to identify the fraud process and issue a warning.

Benefits of technology

It ensures timely warning while improving the comprehensiveness and reliability of fraud warnings, can quickly identify and prevent potential fraudulent behaviors, and improves the efficiency and accuracy of warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117271723B_ABST
    Figure CN117271723B_ABST
Patent Text Reader

Abstract

The present invention provides a fraud early warning method, device, electronic device, and storage medium, wherein the method comprises: segmenting conversation data to be warned to obtain at least one round of conversation data; classifying each round of conversation data to obtain the speech type to which each round of conversation data belongs; matching the order of the at least one round of conversation data in the conversation data and the speech type to which each round of conversation data belongs with a preset speech template, and issuing a fraud early warning based on the matching result; the preset speech template is constructed based on the fraudulent conversation data. The method, device, electronic device, and storage medium provided by the present invention can improve the comprehensiveness and reliability of fraud early warnings while ensuring timely warnings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a fraud warning method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of information technology, there are more and more online frauds through mobile phones or the Internet. To combat this type of online fraud, the current method is mainly to identify the incoming call number or account. When the incoming call number or account is identified as a fraudulent call or account, a reminder is issued through marking, text message or phone call.

[0003] Although the above methods can prevent online fraud to a certain extent, they usually mark the phone number or account as a fraudulent phone number or account only after many people have been deceived. They cannot prevent fraud in the early stages and are not timely. In addition, criminals can also change phone numbers, account numbers, or use virtual phones to commit fraud. Using the above methods, it is difficult to accurately and quickly identify such fraudulent behaviors, resulting in poor early warning effects. Summary of the Invention

[0004] The present invention provides a fraud early warning method, device, electronic device and storage medium, which are used to solve the problems of poor early warning effect and timeliness of fraud early warning methods in the prior art.

[0005] The present invention provides a fraud early warning method, comprising:

[0006] Split the conversation data to be warned to obtain at least one round of conversation data;

[0007] Classify each round of dialogue data to obtain the speech type to which each round of dialogue data belongs;

[0008] Based on the arrangement order of at least one round of conversation data in the conversation data, and the type of speech to which each round of conversation data belongs, matching is performed with a preset speech template, and a fraud warning is issued based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0009] According to a fraud warning method provided by the present invention, segmenting the conversation data to be warned to obtain at least one round of conversation data includes:

[0010] determining a key sentence indicating an affirmative response from each sentence in the conversation data;

[0011] The conversation data is segmented using the key sentence as the ending sentence of the conversation data to obtain at least one round of conversation data.

[0012] According to a fraud early warning method provided by the present invention, determining a key sentence indicating an affirmative response from each sentence in the conversation data includes:

[0013] Determining candidate sentences in which the speaker is the listener from each sentence in the conversation data;

[0014] determining a key phrase indicating an affirmative response from the candidate sentences;

[0015] Determining a length ratio of the key phrase in the candidate sentence based on the phrase length of the key phrase and the sentence length of the candidate sentence in which the key phrase is located;

[0016] Based on the length ratio, determine whether the candidate sentence is the key sentence.

[0017] According to a fraud early warning method provided by the present invention, the process of classifying each round of conversation data to obtain the speech type of each round of conversation data includes:

[0018] Extract keywords from each round of conversation data;

[0019] Based on the keywords in each round of dialogue data and the preset keywords under each preset dialogue type, the speech type to which each round of dialogue data belongs is determined.

[0020] According to a fraud early warning method provided by the present invention, the matching of the arrangement order of at least one round of dialogue data in the conversation data and the speech type of each round of dialogue data with a preset speech template includes:

[0021] Constructing a speech evolution trajectory of the conversation data based on the arrangement order of the at least one round of conversation data in the conversation data and the speech type of each round of conversation data;

[0022] The trajectory of the speech evolution is matched with the trajectory of the fraud speech evolution corresponding to the preset speech template.

[0023] According to a fraud early warning method provided by the present invention, the steps of constructing the preset speech template include:

[0024] Segmenting each fraudulent conversation data to obtain at least one round of sample conversation data under each fraudulent conversation data;

[0025] Classify each round of sample conversation data to obtain the speech type to which each round of sample conversation data belongs;

[0026] The preset speech template is constructed based on the arrangement order of at least one round of sample conversation data in each fraud conversation data, and the speech type of each round of sample conversation data.

[0027] According to a fraud early warning method provided by the present invention, the preset speech template is constructed based on the arrangement order of at least one round of sample conversation data in each fraud session data and the speech type of each round of sample conversation data, including:

[0028] Determining the evolution trajectory of the fraudulent speech corresponding to each fraudulent conversation data based on the arrangement order of the at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data;

[0029] The number of times each fraudulent speech evolution track is applied in each fraudulent session data is counted, and the preset speech template is determined from each fraudulent speech evolution track based on the number of applications.

[0030] The present invention also provides a fraud early warning device, comprising:

[0031] A segmentation unit, configured to segment the conversation data to be warned to obtain at least one round of conversation data;

[0032] The classification unit is used to classify each round of dialogue data and obtain the speech type to which each round of dialogue data belongs;

[0033] An early warning unit is used to match the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs with a preset speech template, and to issue a fraud early warning based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the fraud warning method described above is implemented.

[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the fraud warning methods described above.

[0036] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described fraud warning methods.

[0037] The fraud warning method, device, electronic device and storage medium provided by the present invention divide the conversation data to be warned to obtain at least one round of conversation data, and classify the conversations of each round of conversation data to obtain the type of speech to which each round of conversation data belongs. Therefore, based on the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs, they can be matched with a preset speech template, and a fraud warning can be issued based on the matching result. This realizes fraud warning based on fraud process comparison, without relying on identification and judgment of incoming call numbers or accounts, and can improve the comprehensiveness and reliability of fraud warning while ensuring timely warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 It is a flowchart of the fraud early warning method provided by the present invention;

[0040] Figure 2 1 is a flow chart of step 110 in the fraud warning method provided by the present invention;

[0041] Figure 3 1 is a flow chart of step 111 in the fraud warning method provided by the present invention;

[0042] Figure 4 1 is a flow chart of step 120 in the fraud warning method provided by the present invention;

[0043] Figure 5 This is a flow chart of the steps for constructing a preset speech template provided by the present invention;

[0044] Figure 6 It is a structural diagram of the fraud early warning device provided by the present invention;

[0045] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0047] With the rapid development of information technology, online fraud via mobile phones or the internet is becoming increasingly common. Current methods for combating this type of online fraud primarily rely on identifying incoming caller numbers or accounts. When a caller number or account is identified as fraudulent, a warning message, text message, or phone call is sent. While this method can prevent online fraud to a certain extent, it primarily targets a number or account after many people have fallen for it, marking it as fraudulent. This prevents fraud in its early stages and is therefore ineffective. Furthermore, criminals can also employ different phone numbers and accounts, or use virtual phones to commit fraud. Existing methods struggle to accurately and quickly identify this type of fraud, resulting in limited early warning effectiveness.

[0048] In addition, the existing technology also has a solution that uses keyword recognition during a call to determine whether there is suspicion of fraud. However, considering that fraud tactics are changeable, criminals may consciously avoid some words with obvious tendencies. Fraud warnings based on keyword recognition are limited by the fact that fraud keywords cannot be enumerated, resulting in a lack of comprehensiveness and poor reliability of fraud warnings.

[0049] In summary, how to improve the comprehensiveness and reliability of fraud warning while ensuring timely warning is still an urgent problem to be solved. To this end, the embodiment of the present invention provides a fraud warning method to overcome the above-mentioned shortcomings.

[0050] Figure 1 It is a flow chart of the fraud early warning method provided by the present invention, such as Figure 1 As shown, the method includes:

[0051] Step 110, dividing the conversation data to be warned to obtain at least one round of conversation data;

[0052] Specifically, conversation data subject to early warning refers to conversations or communication records that require analysis to determine whether fraudulent activity exists. These conversations can be between users or between users and automated systems (such as robots or virtual assistants). These conversations can include audio, video, and text.

[0053] During a real-time conversation or online meeting, the conversation data to be warned can be collected and obtained through text recording, voice recording, video recording, data stream capture, etc. After obtaining the conversation data, the conversation data can be segmented to divide the conversation data into at least one round of conversation data.

[0054] It is understandable that before segmenting the conversation data, the conversation voice in the conversation data can be converted into text through voice recognition, and the conversation voice in the conversation video in the conversation data can be extracted from it and converted into text through voice recognition.

[0055] After obtaining text-based conversation data, the conversation data can be segmented into at least one round of conversation data by identifying different rounds or interactions within the conversation data. Conversation data herein refers to the interaction information between the user and other conversation participants (or the system), including questions, answers, and instructions. The at least one round of conversation data refers to conversation data that includes one or more rounds or interactions.

[0056] For example, by analyzing the text content, semantics, and context of conversation data, different questions and answers can be identified, thereby segmenting the conversation data into multiple turns. Natural language processing and text analysis techniques can be used to segment the conversation data. For example, the switching points between different turns can be determined based on the sentence structure, keyword occurrence, semantic similarity, etc. in the conversation data, and at least one round of conversation data can be segmented based on the switching points.

[0057] Step 120: Classify each round of conversation data to obtain the speech type to which each round of conversation data belongs;

[0058] Specifically, after segmenting and obtaining at least one round of conversation data, each round of conversation data can be classified to determine the speech type to which each round of conversation data belongs. Here, speech type refers to the specific method or pattern used to commit fraud or deception, representing the speech and strategies used by fraudsters when communicating or interacting with victims. For example, speech types may include gaining trust, obtaining information, and guiding actions. Information acquisition can be further categorized based on content, such as obtaining names, obtaining mobile phone numbers, obtaining bank card information, obtaining deposit balances and loan limits, etc. Guidance actions can be further categorized based on the content of the operation, such as guiding the user to open an online banking app, guiding the user to make an online loan, or guiding the user to make a transfer.

[0059] In order to classify each round of dialogue data, a classification model can be established through machine learning or deep learning technologies, and each round of dialogue data can be input into the established classification model respectively. The classification model can be used to extract features of each round of dialogue data. For example, in the dialogue data, features can be extracted based on the content, semantics and context of the text. These features may include keywords, phrases, sentence structure, emotional color, etc. Classification prediction is performed based on the extracted features, thereby determining the type of speech to which each round of dialogue data belongs.

[0060] Step 130: Based on the arrangement order of at least one round of conversation data in the conversation data and the speech type of each round of conversation data, matching is performed with a preset speech template, and a fraud warning is issued based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0061] It should be noted that fraudulent behavior typically follows a specific fraud process, which refers to a series of steps or actions taken by fraudsters to achieve their deceptive purposes. It describes the evolution of fraudulent activities and the strategies and means employed during their progress. Therefore, to address the issues of poor effectiveness and timeliness of fraud warnings in the prior art, the embodiments of the present invention match the order of at least one round of conversation data in the session data, as well as the speech type of each round of conversation data, with a preset speech template. This achieves fraud warning based on fraud process comparison, which not only ensures the timeliness of the warning but also improves the comprehensiveness and reliability of the fraud warning.

[0062] Specifically, the arrangement order of each round of conversation data in the conversation data is usually arranged in chronological order, that is, arranged in the order in which the conversations occurred. Therefore, by matching the arrangement order of each round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs with the preset speech template, it is possible to compare the conversation data with the preset speech template based on the fraud process, so as to determine whether there is a fraud risk.

[0063] Here, preset script templates refer to multiple common and general fraud script templates pre-built based on fraudulent conversation data. These templates include various preset script types and the order in which they are implemented. For example, the preset script types may include gaining trust, obtaining information, and guiding actions. Gaining trust means that the fraudster establishes a trusting relationship with the victim by posing as a trustworthy individual, institution, or organization. Once the victim's trust is gained, the fraudster uses various methods to obtain the victim's personal information, account information, passwords, and other sensitive data. After obtaining the victim's information, the fraudster will guide the victim to take specific actions, such as transferring money or downloading malware. Therefore, gaining trust, obtaining information, and guiding actions constitute a typical fraud process.

[0064] After obtaining the speech type associated with each round of conversation data, the order in which the speech types are advanced within each round of conversation data can be determined based on the order in which the conversation data is arranged within the conversation data. For example, after segmenting the conversation data to obtain the first, second, and third rounds of conversation data, each round of conversation data is classified into conversation types. The speech type associated with the first round of conversation data is determined to be gaining trust, the speech type associated with the second round of conversation data is obtaining information, and the speech type associated with the third round of conversation data is guiding operations. Based on the order in which the conversation data is arranged, the order in which the speech types are advanced can be determined to be: gaining trust → obtaining information → guiding operations.

[0065] The speech types and advancement order of each round of dialogue data in the session data are matched with the preset speech types and advancement order in the preset speech template. When all speech types and advancement orders are matched successfully, it indicates that the session data has applied the fraud speech template and conforms to the rules of the fraud process. Therefore, it can be determined that the session data has a high fraud risk, and corresponding measures can be taken to issue fraud warnings to users. For example, risk warnings can be sent to users, specific operations or transactions can be blocked, etc., to prevent users from further fraud.

[0066] The method provided by the embodiment of the present invention divides the conversation data to be warned to obtain at least one round of conversation data, and classifies the conversations of each round of conversation data to obtain the type of speech to which each round of conversation data belongs. Therefore, based on the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs, it can be matched with a preset speech template, and a fraud warning can be issued based on the matching result. This realizes fraud warning based on fraud process comparison, without relying on the identification and judgment of the incoming call number or account number, and can improve the comprehensiveness and reliability of fraud warning while ensuring timely warning.

[0067] In addition, by using pre-built preset speech templates, conversation data can be quickly and automatically matched with the templates, which can greatly improve the efficiency of fraud identification. By implementing fraud warning based on fraud process comparison, the ability to identify and prevent potential fraud can be improved, thereby further improving the effectiveness of fraud warning.

[0068] Based on the above embodiments, Figure 2 is a flow chart of step 110 in the fraud warning method provided by the present invention, such as Figure 2 As shown, step 110 specifically includes:

[0069] Step 111, determining a key sentence indicating an affirmative response from each sentence in the conversation data;

[0070] It should be noted that, considering that fraudulent conversations often consist of a series of interactive dialogue turns, the key sentences that express affirmative responses in each sentence in the conversation data usually mark different stages or turning points of the conversation. Therefore, the conversation data can be segmented based on these key sentences, which helps to identify different stages of the conversation and thus better identify and prevent fraudulent behavior.

[0071] Specifically, each sentence in conversation data refers to the various sentences exchanged or communicated between the two parties in a conversation. Conversation data can include at least one round of conversation data, and each round of conversation data can be composed of one or more sentences. The above-mentioned key sentences indicating affirmative responses refer to sentences in which the recipient gives a positive and affirmative answer or response to the speaker's questions, requests, or statements in the conversation. These sentences indicate that the recipient agrees with or confirms the speaker's views, requests, or actions, and imply the progress or turning point of the conversation. For example, key sentences indicating affirmative responses can be "My bank account number is ***", "Okay, I'll open the ** bank app", "Okay, I'll do this", "Okay, I can provide this information", "No problem", etc.

[0072] Here, to identify key sentences indicating affirmative responses from each sentence in the conversation data, contextual analysis can be used. By analyzing and understanding the context of the conversation data, particularly the relationship between the preceding and following sentences, it is possible to identify whether a sentence indicates an affirmative response. For example, if the preceding sentence is a question or request, and the following sentence gives a clear affirmative answer to the question or request, then the following sentence can be determined to be a key sentence indicating an affirmative response. Alternatively, this can be achieved by searching for key phrases that are often associated with affirmative responses, such as "yes," "okay," "okay," and "understood." If these key phrases appear in a sentence, then the sentence is likely a key sentence indicating an affirmative response.

[0073] Step 112 : Using the key sentence as the ending sentence of the conversation data, the conversation data is segmented to obtain at least one round of conversation data.

[0074] Specifically, after determining each key sentence in the conversation data, each sentence in the conversation data can be traversed one by one. When a key sentence is encountered, it can be regarded as the end of a round of conversation data. That is, the key sentence and the sentence before it can be regarded as a round of conversation data. The subsequent sentences can be traversed to find the next key sentence for segmentation until the entire conversation data is traversed and each round of conversation data in the conversation data is segmented, thereby obtaining at least one round of conversation data. It is understood that each round of conversation data includes both the fraudster's sentences and the victim's sentences.

[0075] The method provided in an embodiment of the present invention divides conversation data into multiple rounds of conversation data based on key sentences indicating affirmative responses, which helps to identify different conversation stages and better analyze and understand the conversation process so as to subsequently compare it with the fraud process, thereby better identifying and preventing fraudulent behavior.

[0076] Based on the above embodiments, Figure 3 1 is a flow chart of step 111 in the fraud warning method provided by the present invention, as shown in FIG. Figure 3 As shown, step 111 specifically includes:

[0077] Step 1111, determining candidate sentences in which the speaker is the listener from each sentence in the conversation data;

[0078] It's important to note that fraudulent conversations often consist of a series of interactive rounds, in which scammers use specific rhetoric to guide victims into certain responses. When victims answer the scammers' key questions or express a positive attitude, it typically signifies the conversation has entered a new phase or reached a turning point. Therefore, embodiments of the present invention filter the recipient's statements from conversation data and identify key statements that represent positive responses for each statement, thereby enabling more accurate identification of key statements.

[0079] Specifically, the speaker-receiver refers to the person who answers a call or responds to a message during a phone call or other communication method. In the context of fraud, this typically refers to the victim or target of the fraud. Conversational data is typically presented as a two-party conversation. When converting audio or video conversations into text, roles are tagged. For example, the two parties in the conversation might be labeled as Speaker 1, Speaker 2, and so on. Based on these tags, candidate sentences in which the speaker is the recipient are identified. These candidate sentences are the sentences corresponding to the recipient in the conversational data.

[0080] Step 1112, determining a key phrase indicating an affirmative response from the candidate sentences;

[0081] Specifically, after determining the candidate sentences of the answering party, a keyword matching method can be used to determine the key phrases expressing the affirmative response from the candidate sentences. Here, the key phrase refers to the phrase or word group expressing the affirmative response in the candidate sentences.

[0082] Step 1113: determining a length ratio of the key phrase in the candidate sentence based on the length of the key phrase and the sentence length of the candidate sentence in which the key phrase is located;

[0083] Specifically, for each candidate sentence, after determining all key phrases in the candidate sentence, the phrase lengths of all key phrases in the candidate sentence can be calculated, thereby calculating the sentence length of the candidate sentence. Here, the phrase length of a key phrase refers to the number of characters in all key phrases in each candidate sentence, and the sentence length of the candidate sentence containing a key phrase refers to the total number of characters in all characters in the candidate sentence.

[0084] Step 1114: Based on the length ratio, determine whether the candidate sentence is the key sentence.

[0085] Specifically, for each candidate sentence, after statistically obtaining the phrase lengths of all key phrases in the candidate sentence and the sentence length of the candidate sentence, the length proportion of all key phrases in the candidate sentence can be determined based on the phrase length and sentence length, so that whether the candidate sentence is a key sentence can be judged based on the length proportion.

[0086] For example, a preset proportion threshold can be set, and the obtained length proportion can be compared with the preset proportion threshold. If the length proportion exceeds the preset proportion threshold, it indicates that the proportion of key phrases appearing in the candidate sentence exceeds a certain proportion of the total number of words in the candidate sentence. Therefore, it can be determined that the candidate sentence as a whole expresses the semantics of an affirmative response, and thus the candidate sentence can be determined as a key sentence. Here, the preset proportion threshold can be set to 70%, 80%, etc., and can also be set according to actual needs. The embodiment of the present invention does not make specific limitations on this.

[0087] Based on any of the above embodiments, Figure 4 1 is a flow chart of step 120 in the fraud warning method provided by the present invention, as shown in FIG. Figure 4 As shown, step 120 specifically includes:

[0088] Step 121, extracting keywords from each round of conversation data;

[0089] Step 122: Determine the speech type to which each round of dialogue data belongs based on the keywords in each round of dialogue data and the preset keywords under each preset speech type.

[0090] Specifically, when classifying each round of conversation data, keywords can be extracted based on the text content, semantics, and context of the conversation data. Keywords here refer to words or phrases that are of special importance and representativeness in each round of conversation data, which can usually accurately summarize the theme, content, or key information of the conversation data.

[0091] When extracting keywords, you can use a frequency-based approach. This involves segmenting each round of conversation data into Chinese words, calculating the frequency of each word in each round, sorting the words in descending order by frequency, and selecting the most frequently occurring words as keywords. Alternatively, you can use a machine learning approach. This involves training a model using machine learning algorithms (such as text classification and clustering) and then using the model to extract keywords.

[0092] After extracting the keywords of each round of conversation data, for any round of conversation data, the keywords of that round of conversation data can be matched with the preset keywords. For example, string matching or semantic matching methods can be used for matching. Based on the matching results, the type of speech to which the conversation data belongs can be determined. If the keyword matches a preset keyword, it can be determined that the type of speech to which the conversation data belongs is the preset type of speech corresponding to the preset keyword that successfully matched.

[0093] Here, the preset speech type refers to various pre-defined fraud speech patterns, which can be defined based on known fraud cases, expert experience, user reports and other information. Each preset speech type usually corresponds to a specific semantic meaning or function. For example, the preset speech type may include gaining trust, obtaining information and guiding operations.

[0094] It is understandable that before determining the speech type of each round of dialogue data, preset speech types can be defined in advance, and corresponding preset keywords can be constructed based on each preset speech type. These preset keywords can be created manually or automatically generated based on domain knowledge and experience.

[0095] Furthermore, a classification model, such as a support vector machine, naive Bayes, or deep learning model (such as a recurrent neural network or a convolutional neural network), can be trained using an annotated conversation dataset to extract and match keywords for each round of conversation data, and automatically classify the conversation data into the type of speech it belongs to.

[0096] In order to further improve the accuracy of conversation classification, manual classification can be added on the basis of machine pre-classification. By manually adjusting the classification results, the accuracy and reliability of the classification results can be ensured, which will help identify and prevent various types of fraud and improve the effectiveness of fraud warning.

[0097] Based on any of the above embodiments, in step 130, matching the preset speech template based on the arrangement order of the at least one round of dialogue data in the conversation data and the speech type of each round of dialogue data includes:

[0098] Constructing a speech evolution trajectory of the conversation data based on the arrangement order of the at least one round of conversation data in the conversation data and the speech type of each round of conversation data;

[0099] The trajectory of the speech evolution is matched with the trajectory of the fraud speech evolution corresponding to the preset speech template.

[0100] Specifically, a speech evolution trajectory records and tracks the evolution of different speech types during a conversation, based on the content and sequence of the conversation. It demonstrates the changes and transitions of the different speech types involved in the conversation. After determining the speech type associated with each round of conversation data, the speech types associated with each round of conversation data can be organized into an ordered speech evolution trajectory based on the order of the conversation data, thereby obtaining the speech evolution trajectory corresponding to the conversation data set.

[0101] For example, after segmenting the conversation data, we obtain the first-round conversation data, the second-round conversation data, and the third-round conversation data. The types of speech used in the three rounds of conversation data are gaining trust, gaining information, and guiding operations, respectively. Based on the arrangement order and speech type of each round of conversation data, we can construct the speech evolution trajectory of the conversation data as follows: gaining trust → gaining information → guiding operations.

[0102] After constructing the speech evolution trajectory of the session data, the speech evolution trajectory of the session data can be matched with the speech evolution trajectory corresponding to the preset speech template. First, each speech type in the speech evolution trajectory is matched with each speech type in the fraud speech evolution trajectory. If the speech type match is successful, the arrangement order of each speech type in the speech evolution trajectory is compared with the arrangement order of each speech type in the fraud speech evolution trajectory. If the arrangement order is consistent, it can be determined that the trajectory matching result is a successful match, indicating that the fraud speech template has been applied to this session data, there is a high risk of fraud, and the user can be warned of fraud.

[0103] The method provided by the embodiment of the present invention can better understand and analyze the development process and changing trends of the conversation by constructing a speech evolution trajectory, which helps to realize fraud warning based on fraud process comparison, thereby improving the ability to identify and prevent potential fraudulent behaviors.

[0104] Based on any of the above embodiments, Figure 5 This is a flow chart of the steps for constructing the preset speech template provided by the present invention, such as Figure 5 As shown, the steps of constructing the preset speech template include:

[0105] Step 510: Segment each fraudulent session data to obtain at least one round of sample conversation data for each fraudulent session data;

[0106] Step 520: Classify each round of sample conversation data to obtain the speech type to which each round of sample conversation data belongs.

[0107] Step 530: construct the preset speech template based on the arrangement order of the at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data.

[0108] Specifically, fraudulent session data refers to pre-collected multimedia session data used for fraudulent purposes, which may include data in the form of conversational voice, conversational video, and conversational text. After collecting each fraudulent session data, it can be segmented into at least one round of sample conversation data to better analyze and understand the fraudster's rhetoric and fraudulent progress, providing more targeted measures and strategies for preventing and combating fraud, and thus better identifying and preventing fraudulent behavior. In this embodiment of the present invention, the steps for segmenting each fraudulent session data can be the same as the steps for segmenting the session data to be warned in the above-mentioned embodiment, and will not be repeated here.

[0109] For any fraudulent conversation data, after segmenting and obtaining each round of sample conversation data under the fraudulent conversation data, each round of sample conversation data can be classified to obtain the type of speech to which each round of sample conversation data belongs. Here, the steps of classifying each round of sample conversation data can be the same as the steps of classifying each round of conversation data in the above embodiment, and will not be repeated here.

[0110] For any fraud conversation data, after obtaining the speech type of each round of sample conversation data under the fraud conversation data, a preset speech template can be constructed based on the arrangement order of each round of sample conversation data and the speech type it belongs to. Therefore, based on each fraud conversation data, multiple preset speech templates can be constructed.

[0111] It is understandable that fraudulent behavior is constantly evolving and changing. The data of each fraud conversation can be updated regularly, and the preset speech templates can be updated and optimized to maintain their accuracy and practicality, adapt to emerging fraud speech and patterns, and improve the effectiveness of fraud warning.

[0112] The method provided by the embodiment of the present invention can cover various common fraud speech and patterns by pre-building preset speech templates, so as to match the conversation data to be warned with the preset speech templates, and can more comprehensively identify and warn different types of fraud behaviors, thereby improving security and protecting users from potential threats.

[0113] Based on the above embodiment, step 530 specifically includes:

[0114] Determining the evolution trajectory of the fraudulent speech corresponding to each fraudulent conversation data based on the arrangement order of the at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data;

[0115] The number of times each fraudulent speech evolution track is applied in each fraudulent session data is counted, and the preset speech template is determined from each fraudulent speech evolution track based on the number of applications.

[0116] Specifically, the evolution trajectory of fraudulent rhetoric refers to the changing trajectory of the rhetoric patterns and strategies that the fraudster gradually develops and adjusts based on the victim's response and situation during the fraudulent conversation. It can demonstrate the evolution of different rhetoric types in the fraudulent conversation. For any fraudulent conversation data, after determining the rhetoric types of each round of sample conversation data within the fraudulent conversation data, the rhetoric types of each round of sample conversation data can be organized into an orderly fraudulent rhetoric evolution trajectory according to the order of the sample conversation data, thereby obtaining the fraudulent rhetoric evolution trajectory corresponding to the fraudulent conversation data.

[0117] After obtaining the fraudulent speech evolution trajectory corresponding to each fraudulent session data, the number of times each fraudulent speech evolution trajectory is used in each fraudulent session data can be counted. Based on the counted number of uses, the fraudulent speech evolution trajectory with the highest number of uses can be selected as the preset speech template. For example, the fraudulent speech evolution trajectory can be sorted by the number of uses, and the top few fraudulent speech evolution trajectories can be selected as common and general speech templates, thereby obtaining the preset speech template. It should be understood that the fraudulent speech evolution trajectory in each preset speech template can be based on speech type or on sub-types of the speech type.

[0118] It can be understood that the above-mentioned number of applications refers to the frequency or number of times each fraudulent speech evolution trajectory is applied in each fraud conversation data. The more times it is applied, the more widely the fraudulent speech evolution trajectory is used in fraudulent activities.

[0119] The method provided by the embodiment of the present invention can understand the frequency and prevalence of use of each fraudulent speech evolution trajectory in actual fraudulent behavior by counting the number of times each fraudulent speech evolution trajectory is used in each fraudulent session data, thereby helping to predict and prevent future fraudulent behavior and improve the ability to identify and prevent fraud.

[0120] The fraud warning device provided by the present invention is described below. The fraud warning device described below and the fraud warning method described above can be referenced to each other.

[0121] Based on any of the above embodiments, Figure 6 This is a schematic diagram of the structure of the fraud warning device provided by the present invention. Figure 6As shown, the device includes:

[0122] A segmentation unit 610 is configured to segment the conversation data to be warned to obtain at least one round of conversation data;

[0123] A classification unit 620 is used to classify each round of dialogue data to obtain the speech type of each round of dialogue data;

[0124] The early warning unit 630 is used to match the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs with a preset speech template, and to issue a fraud early warning based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0125] The device provided by the embodiment of the present invention divides the conversation data to be warned to obtain at least one round of conversation data, and classifies the conversations of each round of conversation data to obtain the type of speech to which each round of conversation data belongs. Therefore, based on the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs, it can match with the preset speech template and issue a fraud warning based on the matching result, thereby realizing fraud warning based on fraud process comparison, without relying on identification and judgment of incoming call numbers or accounts, and can improve the comprehensiveness and reliability of fraud warning while ensuring timely warning.

[0126] Based on any of the above embodiments, the segmentation unit 610 specifically includes:

[0127] a determination subunit, configured to determine a key sentence indicating an affirmative response from each sentence in the conversation data;

[0128] The segmentation subunit is used to segment the conversation data using the key sentence as the ending sentence of the conversation data to obtain at least one round of conversation data.

[0129] Based on any of the above embodiments, the determination subunit is specifically configured to:

[0130] Determining candidate sentences in which the speaker is the listener from each sentence in the conversation data;

[0131] determining a key phrase indicating an affirmative response from the candidate sentences;

[0132] Determining a length ratio of the key phrase in the candidate sentence based on the phrase length of the key phrase and the sentence length of the candidate sentence in which the key phrase is located;

[0133] Based on the length ratio, determine whether the candidate sentence is the key sentence.

[0134] Based on any of the above embodiments, the classification unit 620 is specifically configured to:

[0135] Extract keywords from each round of conversation data;

[0136] Based on the keywords in each round of dialogue data and the preset keywords under each preset dialogue type, the speech type to which each round of dialogue data belongs is determined.

[0137] Based on any of the above embodiments, the early warning unit 630 is specifically configured to:

[0138] Constructing a speech evolution trajectory of the conversation data based on the arrangement order of the at least one round of conversation data in the conversation data and the speech type of each round of conversation data;

[0139] The trajectory of the speech evolution is matched with the trajectory of the fraud speech evolution corresponding to the preset speech template.

[0140] Based on any of the above embodiments, the further embodiment includes a template construction unit, which specifically does not include:

[0141] A session segmentation subunit, configured to segment each fraud session data to obtain at least one round of sample conversation data under each fraud session data;

[0142] The dialogue classification subunit is used to classify each round of sample dialogue data and obtain the speech type to which each round of sample dialogue data belongs;

[0143] The template construction subunit is used to construct the preset speech template based on the arrangement order of at least one round of sample conversation data in each fraud conversation data and the speech type to which each round of sample conversation data belongs.

[0144] Based on any of the above embodiments, the template construction subunit is specifically used for:

[0145] Determining the evolution trajectory of the fraudulent speech corresponding to each fraudulent conversation data based on the arrangement order of the at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data;

[0146] The number of times each fraudulent speech evolution track is applied in each fraudulent session data is counted, and the preset speech template is determined from each fraudulent speech evolution track based on the number of applications.

[0147] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute a fraud warning method, which includes: segmenting the conversation data to be warned to obtain at least one round of conversation data; classifying each round of conversation data to obtain the speech type to which each round of conversation data belongs; matching the order of arrangement of the at least one round of conversation data in the conversation data and the speech type to which the each round of conversation data belongs with a preset speech template, and issuing a fraud warning based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0148] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0149] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the fraud warning method provided by the above methods, which includes: dividing the conversation data to be warned to obtain at least one round of conversation data; classifying each round of conversation data to obtain the type of speech to which each round of conversation data belongs; matching with a preset speech template based on the arrangement order of at least one round of conversation data in the conversation data and the type of speech to which each round of conversation data belongs, and issuing a fraud warning based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0150] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the fraud warning method provided by the above-mentioned methods, the method comprising: dividing the conversation data to be warned to obtain at least one round of conversation data; classifying each round of conversation data to obtain the type of speech to which each round of conversation data belongs; matching with a preset speech template based on the arrangement order of the at least one round of conversation data in the conversation data and the type of speech to which the each round of conversation data belongs, and issuing a fraud warning based on the matching result; the preset speech template is constructed based on the fraud conversation data.

[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A fraud early warning method, characterized in that: include: Split the conversation data to be warned to obtain at least one round of conversation data; Classify each round of dialogue data to obtain the speech type to which each round of dialogue data belongs; Based on the arrangement order of at least one round of conversation data in the conversation data and the speech type of each round of conversation data, matching is performed with a preset speech template, and a fraud warning is issued based on the matching result; The preset speech template is constructed based on fraud conversation data; The matching of the arrangement order of the at least one round of dialogue data in the conversation data and the speech type of each round of dialogue data with a preset speech template includes: Constructing a speech evolution trajectory of the conversation data based on the arrangement order of each round of conversation data in the conversation data and the speech type of each round of conversation data; The trajectory of the speech evolution is matched with the trajectory of the fraud speech evolution corresponding to the preset speech template.

2. The fraud early warning method according to claim 1, characterized in that: The segmenting of the conversation data to be warned to obtain at least one round of conversation data includes: determining a key sentence indicating an affirmative response from each sentence in the conversation data; The conversation data is segmented using the key sentence as the ending sentence of the conversation data to obtain at least one round of conversation data.

3. The fraud early warning method according to claim 2, characterized in that: Determining a key sentence indicating an affirmative response from each sentence in the conversation data includes: Determining candidate sentences in which the speaker is the listener from each sentence in the conversation data; determining a key phrase indicating an affirmative response from the candidate sentences; Determining a length ratio of the key phrase in the candidate sentence based on the phrase length of the key phrase and the sentence length of the candidate sentence in which the key phrase is located; Based on the length ratio, determine whether the candidate sentence is the key sentence.

4. The fraud early warning method according to claim 1, characterized in that: The dialog classification of each round of dialog data to obtain the speech type of each round of dialog data includes: Extract keywords from each round of conversation data; Based on the keywords in each round of dialogue data and the preset keywords under each preset dialogue type, the speech type to which each round of dialogue data belongs is determined.

5. The fraud early warning method according to any one of claims 1 to 4, characterized in that: The steps of constructing the preset speech template include: Segmenting each fraudulent conversation data to obtain at least one round of sample conversation data under each fraudulent conversation data; Classify each round of sample conversation data to obtain the speech type to which each round of sample conversation data belongs; The preset speech template is constructed based on the arrangement order of at least one round of sample conversation data in each fraud conversation data, and the speech type of each round of sample conversation data.

6. The fraud early warning method according to claim 5, characterized in that: The step of constructing the preset speech template based on the arrangement order of at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data includes: Determining the evolution trajectory of the fraudulent speech corresponding to each fraudulent conversation data based on the arrangement order of the at least one round of sample conversation data in each fraudulent conversation data and the speech type of each round of sample conversation data; The number of times each fraudulent speech evolution track is applied in each fraudulent session data is counted, and the preset speech template is determined from each fraudulent speech evolution track based on the number of applications.

7. A fraud early warning device, characterized in that: include: A segmentation unit, configured to segment the conversation data to be warned to obtain at least one round of conversation data; The classification unit is used to classify each round of dialogue data and obtain the speech type to which each round of dialogue data belongs; an early warning unit, configured to match the order of at least one round of conversation data in the conversation data and the speech type of each round of conversation data with a preset speech template, and issue a fraud early warning based on the matching result; the preset speech template is constructed based on the fraudulent conversation data; The early warning unit is specifically used for: Constructing a speech evolution trajectory of the conversation data based on the arrangement order of each round of conversation data in the conversation data and the speech type of each round of conversation data; The trajectory of the speech evolution is matched with the trajectory of the fraud speech evolution corresponding to the preset speech template.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the fraud warning method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the fraud warning method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Model-based verbal skill recommendation method, device, computer equipment and storage medium

    CN113688221A

  • Depth analysis method and system for fraud data

    CN114579692A