Abnormal communication identification method and device, medium and product

By training a large language model in multiple stages, combining an abnormal communication knowledge base and a risk level classification system, and utilizing thought chain reasoning and reinforcement learning, the problem of low accuracy in abnormal communication identification in existing technologies has been solved, and efficient identification of new and variant abnormal communications has been achieved.

CN120979906AActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511468373.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-18
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify complex and ever-changing abnormal communication behaviors, especially new or variant types of abnormal communication, resulting in low identification accuracy and high costs associated with maintaining complex structured problem templates and rule bases.

Method used

By training a large language model with known and unknown anomalous communication types, an anomalous communication identification model is established. By utilizing an anomalous communication knowledge base and risk level classification system, combined with thought chain reasoning mechanism and reinforcement learning, the model's ability to identify anomalous communication is improved.

Benefits of technology

It achieves accurate identification of known abnormal communication types and has the ability to generalize the identification of new or variant abnormal communication types, thereby improving the accuracy and adaptability of abnormal communication identification and reducing development and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979906A_ABST
    Figure CN120979906A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal communication identification method and device, a medium and a product, and the method comprises the steps: determining a target risk level of communication content between a first object and a second object through an abnormal communication identification model according to an abnormal communication knowledge base and a risk level classification system, the abnormal communication identification model is obtained based on training tasks of known abnormal communication type identification and non-known abnormal communication type identification; the training task of unknown abnormal communication type recognition comprises the steps of obtaining an initial abnormal communication recognition model through the training task of known abnormal communication type recognition, and performing reasoning according to an abnormal communication knowledge base and a thinking chain reasoning mechanism to obtain a predicted abnormal communication recognition result; based on the difference between the abnormal communication identification result and a real abnormal communication identification result, determining a reward signal and updating model parameters, and obtaining an abnormal communication identification model; and determining an abnormal communication identification result according to the target risk level and an abnormal communication identification rule. Therefore, the abnormal communication identification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to an abnormal communication identification method, device, medium and product. BACKGROUND

[0002] With the development of mobile communication technology and voice interaction technology, the number of voice calls received by user terminals has increased significantly.

[0003] Among a large number of communication flows, there are many non-requesting communication behaviors, including marketing promotion, harassment calls, and abnormal communication activities with potential security risks. Such communication often simulates legal service subjects or uses social engineering means to induce user responses, which may then lead to security incidents such as information leakage and property loss.

[0004] Therefore, there is an urgent need for a risk communication behavior identification mechanism with high accuracy to effectively identify abnormal communication behaviors and thereby improve the security protection mechanism of the communication system. SUMMARY

[0005] Based on the above technical status, the present application provides an abnormal communication identification method, device, medium and product, which can improve the identification accuracy of abnormal communication, especially the identification accuracy of new or variant abnormal communication types.

[0006] In order to achieve the above technical purpose, the present application specifically proposes the following technical solutions: According to a first aspect of the embodiments of the present application, a method for identifying abnormal communication is provided, which comprises: obtaining communication content between a first object and a second object; the communication content is obtained by obtaining revisit data between the first object and the second object, and the revisit data is obtained by asking a large language model based on a preset revisit target; determining a target risk level corresponding to the communication content according to a preset abnormal communication knowledge base and a preset risk level classification system through an abnormal communication identification model, wherein the abnormal communication knowledge base includes a plurality of known abnormal communication types and corresponding rhetoric characteristics, the abnormal communication identification model is obtained by performing a multi-stage training task on a large language model based on communication content samples and the abnormal communication knowledge base, the multi-stage training task includes a training task for identifying known abnormal communication types and a training task for identifying unknown abnormal communication types, the unknown abnormal communication type is an abnormal communication type outside the abnormal communication knowledge base; the communication content samples include first communication content samples and second communication content samples, the initial abnormal communication identification model is trained by the first communication content samples through the training task for identifying known abnormal communication types, the training task for identifying unknown abnormal communication types includes: reasoning from the second communication content samples to the real abnormal communication identification result by using a thought chain reasoning mechanism according to the second communication content samples and the abnormal communication knowledge base through the initial abnormal communication identification model, obtaining a predicted abnormal communication identification result; determining a reward signal according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; updating the parameters of the initial abnormal communication identification model according to the reward signal until the training is completed, obtaining the abnormal communication identification model; determining an abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and a preset abnormal communication identification rule, the abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

[0007] In some implementations, the first communication content sample corresponds to a real abnormal communication identification result, and the training task for identifying known abnormal communication types includes: identifying the first content sample for known abnormal communication types by the large language model according to the abnormal communication knowledge base, obtaining a predicted abnormal communication identification result; determining an abnormal communication identification prediction loss according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; converging the large language model based on the abnormal communication identification prediction loss, obtaining an initial abnormal communication identification model.

[0008] In some embodiments, the obtaining the communication content between the first object and the second object comprises: obtaining return visit data corresponding to each of preset return visit slots, to obtain the communication content between the first object and the second object; wherein the return visit slots are obtained based on a return visit target, and the return visit target comprises at least one of identity information of the first object, a communication motive, a communication topic, and a risk operation.

[0009] In some embodiments, the obtaining the return visit data corresponding to each of the return visit slots comprises: generating an abnormal communication question corresponding to each of the return visit slots based on the return visit slot; obtaining reply content of the first object to the abnormal communication question; and filling the return visit slot based on the reply content, to obtain the return visit data corresponding to the return visit slot.

[0010] In some embodiments, the method further comprises: extracting a plurality of dimensional feature labels from the communication content by using the abnormal communication identification model, wherein the plurality of dimensional feature labels comprise at least two of a content summary, a call topic, a risk operation, and identity information of the first object and the second object; and determining the abnormal communication identification result based on the target risk level corresponding to the communication content and a preset abnormal communication identification rule, comprises: determining the abnormal communication identification result based on the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0011] In some embodiments, the determining the abnormal communication identification result based on the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule comprises: determining weights of the plurality of dimensional feature labels based on a mapping relationship between preset scene types and the weights of the plurality of dimensional feature labels; and determining the abnormal communication identification result of the communication content based on the plurality of dimensional feature labels, the weights corresponding to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0012] In some implementations, the communication content includes a plurality of communication contents, the plurality of communication contents respectively correspond to the first object and a plurality of second objects, the plurality of communication contents respectively correspond to target risk levels, the first object is a potential object with abnormal communication behavior, and the second object is a potential object facing a security risk; and the determining of the abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and the preset abnormal communication identification rule includes: determining a risk level distribution result corresponding to the plurality of communication contents according to the target risk levels respectively corresponding to the plurality of communication contents; and determining the abnormal communication identification result of the communication content according to the risk level distribution result corresponding to the plurality of communication contents and the preset abnormal communication identification rule.

[0013] According to a second aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor; the memory is connected with the processor, and is configured to store a program; the processor is configured to realize the abnormal communication identification method according to the first aspect by running the program in the memory.

[0014] According to a third aspect of the embodiments of the present application, a storage medium is provided, and the storage medium stores a computer program; when the computer program is run by a processor, the abnormal communication identification method according to the first aspect is realized.

[0015] According to a fourth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions; when the computer program instructions are run by a processor, the processor executes the abnormal communication identification method according to the first aspect.

[0016] The embodiment of the application provides a method, device, medium and product for identifying abnormal communication, which comprises the following steps: obtaining the communication content between the first object and the second object; the communication content is obtained by obtaining the revisit data between the first object and the second object, and the revisit data is obtained by asking the preset revisit target based on a large language model; determining the target risk level corresponding to the communication content according to the preset abnormal communication knowledge base and the preset risk level classification system through an abnormal communication identification model, wherein the abnormal communication knowledge base comprises a plurality of known abnormal communication types and corresponding rhetoric characteristics, and the abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on the communication content sample and the abnormal communication knowledge base, wherein the multi-stage training task comprises a training task for identifying a known abnormal communication type and a training task for identifying a non-known abnormal communication type, and the non-known abnormal communication type is an abnormal communication type outside the abnormal communication knowledge base; the communication content sample comprises a first communication content sample and a second communication content sample; the initial abnormal communication identification model is obtained by training the first communication content sample through the training task for identifying the known abnormal communication type; the inference from the second communication content sample to the real abnormal communication identification result is performed by using a thinking chain inference mechanism according to the second communication content sample and the abnormal communication knowledge base through the initial abnormal communication identification model, and a predicted abnormal communication identification result is obtained; a reward signal is determined according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; the parameters of the initial abnormal communication identification model are updated according to the reward signal until the training is completed, and the abnormal communication identification model is obtained; the abnormal communication identification result of the communication content is determined according to the target risk level corresponding to the communication content and the preset abnormal communication identification rule, and the abnormal communication identification result is used to indicate whether the communication content is abnormal communication. Since the abnormal communication identification model is trained through the two stages of the training task for identifying the known abnormal communication type and the training task for identifying the non-known abnormal communication type, and the training task for identifying the known abnormal communication type can enable the model to establish a mapping relationship between the communication content sample and the real label, the known abnormal communication type in the abnormal communication knowledge base can be accurately identified, and the model has a preliminary generalization ability, that is, a preliminary ability to capture abnormal information, such as analogy and generalization of known patterns. In the training task for identifying the non-known abnormal communication type, the training target is adjusted to be based on the modeling of the inference process. By introducing the mechanism of reinforcement learning or thinking chain, the inference steps of the above mapping relationship are guided by the large language model to generate the reward signal, which further improves the generalization ability of the large model to capture abnormal behaviors, thereby improving the generalization identification ability of the model to new fraud methods and rapidly evolving variants of fraud. Therefore, the model can not only accurately identify the known abnormal communication type, but also effectively identify the non-known abnormal communication type outside the abnormal communication knowledge base based on the thinking ability of the abnormal communication type learned from the known abnormal communication type. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.

[0018] Figure 1 It is an architecture schematic diagram of an abnormal call research and judgment system in the related art.

[0019] Figure 2 It is a running mechanism principle diagram of a follow-up call system in the related art.

[0020] Figure 3 It is a flowchart of an abnormal communication identification method provided by an embodiment of the present application.

[0021] Figure 4 It is a principle diagram of abnormal communication identification provided by an embodiment of the present application.

[0022] Figure 5 It is a principle diagram of determining an abnormal communication identification result of communication content provided by an embodiment of the present application.

[0023] Figure 6 It is a structure schematic diagram of a follow-up call system provided by an embodiment of the present application.

[0024] Figure 7 It is a flowchart of obtaining communication content provided by an embodiment of the present application.

[0025] Figure 8 It is a structure schematic diagram of an abnormal communication identification device provided by an embodiment of the present application.

[0026] Figure 9 It is a structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions provided by the embodiments of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, servers (including cloud servers), or packaged into software programs to be run. When the hardware devices execute the processing process of the technical solutions of the present application, or the above software programs are run, the automatic splitting of the target task and the automatic calling of the application program interface required by the task can be realized, and the purpose of completing the target task can be achieved. The embodiments of the present application only exemplarily introduce the specific processing process of the technical solutions of the present application, and do not limit the specific implementation form of the technical solutions of the present application. Any technical implementation form that can execute the processing process of the technical solutions of the present application can be adopted by the embodiments of the present application.

[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the scope of protection of the present application.

[0029] Before introducing the solutions of the present application, the related art is introduced first. In recent years, abnormal calling activities using telephone have shown a high incidence trend. Precise identification of improper numbers and taking measures such as closing and blocking the numbers have become one of the important means to crack down on abnormal calling behaviors.

[0030] Based on the behavior characteristics of abnormal calling, which usually calls high-frequency strange numbers, the numbers of potential improper behaviors can be preliminarily screened and circled through communication behavior analysis. On this basis, the revisit technology is further introduced to revisit the users who may be affected, obtain their communication with the improper numbers, and analyze the content of the revisit conversation combined with the text semantic understanding technology, so as to realize the comprehensive research and judgment of whether the improper numbers belong to abnormal calling behavior numbers.

[0031] Figure 1 The schematic diagram of the architecture of the abnormal calling research and judgment process in the related art is shown in FIG. 1. Figure 1 As shown in FIG. 1, the abnormal calling research and judgment process includes: collecting the telephone communication content between the improper behavior numbers (i.e. potential abnormal numbers) and the users who may be affected (i.e. potential victims) by the revisit system to generate the revisit call result; the research and judgment system then performs comprehensive analysis based on the revisit call results to determine whether the improper behavior numbers belong to abnormal calling behaviors or other legal identities, such as express delivery personnel or online taxi drivers, etc.

[0032] Figure 2 The principle diagram of the operation mechanism of the revisit system in the related art is shown in FIG. 2. Figure 2 As shown in FIG. 2, the traditional revisit system usually sets structured questions to ensure the effectiveness and efficiency of information collection.

[0033] Specifically, first, a library of question templates with explicit options or openness needs to be designed based on the target of the follow-up. For example, whether the user recognizes the scam number (i.e., a potential abnormal number), and the corresponding options include recognition, non-recognition, and others. When the user selects non-recognition, further understand the call content and determine which pre-set option the call content belongs to. Next, further understand whether there is a risk operation in the call content, and give fixed options of the risk operation. These question templates aim to guide the potentially affected user to provide key information. Then, based on the user's replies to different questions, a tree-like dialogue control logic structure is organized to dynamically adjust the subsequent questioning path, ensuring the comprehensiveness and accuracy of information collection. Finally, the understanding of the user's replies relies on various natural language processing techniques, such as named entity recognition (NER), syntax analysis, intent recognition, semantic slot filling, etc., to extract the core content and key demands of the user's feedback.

[0034] Existing research and judgment schemes related to irregular behavior are deeply bound to the follow-up system. For single-pass follow-up calls, the research and judgment system related to irregular behavior needs to obtain the key content of abnormal communication based on the follow-up dialogue results, including the user identity of the abnormal communication, the theme of the call, risk operations, and other content, and determine the abnormal risk level corresponding to the follow-up call through artificially defined fusion rules. For example: if the communication content involves education and training, and the risk operation is voluntary refund, then the abnormal risk level corresponding to the call is high risk. When multiple follow-up calls are judged as high risk, the number related to irregular behavior will be determined as an abnormal call number, and the number will be shut down.

[0035] Although the existing research and judgment scheme related to irregular behavior is effective to some extent, it still has the following defects: In real dialogue scenarios, users have various expression habits, including diction habits, colloquial expressions, dialect accents, and transcription errors, which can cause the follow-up system to deviate in understanding user replies, affecting the quality of follow-up.

[0036] Secondly, in order to support the identification of diversified fraud means, it is necessary to maintain complex structured question templates, corresponding option templates, and corresponding transition relationships. However, as the number of intents increases and the complexity of expression increases, designing, writing, testing, maintaining, and managing a large number of grammar rules and template libraries becomes extremely tedious and time-consuming. Adding new intents or modifying existing intents involves extensive adjustments and testing, which not only has poor scalability but also increases development and maintenance costs.

[0037] Moreover, the existing follow-up system cannot solve the context relationship and reasoning semantic problems in the follow-up dialogue, leading to an inability to accurately understand abnormal behavior, and even missing known abnormal means.

[0038] Moreover, existing identification schemes related to irregular behavior rely on the pre-designed relevant question templates and option templates of the callback system, which limits their ability to identify new types of abnormal means. Due to the defects of the callback system and the complexity of user expression, existing schemes cannot fully understand the abnormal behavior exposed in the callback conversation.

[0039] In summary, the current abnormal call research and judgment system has certain functions, but still faces many challenges in dealing with complex and variable user expressions, maintaining complex structured question templates, and handling context relationships, resulting in low accuracy in identifying abnormal calls.

[0040] Therefore, the embodiments of the present application aim to provide an abnormal communication method, device, medium and product. Through the training tasks of known abnormal communication type identification and unknown abnormal communication type identification of the large language model, an abnormal communication identification model is obtained, so that the model can not only accurately identify known abnormal communication types, but also learn common rules and features from known abnormal communication types, and then effectively identify unknown abnormal communication types with similar abnormal characteristics through reasoning. In the following embodiments, each is described in detail.

[0041] Exemplary method Figure 3 A flowchart of an abnormal communication identification method provided by the embodiments of the present application. As shown in Figure 3 The abnormal communication identification method provided by the embodiments of the present application includes steps S101-S103: S101, obtaining the communication content between the first object and the second object.

[0042] Among them, the first object is a potential object with abnormal communication behavior, that is, an object that may have abnormal communication behavior after preliminary analysis. These abnormal communication behaviors include: high-frequency calls to unknown numbers, calls to multiple different regions in a short period of time, and other abnormal behavior patterns. These behaviors often indicate that the object may have engaged in some irregular activities, such as fraudulent operations through telephone (i.e. telephone fraud). Since there is no concrete evidence, it is temporarily referred to as a potential object with abnormal communication behavior.

[0043] The second object is an object that may face security risks. The second object can be determined based on the interaction with the first object. For example, if a number is marked as frequently receiving calls, messages or instant communications from objects with abnormal communication behavior, the number will be considered as a potential security risk object. This indicates that the second object may be the target or victim of abnormal communication behavior.

[0044] In some embodiments, the communication content between the first object and the second object can be directly obtained, or the second object can be interviewed using a revisit technique to understand the specific background, theme and intention of the communication between the two parties, so as to obtain the communication content between the first object and the second object by obtaining the revisit data between the first object and the second object.

[0045] In order to improve the quality of the communication content, the revisit data can also be obtained by asking the large language model based on the preset revisit target. The revisit target is obtained by setting a series of targets to guide the large language model to ask questions to obtain the information required by the revisit target, without limiting the specific question form, that is, how to ask questions. The specific implementation process of obtaining the communication content through the revisit data will be described in detail in subsequent embodiments.

[0046] Among them, the communication content not only includes the actual dialogue audio, but also includes some possible text information such as short message or instant message. In some examples, the communication content can be the communication content of telephone fraud.

[0047] It should be noted that the acquisition of the communication content between the first object and the second object needs to comply with laws and regulations and ensure that privacy protection measures are in place. Specifically, relevant call records and voice data can be obtained from telecom operators through legal authorization. And when performing communication content acquisition, appropriate encryption and protection measures should be taken to ensure the security and privacy of the data. At the same time, all operations must comply with local data protection regulations and privacy policies to ensure that the legal rights and interests of users are not infringed.

[0048] S102, determining the target risk level corresponding to the communication content according to the preset abnormal communication knowledge base and the preset risk level classification system through the abnormal communication identification model.

[0049] Among them, the abnormal communication knowledge base includes a plurality of known abnormal communication types and corresponding dialogue characteristics. For example, common abnormal communication types such as brushing single rebate scams, loan frauds, impersonating customer service personnel, emotional frauds (“pig killing disc”), and false investment and financial management, as well as typical dialogue characteristics and behavior patterns of each type.

[0050] The risk level classification system defines different levels of risk levels, including a first risk level, a second risk level, a third risk level, and a fourth risk level, which can help the model to classify the risk level of the communication content.

[0051] The first risk level can be high risk, and a core fraud link is explicitly mentioned or implemented, such as asking for a bank card password, a verification code, requiring a money transfer to a so-called "safe account", inducing the download of a high-risk application (APP), and meeting a specific fraud speech pattern. The second risk level can be medium risk, which has highly suspicious guiding speech, may be in the early stage of fraud induction (such as obtaining identity information and establishing emotional trust), but has not yet directly requested a transfer or leaked a core password. The third risk level can be low risk, which contains some sensitive words or slight induction behavior, but the overall intention is not clear or has not constituted a direct threat. The fourth risk level can be no risk, which means that the communication content is normal and has no suspicious abnormal communication characteristics.

[0052] Figure 4 A schematic diagram of an abnormal communication identification provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, by re-visiting a plurality of second objects associated with a potential abnormal number, a plurality of communication contents can be obtained. Then, the communication contents, the preset abnormal communication knowledge base and the preset risk level classification system are input into an abnormal communication identification model in the form of prompt words, so that the abnormal communication identification model performs risk identification on the communication contents to obtain a target risk level corresponding to the communication contents. Figure 4

[0053] The abnormal communication identification model can be obtained by performing a multi-stage training task on a large language model based on the communication content sample, the preset abnormal communication knowledge base and the preset risk level classification system. The multi-stage training task includes a known abnormal communication type identification training task and a non-known abnormal communication type identification training task. The non-known abnormal communication type is an abnormal communication type outside the abnormal communication knowledge base, that is, an abnormal communication type not included in the abnormal communication knowledge base.

[0054] Specifically, the communication content sample includes a first communication content sample and a second communication content sample. The known abnormal communication type identification training task can be used to train an initial abnormal communication identification model by using the first communication content sample.

[0055] The first communication content sample and the second communication content sample can be partially overlapped, or completely different, or completely the same. This embodiment does not limit this.

[0056] In this embodiment, the known abnormal communication type identification training task can be first performed on the large language model based on the first communication content sample and the preset abnormal communication knowledge base to obtain an initial abnormal communication identification model; and the non-known abnormal communication type identification training task can be performed on the large language model based on the second communication content sample and the abnormal communication knowledge base to obtain the abnormal communication identification model.

[0057] ​The two training tasks are described in detail as follows: In some embodiments, the first communication content sample corresponds to a real abnormal communication recognition result. The training task of known abnormal communication type recognition includes: identifying the known abnormal communication type of the first communication content sample according to the abnormal communication knowledge base through the large language model, obtaining a predicted abnormal communication recognition result; determining the prediction loss of abnormal communication recognition according to the difference between the predicted abnormal communication recognition result and the real abnormal communication recognition result; and converging the large language model based on the prediction loss of abnormal communication recognition to obtain an initial abnormal communication recognition model.

[0058] Specifically, the first communication content sample can be obtained through public data sets, web crawling or direct collection in real scenarios. The first communication content sample can be obtained by obtaining the actual communication content between the first object and the second object. The actual communication content includes call recording, short message or instant message, etc. In addition, the actual communication content between the first object and the second object can also be obtained by interviewing the second object.

[0059] Among them, the first communication content sample corresponds to a labeled real abnormal communication recognition result, which indicates the abnormal communication type and / or risk level corresponding to the first communication content sample, thereby providing a basic label for model training and ensuring the accuracy of subsequent evaluation.

[0060] During the training process, the preset abnormal communication knowledge base and the preset risk level classification system can be input into the large language model in the form of prompt words. The abnormal communication knowledge base can help the large language model learn how to identify these abnormal communications. The risk level classification system helps the model to classify the identified abnormal communications more accurately.

[0061] At the same time, the first communication content sample is input into the large language model, which identifies the risk level classification corresponding to the first communication content sample and whether it belongs to abnormal communication according to the multiple known abnormal communication types in the abnormal communication knowledge base, thereby generating a predicted abnormal communication recognition result.

[0062] In order to ensure the accuracy and robustness of the model, the predicted abnormal communication recognition result needs to be compared with the real abnormal communication recognition result, the difference between the two is calculated, and the abnormal communication prediction loss is determined. And use the supervised fine-tuning (Supervised Fine-Tuning, SFT) strategy to iteratively optimize the large language model according to the loss value, gradually adjust the model parameters until the predetermined performance indicators are reached.

[0063] After the large language model is fully trained and optimized through the first-stage training task, an initial abnormal communication identification model can be obtained. This model not only can accurately identify various known abnormal communication types, but also can accurately classify their risk levels, and initially has the ability to capture abnormal information related to improper behavior in the call content.

[0064] In order to further improve the model's ability to identify abnormal communication behavior in complex real scenarios, and enhance its generalization and adaptability to other non-known abnormal communication types other than known abnormal communication types, such as new improper communication behavior and rapidly evolving variants, the second-stage training task can be introduced on the basis of the completion of the first-stage training task, to break the model's dependence on known abnormal communication patterns, so that it can cope with new improper communication means evolving constantly, including semantic camouflage, cross-scene fusion, and speech variation, etc. "variant" behaviors with strong concealment. The second-stage training task will be described in detail below: In some embodiments, the initial abnormal communication identification model can be obtained through the known abnormal communication type identification training task, and the non-known abnormal communication type identification training task includes: obtaining a predicted abnormal communication identification result by using a thought chain reasoning mechanism to infer the second communication content sample to the real abnormal communication identification result according to the second communication content sample and the abnormal communication knowledge base through the initial abnormal communication fraud identification model; determining a reward signal according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; updating the parameters of the initial abnormal communication identification model according to the reward signal until the training is completed, to obtain an abnormal communication identification model.

[0065] The abnormal communication identification model has the ability to identify known abnormal communication types and other non-known abnormal communication types (such as new abnormal communication types or variant abnormal communication types) other than known abnormal communication types.

[0066] During the second-stage training process, the model can use the Chain-of-Thought Reasoning (CoT) mechanism to logically infer the input second communication content sample based on the knowledge system and context understanding ability learned in the first stage, combined with the built-in abnormal communication knowledge base, and output a predicted abnormal communication identification result for the second communication content sample by analyzing the semantic clues, behavior sequences and potential intentions in the second communication content sample. The coping strategies include: suggesting to ban, limiting communication rights, starting an artificial review process, triggering secondary identity verification, suggesting to revisit confirmation, strengthening monitoring or pushing safety reminders, etc.

[0067] Subsequently, a reward signal can be determined according to a difference between the predicted abnormal communication identification result and the real abnormal communication identification result. In some examples, the reward signal includes: a positive reward is given if a non-known high-risk behavior is successfully identified and an effective treatment suggestion is proposed; a negative reward is given if a missed judgment (e.g., a real fraud is judged as no risk) or a misjudgment (e.g., a normal service is misjudged as a high risk) occurs; a neutral or slightly positive reward is given if a conservative but reasonable cautious strategy (e.g., "suggestion for return visit confirmation" or "transfer into manual review") is proposed for a fuzzy scene.

[0068] Further, the reward signal is used as a target function, and the model parameters are updated by a policy gradient method. With the dynamic expansion of the training data and the continuous accumulation of the feedback signal, the sensitivity and discrimination ability of the model to atypical and highly concealed abnormal behaviors in a complex semantic space can be improved, and the sensitivity and judgment ability to atypical abnormal behaviors can be gradually mastered. When the comprehensive performance indicators (such as risk identification accuracy, strategy adoption rate, false alarm rate control, etc.) of the model on an independent verification set tend to be stable and reach a preset threshold, the training process ends, and an abnormal communication identification model with strong generalization ability is obtained. The model not only has accurate identification ability for known abnormal communication patterns, but also has generalization identification, context reasoning and intelligent decision-making ability for non-known irregular communication behaviors and their rapid evolution variants.

[0069] In summary, in the first stage of the training task, the model can establish a mapping relationship from the communication content sample to the real label by learning a large number of real cases, so it can accurately identify known abnormal communication types in the knowledge base. In the second stage of the training task, the training target is adjusted from label prediction to modeling based on the reasoning process. By introducing reinforcement learning or thinking chain mechanism, the reward signal is used to guide the large language model to generate the reasoning steps of the above mapping relationship. Under this mechanism, the large language model is no longer limited to memorizing the static mapping relationship between the sample and the label, but learns the internal reasoning ability of identifying abnormal communication types, i.e., "how to think to make correct judgments". This ability enables it to abstract the essential rules from known abnormal patterns, so that when faced with new or variant abnormal communications that have not been seen before, it can still recognize their deep association with known abnormal communication types through logical deduction, and then make accurate identification judgments, achieving accurate identification of new or variant abnormal communication types.

[0070] S103、According to the target risk level corresponding to the communication content and the preset abnormal communication identification rule, determine the abnormal communication identification result of the communication content, and the abnormal communication identification result is used to indicate whether the communication content is an abnormal communication.

[0071] Continuing to refer to Figure 4In some embodiments, in order to more accurately determine the abnormality of the communication content, the method of the embodiment further includes: extracting a plurality of dimensional feature labels from the communication content by the abnormal communication identification model; and determining an abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and a preset abnormal communication identification rule. The plurality of dimensional feature labels include at least two of a content summary of the communication content, a call topic, a risk operation, and identity information of the first object and the second object.

[0072] The content summary of the communication content represents the summary of the communication content; the identity information includes an operator, a courier, a delivery man, a ride-hailing driver, a hotel, a marriage platform, a tour guide, a travel agency, a teacher, a classmate, etc.; the call topic includes flight delay, loan problem, notification reminder, travel-related, insurance-related, education-related, operator business, etc.; and the risk operation includes downloading an app, asking for a verification code, actively refunding, transferring / recharging, logging in and clicking a link website, etc.

[0073] In some embodiments, determining the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule includes: determining the weights of the plurality of dimensional feature labels; and determining the abnormal communication identification result of the communication content according to the plurality of dimensional feature labels and their respective weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0074] Since the importance of each dimensional feature label is different in different business scenarios, different weights can be configured for the plurality of dimensional feature labels for different scene types to reflect the degree of attention to each feature label in different scenarios. On this basis, determining the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule includes: determining the weights of the plurality of dimensional feature labels according to the mapping relationship between the scene type and the weights of each dimensional feature label; and determining the abnormal communication identification result of the communication content according to the plurality of dimensional feature labels and their respective weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0075] The mapping relationship between the scene type and the weights of each dimensional feature label can be as shown in Table 1 below: Table 1 Mapping relationship between scene type and weights of each dimensional feature label The following takes the "express service" scenario as an example to illustrate how the weight configuration affects the final abnormal communication judgment result.

[0076] Assume that the first object is an employee of a certain express company (identity information labeled as "courier"), who has made 10 calls with multiple second objects (recipients) in one day. By identifying the risk of each call content, the following results can be obtained: 9 times of call content summary are "delivery notice", "package has been signed for", etc., with no sensitive words, identified as "no risk"; 1 time of call mentions "package loss can apply for compensation", but no induced operation, identified as "medium risk"; All call topics are "express notice" or "delivery follow-up"; There is no clear risk operation behavior (such as asking for verification code, guiding transfer, providing non-official link, etc.); The identity information of the first object is "courier".

[0077] Although this communication behavior belongs to the "express service" scene, but in view of the fact that abnormal communication behaviors such as impersonating customer service for compensation and inducing refund operation have occurred frequently in the name of express delivery in recent years, the corresponding weights in this scene are set as follows: Risk operation: 0.4 (high weight, although the express industry is a routine business, it is easy to use compensation, refund and other tactics for induced operation) Call topic: 0.3 (medium weight, the topic should be related to express) Content summary: 0.2 (lower weight, allowing certain sensitive words to exist, such as "compensation") Identity information: 0.1 (low weight, identity can be fake or account stolen, only "courier" label is not enough to fully trust) Further, the comprehensive score can be calculated as follows: Identity information score: 1.0 (legal courier) × 0.1 = 0.1 Call topic score: 0.9 (all related to express) × 0.3 = 0.27 Content summary score: 0.8 (1 time of medium risk, the rest are no risk) × 0.2 = 0.16 Risk operation score: 1.0 (no risk operation) × 0.4 = 0.4 Comprehensive risk score = 0.1 + 0.27 + 0.16 + 0.4 = 0.93 (the higher the better) Combined with the preset rules: Rule setting: "If the first object's identity is 'courier' and the call topic is 'express notice', even if there is a small amount of medium risk content summary, as long as no risk operation behavior is triggered, no high risk warning will be triggered." Therefore, the abnormal communication identification result is: the communication behavior of this number is in line with the normal business characteristics, and is considered as legal and compliant business communication.

[0078] For example, in a financial service scenario, even if the identity information is reliable, if the content summary involves "investment rebate" and the risk operation is frequent, the pre-warning of abnormal communication can still be triggered.

[0079] The embodiment can flexibly adjust the weights of various feature labels according to different business scenarios, avoid "one-size-fits-all" risk judgment, and improve the accuracy and business adaptability of identification. In particular, in high-frequency and low-risk communication scenarios such as express delivery and customer service, reasonable configuration of high weights of identity information and call topics can effectively reduce the false positive rate and ensure the smoothness of normal business communication.

[0080] It should be noted that only the weight configuration of the feature label of each dimension in some scenarios is given in the above table, and the weight configuration coefficient therein is exemplary, and those skilled in the art can set reasonable weight coefficients according to actual business scenario requirements, and the embodiment does not make specific limitations.

[0081] Continue to refer to Figure 4 , and refer to Figure 5 In some embodiments, since in actual communication scenarios, the first object usually communicates with multiple second objects. Therefore, for each second object, the first object and it may have one or more communication behaviors in a period of time, forming corresponding communication content. When there are multiple communication contents between the first object and the second object, the integration processing can be performed to generate a unique comprehensive communication content between the first object and each second object. Thus, between the first object and N second objects, N corresponding communication contents will be formed. The integration processing includes text aggregation, time window merging, or session merging, etc.

[0082] For the communication content between the first object and each second object, the corresponding target risk level can be obtained. On this basis, step S103 includes: determining a risk level distribution result corresponding to the multiple communication contents according to the target risk levels corresponding to the multiple communication contents respectively; determining an abnormal communication identification result of the communication content according to the risk level distribution result corresponding to the multiple communication contents and a preset abnormal communication identification rule.

[0083] In the embodiment, first, the risk level distribution result is generated according to the target risk levels of the N communication contents respectively. For example, the number proportion of high-risk communication content, the concentration of medium-high-risk communication object, or the risk level histogram distribution characteristics are counted and generated. The distribution result reflects the risk situation of the overall communication behavior of the first object.

[0084] Subsequently, the risk level distribution results are matched with preset abnormal communication identification rules. For example, preset rules may include: "If the proportion of high-risk communication content exceeds 30%, then abnormal communication behavior is determined to exist"; or "If the communication content with three or more different second objects is identified as high-risk, and the communication content contains the same sensitive keywords, then it is determined to be batch abnormal communication", etc.

[0085] Ultimately, based on the matching results, abnormal communication identification results will be generated, such as outputting "abnormal communication behavior exists" or "communication behavior is normal", and further triggering subsequent processing measures such as alarms, manual review or communication blocking.

[0086] Continue reading Figure 5 In some embodiments, since the judgment rules differ for different regions, the target risk level can be adjusted after the abnormal communication identification model outputs the target risk level. For example, the target risk level can be adjusted through manual intervention (see the bold and black text in the figure), and a risk level distribution result can be generated based on the adjusted target risk level.

[0087] As mentioned earlier, the communication content between the first and second objects can be obtained through follow-up technology. In this case, the completeness and richness of the information contained in the follow-up dialogue directly affect the quality of the follow-up dialogue, and thus affect the effect of subsequent comprehensive judgment. In order to fully explore the key information in the communication process and improve the information density and interaction quality of the follow-up dialogue, in some embodiments, when obtaining the communication content between the first and second objects in step S101, it specifically includes: obtaining the follow-up data corresponding to each preset follow-up slot to obtain the communication content between the first and second objects; wherein, each follow-up slot is set based on the follow-up target, and the follow-up target includes at least one of the following: the identity information of the second object, the communication motivation, the communication topic, and the risky operation.

[0088] Figure 6 This is a schematic diagram of the structure of the follow-up system provided in an embodiment of this application. Figure 6 As shown, the follow-up system includes a predefined slot module 201, a target status tracking and management module 202, a dialogue engine module 203, a slot parsing and filling module 204, and a follow-up questioning and process control logic module 205. The slot marking predefinition module 201 is configured to provide a setting function of the revisit target. A user can define a complete information set to be collected in the current revisit in advance through the slot marking predefinition module 201, so as to set the revisit target. The revisit target includes a series of structured slots, each of which represents a key data item to be obtained, for example, identity information, communication motivation, communication topic, risk operation, etc. of the first object. In addition, each slot can be configured with a data type and a necessity degree. The data type includes text, number, date, or option list, etc. The necessity degree includes mandatory or optional.

[0089] The target state tracking and management module 202 is configured to maintain the state of the current revisit session in real time. It dynamically tracks the filling state (filled / unfilled) of all preset slots, the obtained slot values, and whether the slot values pass the verification (such as format verification, range verification). The manager is the control center of the whole system.

[0090] Figure 7 A flowchart for obtaining communication content is provided for the embodiments of the present application. As shown in Figure 7 The preset respective revisit slots are obtained, and the respective revisit data corresponding to the respective revisit slots is obtained, including the following steps S701-S703: S701, based on each of the respective revisit slots, an abnormal communication question corresponding to the revisit slot is generated.

[0091] This step is to collect information from the respective revisit slots in sequence to obtain key context information for identifying abnormal communication behavior. Continue to refer to Figure 6 After completing the information filling of the i th revisit slot, the next revisit slot that needs to collect information can be selected from the unfilled revisit slots as the i+1 th revisit slot based on the preset scheduling rule and the current filling state of the respective revisit slots through the dialogue engine module 203. Subsequently, a natural, fluent, and semantically accurate question sentence is dynamically generated for the selected i+1 th revisit slot by using the capability of the large language model and combining the context of the current dialogue. The question sentence can accurately focus on the information demand corresponding to the revisit slot, and is used to guide the revisited object to provide effective feedback related to the abnormal communication behavior.

[0092] The preset rule includes a priority order of the respective revisit slots and a logical dependency relationship between the respective revisit slots. The priority order of the respective revisit slots can ensure the priority acquisition of key information, and the logical dependency relationship between the respective revisit slots can ensure that the question process conforms to a reasonable reasoning path and business logic, thereby improving the efficiency and integrity of information collection.

[0093] S702, the reply content of the first object to the abnormal communication question is obtained.

[0094] With reference to Figure 6 , this step can be implemented by the slot parsing and filling module 204. This module can utilize the powerful natural language understanding capabilities of large language models to perform deep semantic analysis on the reply content of the first object, identifying key information related to the revisit goal.

[0095] S703, based on the reply content, fill the revisit slot to obtain the revisit data corresponding to the revisit slot.

[0096] Further, the identified key information can be structured by the slot parsing and filling module 204, and then filled into the corresponding revisit slot to complete the automatic filling of information.

[0097] To evaluate the completeness and effectiveness of the information extraction, the slot parsing and filling module 204 can also perform basic slot value verification on the filling results. The basic slot value verification includes the following states: Successful filling: the key information is complete and the semantics is clear, and has been accurately filled into the corresponding slot; Partial filling: only part of the valid information is extracted, and subsequent follow-up is needed to complete the content; Invalid slot value: the identified information does not meet the pre-set semantic range or logical rationality, and is considered as invalid input; Unidentified: the reply content does not contain the information required by the target slot, or the information is expressed ambiguously and cannot be parsed.

[0098] The above filling states and verification results will be fed back to the target state tracking and management module 202 in real time, which is used to update the overall progress state of the revisit dialogue, guide the adjustment of subsequent dialogue strategy and slot scheduling decision, so as to ensure the coherence of the revisit process and the systematicness of information collection.

[0099] Further, if the target state tracking and management module 202 receives the "invalid slot value" or "unidentified" state feedback from the slot parsing and filling module 204, and the current processed revisit slot is a mandatory item, the follow-up and process control logic module 205 will trigger the dialogue engine module 203 to generate a targeted follow-up question. The follow-up question aims to guide the second object to provide valid information again by clarifying semantics, using more explicit expression methods, or adjusting the angle of questioning, so as to improve the success rate of key information acquisition.

[0100] In this mechanism, flexible follow-up strategies can be preset, such as limiting the maximum number of follow-up questions for the same slot (e.g. no more than 3 times), to balance the information collection efficiency and user experience. If the valid filling is completed within the specified number of times, the target state tracking and management module 202 will update the state of the slot to "filled", and simultaneously maintain the overall revisit progress.

[0101] Subsequently, the dialogue engine module 203 selects the next target slot to be collected from the remaining unfilled slots based on the updated target state and in combination with preset scheduling rules (such as slot priority, logical dependency, etc.), and generates a corresponding question sentence, entering a new round of information collection process, i.e., repeating steps S501 to S503.

[0102] The cyclic information collection process continues to iterate until any of the following termination conditions is met: all mandatory interview slots have been successfully filled with valid values, the interview target has reached the preset completeness standard (e.g., all mandatory items are completed, and the filling ratio of optional items reaches the set threshold), the user actively expresses the intention to end the conversation (such as explicitly replying "no longer continue" or being unresponsive for a long time), or the system determines that the current information is sufficient to support subsequent comprehensive research and judgment according to preset rules.

[0103] When any of the termination conditions is met, an ending sentence will be triggered to terminate the current interview process.

[0104] The embodiment sets the interview target based on at least one of the first object's identity information, communication motivation, communication topic, and risk operation, and constructs various interview slots. Each interview slot corresponds to a dimension of information to be collected, and multiple interview slots form an information collection framework, ensuring comprehensive coverage of key information dimensions.

[0105] On this basis, by utilizing the powerful semantic understanding and natural language generation capabilities of large language models, the question sentence is dynamically generated in a semantic-driven manner guided by the interview target, without being limited by preset fixed question templates and limited options, and can adaptively construct natural, coherent, and semantically accurate question content according to the actual dialogue context.

[0106] Further, the question wording and expression method can be flexibly adjusted according to the second object's language style, spoken expression habits, and even dialect characteristics, enhancing the intelligibility of the dialogue, thereby more effectively guiding the first object to provide real and detailed feedback information.

[0107] Therefore, the embodiment can realize the transformation of the interview process from "passive response" to "active inquiry", not only being able to actively track missing information and dynamically adjust the questioning path, but also being able to accurately identify and extract key semantic content in complex expressions, thereby significantly improving the completeness, accuracy, and richness of information collection, providing high-quality data support for subsequent comprehensive risk research and judgment.

[0108] In summary, the embodiments of the present application realize dynamic management of follow-up targets through a large language model. With the generalization of the large model and the advantages in semantic understanding and context modeling, the ability to understand the intent of user responses is significantly enhanced. On this basis, further combined with the current progress state of the follow-up target, the subsequent question content is adaptively generated, and the intelligent promotion of the dialogue process is realized. Compared with the traditional follow-up system based on fixed templates and explicit state transitions, the embodiments effectively reduce the manual configuration and maintenance cost of question templates, answer options and state transition logic. At the same time, due to the deep understanding ability of the large language model to natural language, it can have stronger semantic coverage and scene adaptability, significantly improving the quality of follow-up dialogue.

[0109] Further, by introducing an abnormal communication identification model, since the model performs two-stage training tasks, it not only has accurate identification ability for known abnormal communication types, but also can improve the generalization identification and risk assessment ability for unknown new abnormal communication types and rapidly evolving variant abnormal communication types.

[0110] In addition, in the research and judgment process, the model extracts multiple dimensional feature labels at the same time to form a structured risk feature vector, and cooperates with the pre-configured abnormal communication identification rules to flexibly process the abnormal communication numbers. Specifically, the strictness of the research and judgment logic can be dynamically adjusted according to the business scene demand. For example, in a specific high-risk scene (such as "teaching and training industry refund call"), a strong rule strategy can be configured: once the relevant theme and sensitive language are identified, a high-risk determination is triggered and an automatic shutdown operation is performed. Thus, the fine and scene management of risk prevention and control strategy is realized, while ensuring the normal communication of business, the response ability and disposal efficiency of new and variant fraud behaviors are significantly improved.

[0111] Exemplary device Corresponding to the abnormal communication identification method described above, the embodiments of the present application also provide an abnormal communication identification device. Figure 8 is a structural schematic diagram of an abnormal communication identification device provided by the embodiments of the present application. As shown in Figure 8As shown, the abnormal communication identification device provided by the embodiments of the present application comprises: an acquisition unit 601, a risk level identification unit 602, and an abnormal communication identification unit 603; wherein the acquisition unit 601 is configured to acquire communication content between a first object and a second object; the communication content is obtained by acquiring return visit data between the first object and the second object, and the return visit data is obtained by asking a large language model based on a preset return visit target; the risk level identification unit 602 is configured to determine a target risk level corresponding to the communication content according to a preset abnormal communication knowledge base and a preset risk level classification system through an abnormal communication identification model, the abnormal communication knowledge base comprises a plurality of known abnormal communication types and corresponding dialogue characteristics, the abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on communication content samples and the abnormal communication knowledge base, the multi-stage training task comprises a known abnormal communication type identification training task and a non-known abnormal communication type identification training task, the non-known abnormal communication type is an abnormal communication type outside the abnormal communication knowledge base; the communication content samples comprise first communication content samples and second communication content samples, the known abnormal communication type identification training task is used to train an initial abnormal communication identification model through the first communication content samples, the non-known abnormal communication type identification training task comprises: using a thought chain reasoning mechanism to perform reasoning from the second communication content samples to a real abnormal communication identification result according to the second communication content samples and the abnormal communication knowledge base through the initial abnormal communication identification model, obtaining a predicted abnormal communication identification result; determining a reward signal according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; updating the parameters of the initial abnormal communication identification model according to the reward signal until the training is completed, and obtaining the abnormal communication identification model; the abnormal communication identification unit 603 is configured to determine an abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and a preset abnormal communication identification rule, and the abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

[0112] In some embodiments, the first communication content sample corresponds to a real abnormal communication identification result, and the known abnormal communication type identification training task comprises: identifying the first content sample according to the abnormal communication knowledge base through the large language model to identify the known abnormal communication type, obtaining a predicted abnormal communication identification result; determining an abnormal communication identification prediction loss according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; and converging the large language model based on the abnormal communication identification prediction loss to obtain an initial abnormal communication identification model.

[0113] In some embodiments, the obtaining unit 601 obtains the communication content between the first object and the second object, specifically comprising: obtaining the return visit data corresponding to each of the preset return visit slots, to obtain the communication content between the first object and the second object; wherein the return visit slots are obtained based on a return visit target, and the return visit target includes at least one of the identity information of the first object, the communication motive, the communication theme and the risk operation.

[0114] In some embodiments, the obtaining of the return visit data corresponding to each of the preset return visit slots comprises: generating an abnormal communication question corresponding to each of the return visit slots based on the return visit slot; obtaining the reply content of the first object to the abnormal communication question; and filling the return visit slot based on the reply content to obtain the return visit data corresponding to the return visit slot.

[0115] In some embodiments, the device further comprises an extraction unit 604 configured to extract a plurality of dimensional feature labels from the communication content by the abnormal communication identification model, wherein the plurality of dimensional feature labels include at least two of the content summary, the call theme, the risk operation of the communication content, and the identity information of the first object and the second object; and wherein the abnormal communication identification unit 603 determines the abnormal communication identification result according to the target risk level corresponding to the communication content and the preset abnormal communication identification rule, comprising: determining the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content and the preset abnormal communication identification rule.

[0116] In some embodiments, the abnormal communication identification unit 603 determines the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content and the preset abnormal communication identification rule, comprising: determining the weight of the plurality of dimensional feature labels according to the mapping relationship between the preset scene type and the weight corresponding to the plurality of dimensional feature labels; and determining the abnormal communication identification result of the communication content according to the plurality of dimensional feature labels and the weights corresponding thereto, the target risk level corresponding to the communication content and the preset abnormal communication identification rule.

[0117] In some embodiments, the communication content includes a plurality of communication contents, the plurality of communication contents respectively correspond to communication contents of the first object and a plurality of second objects, the plurality of communication contents respectively correspond to target risk levels, the first object is a potential object with abnormal communication behavior, and the second object is a potential object facing a security risk; wherein the abnormal communication identification unit 603 determines the abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and a preset abnormal communication identification rule, including: determining a risk level distribution result corresponding to the plurality of communication contents according to the target risk level corresponding to each of the plurality of communication contents; and determining the abnormal communication identification result of the communication content according to the risk level distribution result corresponding to the plurality of communication contents and the preset abnormal communication identification rule.

[0118] The abnormal communication identification device provided in the embodiment belongs to the same application concept as the abnormal communication identification method provided in the above embodiments of the present application, can execute the abnormal communication identification method provided in any of the above embodiments of the present application, and has the corresponding function modules and beneficial effects of executing the abnormal communication identification method. Technical details not described in detail in the embodiment can be referred to the specific processing content of the abnormal communication identification method provided in the above embodiments of the present application, which will not be described here.

[0119] The functions implemented by the above acquisition unit 601, risk level identification unit 602, abnormal communication identification unit 603 and extraction unit 604 can be respectively implemented by the same or different processors, and the embodiments of the present application are not limited.

[0120] It should be understood that the units in the above apparatus can be implemented in the form of processor calling software. For example, the apparatus includes a processor connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of the units of the apparatus, wherein the processor can be a general processor such as CPU or microprocessor, and the memory can be an internal memory or an external memory of the apparatus. Alternatively, the units in the apparatus can be implemented in the form of hardware circuit, and the functions of part or all of the units can be realized by the design of the hardware circuit, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is ASIC, and the functions of part or all of the units are realized by the design of the logical relationship of the elements in the circuit. For another example, in another implementation, the hardware circuit can be realized by PLD, and taking FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of part or all of the units. All the units of the above apparatus can be realized in the form of processor calling software, or realized in the form of hardware circuit, or part of them is realized in the form of processor calling software, and the remaining part is realized in the form of hardware circuit.

[0121] In the embodiments of the present application, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as CPU, microprocessor, GPU, or DSP, etc. In another implementation, the processor can realize certain functions through the logical relationship of hardware circuit, which is fixed or can be reconfigured, such as ASIC or PLD implemented hardware circuit, such as FPGA, etc. In the reconfigurable hardware circuit, the process of the processor loading configuration document to realize hardware circuit configuration can be understood as the process of the processor loading instructions to realize the functions of part or all of the units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as NPU, TPU, DPU, etc.

[0122] It can be seen that each unit in the above apparatus can be one or more processors (or processing circuits) configured to implement the above methods, such as CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0123] In addition, all or part of each unit in the above apparatus can be integrated together or can be independently implemented. In one implementation, the units are integrated together to be implemented in the form of a SOC. The SOC can include at least one processor for implementing any of the above methods or functions of the units of the apparatus, and the at least one processor can be of different types, such as including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, and the like.

[0124] Exemplary electronic device An embodiment of the present application provides an electronic device, referring to Figure 9 As shown in the figure, the electronic device includes: a memory 200 and a processor 210; The memory 200 is connected with the processor 210, and is configured to store a program. The processor 210 is configured to realize the abnormal communication identification method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0125] Specifically, the above electronic device can further include a bus, a communication interface 220, an input device 230, and an output device 240.

[0126] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are connected with each other through the bus. Wherein: The bus can include a path for transmitting information between various components of a computer system.

[0127] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or can be an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling program execution of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-use programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0128] The processor 210 can include a main processor, and can also include a baseband chip, a modem, etc.

[0129] The memory 200 stores programs for implementing the technical solutions of the present application, and can also store operating systems and other key services. Specifically, the programs can include program codes, and the program codes include computer operation instructions. More specifically, the memory 200 can include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash, and the like.

[0130] The input device 230 can include devices that receive data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, a gravity sensor, and the like.

[0131] The output device 240 can include devices that allow information to be output to a user, such as a display screen, a printer, a speaker, and the like.

[0132] The communication interface 220 can include devices of the transceiver type or the like to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), and the like.

[0133] The processor 210 executes the programs stored in the memory 200 and calls other devices, which can be used to implement each step of the method for identifying abnormal communication provided by any of the above embodiments.

[0134] The embodiments of the present application also propose a chip including a processor and a data interface, the processor reads and runs programs stored on the memory through the data interface to execute the method for identifying abnormal communication introduced in any of the above embodiments, and the specific processing process and its beneficial effects can be referred to the above embodiments of the method for identifying abnormal communication.

[0135] Exemplary computer program products and storage media In addition to the above method and device, the embodiments of the present application can also be a computer program product, which includes computer program instructions that, when executed by a processor, cause the processor to perform the steps of the method for identifying abnormal communication according to various embodiments of the present application described in any of the above embodiments.

[0136] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0137] In addition, the embodiments of the present application can also be storage media, which stores computer programs, and the computer programs are executed by processors to perform the steps of the method for identifying abnormal communication according to various embodiments of the present application described in the above embodiments of the present application.

[0138] For each of the foregoing method embodiments, in order to simply describe, it is expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0139] It should be noted that each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same and similar parts between each embodiment can be referred to. For device embodiments, since they are basically similar to method embodiments, they are described more simply, and the relevant parts refer to the part of the method embodiment.

[0140] The steps in the method of each embodiment of the present application can be adjusted, combined and reduced in sequence according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0141] The modules and sub-modules in the device and terminal of each embodiment of the present application can be combined, divided and reduced according to actual needs.

[0142] It should be understood that the disclosed terminal, device and method can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is merely a logical function division. In actual implementation, another division manner can be used. For example, a plurality of sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.

[0143] The modules or sub-modules described as separate components can or can not be physically separated, and the components of the modules or sub-modules can or can not be physical modules or sub-modules, i.e. can be located in one place or distributed on a plurality of network modules or sub-modules. Part or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0144] In addition, the functional modules or sub-modules in each embodiment of the present application can be integrated into a processing module, or each module or sub-module can exist physically, or two or more modules or sub-modules can be integrated into one module. The integrated module or sub-module can be realized in the form of hardware or software functional module or sub-module.

[0145] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0146] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software units executed by a processor, or a combination of both. The software units can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0147] Finally, it should be noted that, in this document, the term "only" is used simply to set off from one entity or action to another in order to avoid the use of the term "and / or" or the like for the sake of clarity. In no way should the term "only" be interpreted as implying that there is an implied exclusion of any referenced entity or action. Moreover, the terms "comprising", "including", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.

[0148] The above description of disclosed embodiments provides enabling teaching for making or using the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of identifying anomalous communications, the method comprising: The method comprises: obtaining communication content between a first object and a second object; the communication content is obtained by obtaining revisit data between the first object and the second object, and the revisit data is obtained by asking a large language model based on a preset revisit target; determining the target risk level corresponding to the communication content according to a preset abnormal communication knowledge base and a preset risk level classification system through an abnormal communication identification model, wherein the abnormal communication knowledge base includes a plurality of known abnormal communication types and corresponding rhetoric characteristics, and the abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on communication content samples and the abnormal communication knowledge base, wherein the multi-stage training task includes a known abnormal communication type identification training task and a non-known abnormal communication type identification training task, and the non-known abnormal communication type is an abnormal communication type outside the abnormal communication knowledge base; the communication content samples include first communication content samples and second communication content samples, the known abnormal communication type identification training task is used to train an initial abnormal communication identification model through the first communication content samples, and the non-known abnormal communication type identification training task includes: using a thought chain reasoning mechanism to infer the second communication content sample to a real abnormal communication identification result through the initial abnormal communication identification model based on the second communication content sample and the abnormal communication knowledge base, and obtaining a predicted abnormal communication identification result; determining a reward signal according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; updating the parameters of the initial abnormal communication identification model according to the reward signal until the training is completed, and obtaining the abnormal communication identification model; determining an abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and a preset abnormal communication identification rule, wherein the abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

2. The method of claim 1, wherein, The first communication content sample corresponds to a real abnormal communication identification result, and the known abnormal communication type identification training task includes: identifying the abnormal communication type of the first communication content sample through the large language model based on the abnormal communication knowledge base, obtaining a predicted abnormal communication identification result; determining an abnormal communication identification prediction loss according to the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; converging the large language model based on the abnormal communication identification prediction loss to obtain an initial abnormal communication identification model.

3. The method according to claim 1 or 2, characterized in that, The method comprises: obtaining communication content between a first object and a second object; wherein the communication content is obtained by obtaining revisit data corresponding to each of a plurality of preset revisit slots between the first object and the second object, and the plurality of revisit slots are obtained based on the revisit target, wherein the revisit target includes at least one of the identity information, communication motivation, communication theme and risk operation of the first object.

4. The method of claim 3, wherein, The method comprises: obtaining communication content between a first object and a second object; wherein the communication content is obtained by obtaining revisit data corresponding to each of a plurality of preset revisit slots between the first object and the second object, and the plurality of revisit slots are obtained based on the revisit target, wherein the revisit target includes at least one of the identity information, communication motivation, communication theme and risk operation of the first object. The method comprises: generate an abnormal communication question corresponding to each of the revisit slots based on each of the revisit slots; obtain reply content of the first object to the abnormal communication question; fill the revisit slot based on the reply content to obtain revisit data corresponding to the revisit slot.

5. The method according to claim 1 or 2, characterized in that, The method further comprises: extracting a plurality of dimensional feature labels from the communication content through the abnormal communication identification model, the plurality of dimensional feature labels including at least two of a content summary, a call topic, a risk operation, and identity information of the first object and the second object of the communication content; wherein the determining of the abnormal communication identification result according to the target risk level corresponding to the communication content and the preset abnormal communication identification rule comprises: determining the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

6. The method of claim 5, wherein, determining the abnormal communication identification result according to the plurality of dimensional feature labels, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule. determining the weight of the plurality of dimensional feature labels according to a mapping relationship between a preset scene type and the weight of the plurality of dimensional feature labels; determining the abnormal communication identification result of the communication content according to the plurality of dimensional feature labels, the respective weights corresponding thereto, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

7. The method of claim 1 or 2, wherein, The communication content includes a plurality of communication contents, each of the plurality of communication contents corresponding to a respective communication content of the first object and a plurality of the second objects, each of the plurality of communication contents corresponding to a target risk level, the first object being a potential object with abnormal communication behavior, and the second object being a potential object facing a security risk; wherein the determining of the abnormal communication identification result of the communication content according to the target risk level corresponding to the communication content and the preset abnormal communication identification rule comprises: determining a risk level distribution result corresponding to the plurality of communication contents according to the respective target risk levels corresponding to the plurality of communication contents; determining the abnormal communication identification result of the communication content according to the risk level distribution result corresponding to the plurality of communication contents and the preset abnormal communication identification rule.

8. An electronic device, comprising: comprises a memory and a processor; The memory is connected with the processor and is used for storing programs; The processor is used for realizing the method according to any one of claims 1 to 7 by running the programs in the memory.

9. A storage medium, characterized by The storage medium has a computer program stored thereon, and the computer program is run by a processor to realize the method according to any one of claims 1 to 7.

10. A computer program product, characterised in that, comprises computer program instructions, which, when run by a processor, cause the processor to realize the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Number processing method and device and storage medium

    CN117221896A

  • Question and answer pair construction method and device, electronic equipment and storage medium

    CN117688160A

  • Construction method and device of multi-component data agent

    CN119398092A

  • Intention recognition method and device, electronic equipment, storage medium and computer program product

    CN119441492A

  • Network risk assessment method and system based on multi-modal data pre-training model

    CN119814354A