A method, device, medium, and product for identifying abnormal communication.

By training a large language model in multiple stages, combining an abnormal communication knowledge base and a risk level classification system, and utilizing the thought chain reasoning mechanism and reinforcement learning, the problem of low accuracy in identifying new or variant abnormal communication in existing technologies has been solved, achieving accurate identification and generalization capabilities for abnormal communication.

CN120979906BActive Publication Date: 2026-04-03IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify new or variant abnormal communication behaviors, and suffer from low accuracy in handling complex and ever-changing user expressions and contextual relationships.

Method used

By training a large language model with known and unknown anomalous communication types, an anomalous communication identification model is established. By utilizing an anomalous communication knowledge base and risk level classification system, combined with thought chain reasoning mechanism and reinforcement learning, the model's ability to identify anomalous communication is improved.

Benefits of technology

It achieves accurate identification of known abnormal communication types and has the ability to generalize the identification of new or variant abnormal communication types, thereby improving the accuracy and adaptability of abnormal communication identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979906B_ABST
    Figure CN120979906B_ABST
Patent Text Reader

Abstract

This application provides a method, device, medium, and product for identifying abnormal communication. The method uses an abnormal communication identification model based on an abnormal communication knowledge base and a risk level classification system to determine the target risk level of the communication content between a first object and a second object. The abnormal communication identification model is obtained based on training tasks for identifying known and unknown abnormal communication types. The training task for identifying unknown abnormal communication types includes: using an initial abnormal communication identification model obtained from a training task for identifying known abnormal communication types; reasoning based on an abnormal communication knowledge base and a thought chain reasoning mechanism to obtain a predicted abnormal communication identification result; determining a reward signal and updating model parameters based on the difference between the predicted result and the actual abnormal communication identification result to obtain the abnormal communication identification model; and determining the abnormal communication identification result based on the target risk level and abnormal communication identification rules. This improves the accuracy of abnormal communication identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, device, medium and product for identifying abnormal communication. Background Technology

[0002] With the development of mobile communication technology and voice interaction technology, the number of voice calls received by user terminals has increased significantly.

[0003] Amidst a large volume of communication traffic, numerous unsolicited communication activities exist, including marketing promotions, harassing calls, and unusual communication activities with potential security risks. These communications often induce users to respond by simulating legitimate service providers or employing social engineering techniques, potentially leading to security incidents such as information leaks and financial losses.

[0004] Therefore, there is an urgent need for a highly accurate risk communication behavior identification mechanism to effectively identify abnormal communication behaviors and thus improve the security protection mechanism of communication systems. Summary of the Invention

[0005] Based on the above-mentioned technological status, this application provides a method, device, medium, and product for identifying abnormal communication, which can improve the accuracy of abnormal communication identification, especially the accuracy of identifying new or variant types of abnormal communication.

[0006] To achieve the above-mentioned technical objectives, this application proposes the following technical solution:

[0007] According to a first aspect of the embodiments of this application, a method for identifying abnormal communication is provided, comprising: acquiring communication content between a first object and a second object; the communication content being obtained by acquiring follow-up data between the first object and the second object, the follow-up data being obtained by a large language model asking questions based on a preset follow-up target; determining a target risk level corresponding to the communication content by an abnormal communication identification model based on a preset abnormal communication knowledge base and a preset risk level classification system, wherein the abnormal communication knowledge base includes multiple known abnormal communication types and corresponding verbal features, the abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on communication content samples and the abnormal communication knowledge base, the multi-stage training task including a training task for identifying known abnormal communication types and a training task for identifying unknown abnormal communication types, wherein the unknown abnormal communication types are abnormal communication types outside the abnormal communication knowledge base; the communication content samples include a first object and a second object ... The training task for identifying known abnormal communication types involves using a first communication content sample and a second communication content sample. An initial abnormal communication identification model is trained using the first communication content sample. The training task for identifying unknown abnormal communication types includes: using the initial abnormal communication identification model, based on the second communication content sample and the abnormal communication knowledge base, employing a thought chain reasoning mechanism to infer from the second communication content sample to the actual abnormal communication identification result, obtaining a predicted abnormal communication identification result; determining a reward signal based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result; updating the parameters of the initial abnormal communication identification model based on the reward signal until training ends, obtaining the abnormal communication identification model; and determining the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and preset abnormal communication identification rules, wherein the abnormal communication identification result indicates whether the communication content is abnormal communication.

[0008] In some implementations, the first communication content sample corresponds to a real abnormal communication identification result, and the training task for the known abnormal communication type identification includes: using the large language model to identify the known abnormal communication type of the first content sample based on the abnormal communication knowledge base, and obtaining a predicted abnormal communication identification result; determining the prediction loss for abnormal communication identification based on the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; and converging the large language model based on the prediction loss for abnormal communication identification to obtain an initial abnormal communication identification model.

[0009] In some implementations, obtaining the communication content between the first object and the second object includes: obtaining the callback data corresponding to each preset callback slot to obtain the communication content between the first object and the second object; wherein, each callback slot is set based on the callback target, and the callback target includes at least one of the first object's identity information, communication motivation, communication topic, and risky operation.

[0010] In some implementations, obtaining the callback data corresponding to each preset callback slot includes: generating an abnormal communication question corresponding to each callback slot based on each callback slot; obtaining the response content of the first object to the abnormal communication question; and filling the callback slot based on the response content to obtain the callback data corresponding to the callback slot.

[0011] In some implementations, the method further includes: extracting multi-dimensional feature tags from the communication content using the abnormal communication identification model, wherein the multi-dimensional feature tags include a content summary of the communication content, a call topic, a risky operation, and at least two of the identity information of the first object and the second object; wherein, determining the abnormal communication identification result based on the target risk level corresponding to the communication content and a preset abnormal communication identification rule includes: determining the abnormal communication identification result based on the multi-dimensional feature tags, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0012] In some implementations, determining the abnormal communication identification result based on the feature labels of the multiple dimensions, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules includes: determining the weights of the feature labels of the multiple dimensions based on the mapping relationship between the preset scenario type and the weights corresponding to the feature labels of the multiple dimensions; and determining the abnormal communication identification result of the communication content based on the feature labels of the multiple dimensions and their respective weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules.

[0013] In some implementations, the communication content includes multiple communication contents, which are communication contents corresponding to the first object and multiple second objects respectively. Each of the multiple communication contents corresponds to a target risk level. The first object is an object potentially exhibiting abnormal communication behavior, and the second objects are objects potentially facing security risks. The step of determining the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and a preset abnormal communication identification rule includes: determining the risk level distribution result corresponding to the multiple communication contents based on the target risk level corresponding to each of the multiple communication contents; and determining the abnormal communication identification result of the communication content based on the risk level distribution result corresponding to the multiple communication contents and the preset abnormal communication identification rule.

[0014] According to a second aspect of the embodiments of this application, an electronic device is provided, including a memory and a processor; the memory is connected to the processor and is used to store a program; the processor is used to implement the abnormal communication identification method as described in the first aspect by running the program in the memory.

[0015] According to a third aspect of the embodiments of this application, a storage medium is provided, on which a computer program is stored, and when the computer program is run by a processor, it implements the abnormal communication identification method as described in the first aspect.

[0016] According to a fourth aspect of the embodiments of this application, a computer program product is provided, including computer program instructions, which, when executed by a processor, cause the processor to perform: the abnormal communication identification method as described in the first aspect.

[0017] This application provides a method, device, medium, and product for identifying abnormal communication. The method involves acquiring communication content between a first object and a second object. This communication content is obtained by acquiring follow-up data between the first and second objects. The follow-up data is obtained by a large language model asking questions based on a preset follow-up target. An abnormal communication identification model determines the target risk level corresponding to the communication content based on a preset abnormal communication knowledge base and a preset risk level classification system. The abnormal communication knowledge base includes multiple known abnormal communication types and their corresponding verbal characteristics. The abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on communication content samples and the abnormal communication knowledge base. The multi-stage training task includes training tasks for identifying known abnormal communication types and training tasks for identifying unknown abnormal communication types. Unknown abnormal communication types are those outside the abnormal communication knowledge base. The samples include a first communication content sample and a second communication content sample. The training task for identifying known abnormal communication types is trained using the first communication content sample to obtain an initial abnormal communication identification model. The training task for identifying unknown abnormal communication types includes: using the initial abnormal communication identification model, based on the second communication content sample and the abnormal communication knowledge base, and employing a thought chain reasoning mechanism to infer from the second communication content sample to the actual abnormal communication identification result, obtaining a predicted abnormal communication identification result; determining a reward signal based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result; updating the parameters of the initial abnormal communication identification model based on the reward signal until training ends, obtaining the abnormal communication identification model; and determining the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and the preset abnormal communication identification rules. The abnormal communication identification result is used to indicate whether the communication content is abnormal communication. Because the abnormal communication identification model is trained on two stages—one for identifying known abnormal communication types and the other for identifying unknown abnormal communication types—the training task for identifying known abnormal communication types enables the model to establish a mapping relationship between communication content samples and real labels. Therefore, it can accurately identify known abnormal communication types in the abnormal communication knowledge base and possesses preliminary generalization ability, i.e., a preliminary ability to capture abnormal information, such as analogous generalization of known patterns. In the training task for identifying unknown abnormal communication types, the training objective is adjusted to modeling based on the reasoning process. By introducing mechanisms such as reinforcement learning or thought chains, and using reward signals to guide the large language model to generate the reasoning steps of the aforementioned mapping relationship, the generalization ability of the large model to capture abnormal behavior is further improved, thereby enhancing the model's ability to generalize and identify new fraud methods and rapidly evolving variants of fraud. Therefore, it can not only accurately identify known abnormal communication types but also effectively identify unknown abnormal communication types outside the abnormal communication knowledge base based on the thinking ability about abnormal communication types learned from known abnormal communication types. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the architecture of an abnormal call analysis system in related technologies.

[0020] Figure 2 This is a schematic diagram illustrating the operating mechanism of a return visit system in related technologies.

[0021] Figure 3 A flowchart illustrating an abnormal communication identification method provided in an embodiment of this application.

[0022] Figure 4 This is a schematic diagram illustrating an abnormal communication identification method provided in an embodiment of this application.

[0023] Figure 5 This is a schematic diagram illustrating an abnormal communication identification result for determining communication content, provided in an embodiment of this application.

[0024] Figure 6 This is a schematic diagram of the structure of the return visit system provided in the embodiments of this application.

[0025] Figure 7 A flowchart illustrating the acquisition of communication content provided in an embodiment of this application.

[0026] Figure 8 This is a schematic diagram of the structure of an abnormal communication identification device provided in an embodiment of this application.

[0027] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions provided in this application can be applied, by way of example, to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged as software programs and run. When the hardware device executes the processing procedure of the technical solutions in this application, or when the aforementioned software program is run, the target task can be automatically split and the application programming interfaces required by the task can be automatically invoked to achieve the purpose of the target task. This application only provides illustrative descriptions of the specific processing procedure of the technical solutions in this application and does not limit the specific implementation form of the technical solutions in this application. Any technical implementation form that can execute the processing procedure of the technical solutions in this application can be adopted by this application.

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] Before introducing the solution proposed in this application, the relevant technologies will first be introduced:

[0031] In recent years, the use of telephones to make abnormal calls has become increasingly common. Accurately identifying and shutting down these illegitimate numbers has become one of the important means of combating such calls.

[0032] Based on the characteristic of abnormal calls frequently targeting unknown numbers, communication behavior analysis can be used to initially screen and identify numbers potentially involved in improper behavior. Building on this, follow-up technology is introduced to revisit potentially affected users, obtain their call details with the numbers involved in improper calls, and analyze the content of the follow-up conversations using text semantic understanding technology. This allows for a comprehensive assessment of whether the numbers involved in improper calls constitute abnormal call behavior.

[0033] Figure 1 This is a schematic diagram of the architecture of the abnormal call analysis process in related technologies. For example... Figure 1 As shown, the abnormal call analysis process includes: collecting telephone communication content between numbers involved in improper behavior (i.e., potential abnormal numbers) and potentially affected users (i.e., potential victims) through a callback system, and generating callback call results; the analysis system then conducts a comprehensive analysis based on these callback call results to determine whether the numbers involved in improper behavior belong to abnormal call behavior or other legitimate identities, such as deliverymen or ride-hailing drivers.

[0034] Figure 2 This is a schematic diagram illustrating the operational mechanism of a callback system in related technologies. For example... Figure 2 As shown, traditional follow-up systems typically ensure the effectiveness and efficiency of information collection by setting structured questions.

[0035] Specifically, the first step is to design a question template library with clearly defined options or open-ended questions based on the follow-up goals. For example, the question could ask whether the user recognizes a fraudulent number (i.e., a potentially suspicious number), with options including "recognize," "don't recognize," and "other." When the user selects "don't recognize," the call content is further investigated to determine which pre-defined option it falls under. Next, the call content is analyzed to identify any risky actions and a fixed option for such actions is provided. These question templates aim to guide potentially affected users to provide key information. Then, based on the users' responses to different questions, a tree-like dialogue control logic structure is organized to dynamically adjust subsequent questioning paths, ensuring the comprehensiveness and accuracy of information collection. Finally, understanding user responses relies on various natural language processing technologies, such as Named Entity Recognition (NER), syntactic analysis, intent recognition, and semantic slot filling, to extract the core content and key demands of user feedback.

[0036] Existing systems for assessing improper conduct are deeply integrated with the follow-up call system. For a single follow-up call, the system needs to extract key information about the abnormal communication based on the follow-up dialogue results, including the user's identity, the call's topic, and any risky actions taken. It then uses manually defined fusion rules to determine the abnormal risk level of that follow-up call. For example, if the communication involves educational training and the risky action is proactively offering a refund, the call's abnormal risk level is high. When multiple follow-up calls are assessed as high-risk, the number involved in the improper conduct will be identified as an abnormal caller and will be suspended.

[0037] While existing assessment methods for misconduct are effective to some extent, they still have the following shortcomings:

[0038] In real-world conversations, users have diverse expression habits, including word choice, spoken language, regional accents, and transcription errors. These factors can lead to misunderstandings in the follow-up system's understanding of user responses, affecting the quality of the follow-up.

[0039] Secondly, to support the identification of diverse fraudulent methods, it is necessary to maintain complex structured question templates, corresponding option templates, and corresponding transition relationships. However, as the number of intents and the complexity of their expression increase, designing, writing, testing, maintaining, and managing a large number of grammatical rules and template libraries becomes extremely cumbersome and time-consuming. Adding new intents or modifying the rules involved in existing intents often requires extensive adjustments and testing, resulting in poor scalability and increased development and maintenance costs.

[0040] Furthermore, existing follow-up systems struggle to address contextual relationships and semantic reasoning issues in follow-up dialogues, making it difficult to accurately detect abnormal behavior and easily overlooking even known abnormal methods.

[0041] Furthermore, existing methods for identifying inappropriate behavior rely on pre-designed question and option templates within the follow-up system, which limits their ability to identify novel anomalous tactics. Due to the limitations of the follow-up system and the complexity of user expressions, existing solutions cannot fully reveal anomalous behaviors exposed during follow-up conversations.

[0042] In summary, while current abnormal call analysis systems possess certain functionalities, they still face numerous challenges in dealing with complex and ever-changing user expressions, maintaining complex structured question templates, and handling contextual relationships, resulting in low accuracy in identifying abnormal calls.

[0043] In view of this, the embodiments of this application aim to provide an abnormal communication method, device, medium, and product. By training a large language model with tasks for identifying known abnormal communication types and those for identifying unknown abnormal communication types, an abnormal communication identification model is obtained. This model can not only accurately identify known abnormal communication types, but also, based on common patterns and features learned from known abnormal communication types, effectively identify unknown abnormal communication types with similar abnormal characteristics through reasoning and judgment. Detailed descriptions are provided in the following embodiments.

[0044] Exemplary methods

[0045] Figure 3 A flowchart illustrating an abnormal communication identification method provided in an embodiment of this application. Figure 3 As shown, the abnormal communication identification method provided in this embodiment includes steps S101-S103:

[0046] S101. Obtain the communication content between the first object and the second object.

[0047] The first category comprises individuals with potential abnormal communication behavior, specifically those whose preliminary analysis indicates a possible predisposition to such behavior. These abnormal behaviors include patterns such as frequently calling unknown numbers and making calls to multiple different regions within a short period. These behaviors often suggest that the individual may be engaging in illicit activities, such as telephone fraud. Since there is no concrete evidence, these individuals are tentatively referred to as potential individuals with abnormal communication behavior.

[0048] The second object is an entity that may pose a security risk. The second object can be identified based on interactions with the first object. For example, if a number is flagged for frequently receiving calls, text messages, or instant messages from an entity exhibiting unusual communication behavior, that number would be considered a potential security risk. This indicates that the second object may be the target or victim of unusual communication behavior.

[0049] In some embodiments, the communication content between the first object and the second object can be directly obtained, or a follow-up technique can be used to follow up with the second object to understand the specific background, topic and intention of the communication between the two parties, thereby obtaining the communication content between the first object and the second object by obtaining the follow-up data between the first object and the second object.

[0050] To improve the quality of communication content, this follow-up data can also be obtained by asking questions based on preset follow-up goals using a large language model. The follow-up goals are determined by setting a series of objectives to guide the large language model in asking questions to obtain the information needed for the follow-up goals, without limiting the specific question format, i.e., how to ask the questions. The specific implementation process of obtaining communication content through follow-up data will be described in detail in subsequent embodiments.

[0051] The communication content includes not only actual audio conversations but also potentially text messages, such as SMS or instant messages. In some examples, the communication content may be used in telephone fraud.

[0052] It should be noted that obtaining communication content between the first and second parties must comply with laws and regulations, and ensure that privacy protection measures are in place. Specifically, relevant call records and voice data can be obtained from telecommunications operators through legally authorized means. Furthermore, appropriate encryption and protection measures should be taken when obtaining communication content to ensure data security and privacy. At the same time, all operations must comply with local data protection regulations and privacy policies to ensure that the legitimate rights and interests of users are not infringed.

[0053] S102. Based on the preset abnormal communication knowledge base and the preset risk level classification system, the abnormal communication identification model determines the target risk level corresponding to the communication content.

[0054] The abnormal communication knowledge base includes multiple known abnormal communication types and their corresponding script characteristics. Examples include common abnormal communication types such as fake order scams, loan fraud, impersonation of customer service personnel, romance scams, and fake investment schemes, along with typical script characteristics and behavioral patterns for each type.

[0055] The risk level classification system defines different risk levels, including the first risk level, the second risk level, the third risk level, and the fourth risk level, which can help the model classify the risk level of communication content.

[0056] The first risk level can be high-risk, explicitly mentioning or implementing core fraudulent elements, such as requesting bank card passwords, verification codes, demanding transfers to so-called "safe accounts," inducing the download of high-risk applications (APPs), and conforming to specific fraudulent scripts. The second risk level can be medium-risk, containing highly suspicious leading language, possibly in the early stages of fraudulent inducement (such as obtaining identity information or building trust), but without directly requesting transfers or revealing core passwords. The third risk level can be low-risk, containing some sensitive words or slightly leading behaviors, but the overall intent is unclear or has not yet posed a direct threat. The fourth risk level can be no risk, indicating that the communication content is normal and has no suspicious or abnormal communication characteristics.

[0057] Figure 4 This is a schematic diagram illustrating an abnormal communication identification method provided in an embodiment of this application. Figure 4 As shown, by revisiting multiple second objects associated with potentially abnormal numbers, multiple communication contents can be obtained. Then, the communication contents, a pre-defined abnormal communication knowledge base, and a pre-defined risk level classification system are input into the abnormal communication identification model in the form of prompt words, enabling the model to identify the risks associated with the communication contents and obtain the target risk level corresponding to each content.

[0058] The abnormal communication identification model can be obtained by training a large language model in multiple stages based on communication content samples, a pre-set abnormal communication knowledge base, and a pre-set risk level classification system. These multi-stage training tasks include training for identifying known abnormal communication types and training for identifying unknown abnormal communication types. Unknown abnormal communication types are those outside the abnormal communication knowledge base, i.e., those not included in the knowledge base.

[0059] Specifically, the communication content samples include a first communication content sample and a second communication content sample. The training task for identifying known abnormal communication types can be trained using the first communication content sample to obtain an initial abnormal communication identification model.

[0060] The first communication content sample and the second communication content sample may partially overlap, be completely different, or be completely identical. This embodiment does not impose any limitations on this.

[0061] In this embodiment, the large language model can first be trained to identify known abnormal communication types based on the first communication content sample and the preset abnormal communication knowledge base to obtain an initial abnormal communication identification model; and then the large language model can be trained to identify unknown abnormal communication types based on the second communication content sample and the abnormal communication knowledge base to obtain an abnormal communication identification model.

[0062] The training tasks for the two phases are described in detail below:

[0063] In some embodiments, the first communication content sample corresponds to a real abnormal communication identification result. The training task for identifying known abnormal communication types includes: using a large language model to identify known abnormal communication types of the first communication content sample based on an abnormal communication knowledge base, and obtaining a predicted abnormal communication identification result; determining the prediction loss for abnormal communication identification based on the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; and converging the large language model based on the prediction loss for abnormal communication identification to obtain an initial abnormal communication identification model.

[0064] Specifically, the first communication content sample can be obtained through public datasets, web scraping, or direct collection from real-world scenarios. This sample can be acquired by obtaining the actual communication content between the first and second objects. This actual communication content includes call recordings, text messages, or instant messages. Alternatively, the actual communication content between the first and second objects can be obtained by revisiting the second object.

[0065] The first communication content sample corresponds to the labeled real abnormal communication identification result, which indicates the abnormal communication type and / or risk level corresponding to the first communication content sample, thereby providing basic labels for model training and ensuring the accuracy of subsequent evaluation.

[0066] During training, a pre-defined knowledge base of abnormal communications and a pre-defined risk level classification system can be input into the large language model via prompt words. The abnormal communication knowledge base helps the large language model learn how to identify such abnormal communications. The risk level classification system helps the model to more accurately classify the risks of identified abnormal communications.

[0067] Simultaneously, the first communication content sample is input into the large language model, which identifies the risk level classification corresponding to the first communication content sample and whether it belongs to abnormal communication based on multiple known abnormal communication types in the abnormal communication knowledge base, thereby generating a predicted abnormal communication identification result.

[0068] To ensure the accuracy and robustness of the model, the predicted abnormal communication identification results need to be compared with the actual abnormal communication identification results. The difference between the two is calculated to determine the abnormal communication prediction loss. This abnormal communication prediction loss is then used as a feedback signal, and a supervised fine-tuning (SFT) strategy is employed to iteratively optimize the large language model based on the loss value, gradually adjusting the model parameters until the predetermined performance indicators are achieved.

[0069] After the first phase of training and optimization of the large language model, an initial abnormal communication identification model can be obtained. This model can not only accurately identify various known abnormal communication types, but also accurately classify their risk levels, and has a preliminary ability to capture abnormal information related to improper behavior in the call content.

[0070] To further enhance the model's ability to identify anomalous communication behaviors in complex real-world scenarios and improve its generalization and adaptability to other unknown anomalous communication types, such as novel improper communication behaviors and rapidly evolving variants, this embodiment, based on the completion of the first-stage training task, can introduce a second-stage training task. This aims to break the model's dependence on known anomalous communication patterns and enable it to cope with constantly evolving novel improper communication methods, including semantic spoofing, cross-scenario fusion, and phrasal variation—all highly concealed "variant" behaviors. The second-stage training task will be described in detail below:

[0071] In some embodiments, an initial abnormal communication identification model can be obtained through a training task for identifying known abnormal communication types. The training task for identifying unknown abnormal communication types includes: using the initial abnormal communication fraud identification model, based on a second communication content sample and an abnormal communication knowledge base, and employing a thought chain reasoning mechanism to infer from the second communication content sample to the actual abnormal communication identification result, thereby obtaining a predicted abnormal communication identification result; determining a reward signal based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result; and updating the parameters of the initial abnormal communication identification model based on the reward signal until training is completed, thereby obtaining the abnormal communication identification model.

[0072] The abnormal communication identification model has the ability to identify known abnormal communication types as well as other unknown abnormal communication types (such as new abnormal communication types or variant abnormal communication types).

[0073] During the second phase of training, the model, based on the knowledge system and contextual understanding learned in the first phase, and combined with its built-in abnormal communication knowledge base, employs a Chain-of-Thought Reasoning (CoT) mechanism to perform logical reasoning on the input second communication content sample. By analyzing semantic cues, behavioral sequences, and potential intentions within the second communication content sample, it outputs a predicted abnormal communication identification result for that sample. Response strategies include: suggesting account suspension, restricting communication permissions, initiating manual review processes, triggering secondary authentication, suggesting follow-up confirmation, strengthening monitoring, or sending security alerts.

[0074] Subsequently, reward signals can be determined based on the difference between the predicted abnormal communication identification results and the actual abnormal communication identification results. In some examples, reward signals include: a positive reward for successfully identifying non-known high-risk behaviors and proposing effective handling suggestions; a negative reward for missed detections (such as classifying genuine fraud as risk-free) or misjudgments (such as misclassifying normal services as high-risk); and a neutral or slightly positive reward for proposing conservative but reasonable prudent strategies for ambiguous scenarios (such as "suggesting a follow-up confirmation" or "transferring to manual review").

[0075] Furthermore, the reward signal is used as the objective function, and the model parameters are updated through backpropagation using the policy gradient method. With the dynamic expansion of training data and the continuous accumulation of feedback signals, the model's sensitivity and discrimination ability against atypical and highly concealed abnormal behaviors in complex semantic spaces can be improved, gradually mastering the sensitivity and judgment of atypical abnormal behaviors. When the model's comprehensive performance indicators (such as risk identification accuracy, policy adoption rate, and false alarm rate control) on the independent validation set tend to stabilize and reach a preset threshold, the training process ends, resulting in an abnormal communication identification model with strong generalization ability. This model not only has the ability to accurately identify known abnormal communication patterns, but also possesses the ability to generalize the identification, contextual reasoning, and intelligent decision-making capabilities for unknown improper communication behaviors and their rapidly evolving variants.

[0076] In summary, during the first training phase, by learning from numerous real-world examples, the model establishes a mapping between communication content samples and real labels, thus enabling accurate identification of known anomalous communication types in the knowledge base. In the second training phase, the training objective shifts from label prediction to modeling based on the reasoning process. By introducing mechanisms such as reinforcement learning or thought chains, reward signals guide the large language model in generating the aforementioned mapping relationship through reasoning steps. Under this mechanism, the large language model is no longer limited to memorizing static mappings between samples and labels but learns the inherent reasoning ability to identify anomalous communication types—that is, "how to think to arrive at the correct judgment." This ability allows it to abstract essential patterns from known anomalous patterns, enabling it to identify the deep connections between new or variant anomalous communication types and known anomalous communication types through logical deduction when faced with unfamiliar ones, thereby making accurate identification judgments and achieving accurate identification of new or variant anomalous communication types.

[0077] S103. Based on the target risk level corresponding to the communication content and the preset abnormal communication identification rules, determine the abnormal communication identification result of the communication content. The abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

[0078] Continue reading Figure 4In some embodiments, to more accurately determine the abnormality of communication content, the method of this embodiment further includes: extracting multi-dimensional feature tags from the communication content using an abnormal communication identification model; and determining the abnormal communication identification result based on the multi-dimensional feature tags, the target risk level corresponding to the communication content, and preset abnormal communication identification rules. The multi-dimensional feature tags include at least two of the following: a content summary of the communication content, the call topic, risky operations, and the identity information of the first and second objects.

[0079] The content summary of the communication content represents a general outline of the communication; identity information includes operators, couriers, food delivery workers, ride-hailing drivers, hotels, dating platforms, tour guides, travel agencies, teachers, classmates, etc.; call topics include flight delays, loan issues, notifications and reminders, travel-related topics, insurance-related topics, education and training-related topics, operator services, etc.; risky actions include downloading apps, requesting verification codes, proactively requesting refunds, transferring / recharging money, logging in and clicking on linked websites, etc.

[0080] In some embodiments, the abnormal communication identification result is determined based on feature labels of multiple dimensions, the target risk level corresponding to the communication content, and preset abnormal communication identification rules, including: determining the weights of feature labels of multiple dimensions; and determining the abnormal communication identification result of the communication content based on feature labels of multiple dimensions and their respective weights, the target risk level corresponding to the communication content, and preset abnormal communication identification rules.

[0081] Since the importance of feature tags across different dimensions varies across different business scenarios, different weights can be configured for feature tags across multiple dimensions for different scenario types to reflect the degree of attention paid to each feature tag in different scenarios. Based on this, the abnormal communication identification result is determined according to the feature tags across multiple dimensions, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules. This includes: determining the weights of the feature tags across multiple dimensions based on the mapping relationship between scenario type and the weights corresponding to the feature tags across each dimension; and determining the abnormal communication identification result of the communication content based on the feature tags across multiple dimensions and their respective weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules.

[0082] The mapping relationship between scene type and the weights corresponding to feature labels of each dimension can be shown in Table 1 below:

[0083] Table 1. Mapping relationship between scene type and the weights of feature labels for each dimension

[0084]

[0085] The following example, using the "express delivery service" scenario, illustrates how weight configuration affects the final abnormal communication judgment result.

[0086] Suppose the first contact is an employee of a courier company (identified as "courier"), who makes 10 calls to multiple second contacts (recipients) in one day. By performing risk identification on the content of each call, the following results can be obtained:

[0087] The summaries of the 9 calls were "delivery notification" and "package signed for," containing no sensitive words, and were identified as "risk-free."

[0088] One call mentioned that "compensation can be applied for if the package is lost," but there was no prompting to take any action; therefore, it was identified as "medium risk."

[0089] All call topics were either "delivery notification" or "delivery follow-up".

[0090] No clearly risky operational behaviors (such as requesting verification codes, guiding money transfers, or providing unofficial links);

[0091] The first individual's identity information is "courier".

[0092] Although this communication behavior falls under the "express delivery service" scenario, given the frequent occurrence of abnormal communication behaviors that use the guise of express delivery to commit fraud in recent years, such as impersonating customer service for claims processing or inducing refund operations, the corresponding weights for this scenario are set as follows:

[0093] Risk level: 0.4 (High weighting; although express delivery is a routine business, it is easy to use tactics such as compensation and refunds to induce customers to engage in manipulative practices.)

[0094] Call topic: 0.3 (Medium weight, topic should be related to express delivery)

[0095] Content Summary: 0.2 (lower weight, allowing some sensitive words, such as "compensation")

[0096] Identity information: 0.1 (low weight, identity can be forged or account can be stolen, the "courier" label alone is not enough to fully trust)

[0097] Therefore, the overall score can be calculated as follows:

[0098] Identity information score: 1.0 (legal courier) × 0.1 = 0.1

[0099] Call topic score: 0.9 (all related to express delivery) × 0.3 = 0.27

[0100] Content summary score: 0.8 (1 medium-risk, the rest no-risk) × 0.2 = 0.16

[0101] Risk operation score: 1.0 (no risk operation) × 0.4 = 0.4

[0102] Overall risk score = 0.1 + 0.27 + 0.16 + 0.4 = 0.93 (the higher the score, the safer the risk).

[0103] Judgment based on preset rules:

[0104] The rule states: "If the first party is identified as a 'courier' and the call topic is 'delivery notification,' even if there is a small amount of medium-risk content summary, as long as no risky actions are triggered, a high-risk warning will not be triggered."

[0105] Therefore, the abnormal communication identification result is: the communication behavior of this number conforms to the characteristics of normal business and is regarded as legitimate and compliant business communication.

[0106] For example, in financial business scenarios, even if the identity information is reliable, if the content summary involves "investment rebates" and risky operations are frequent, it can still trigger an alert for abnormal communication.

[0107] This embodiment can flexibly adjust the weight of each feature tag according to different business scenarios, avoiding a "one-size-fits-all" approach to risk assessment and improving the accuracy and adaptability of identification. Especially in high-frequency, low-risk communication scenarios such as express delivery and customer service, reasonably configuring high weights for identity information and call topics can effectively reduce false alarm rates and ensure smooth normal business communication.

[0108] It should be noted that the table above only provides the weight configuration of feature labels for each dimension in some scenarios, and the weight configuration coefficients are illustrative examples. Those skilled in the art can set reasonable weight coefficients according to the actual business scenario requirements. This embodiment does not impose specific limitations on this.

[0109] Continue reading Figure 4 And refer to Figure 5 In some embodiments, since in actual communication scenarios, a first object typically communicates with multiple second objects, the first object may communicate with each second object once or multiple times over a period of time, forming corresponding communication content. When there are multiple communication contents between the first object and the second objects, they can be integrated to generate a unique comprehensive communication content between the first object and each second object. Thus, N corresponding communication contents will be formed between the first object and N second objects. The integration process includes methods such as text aggregation, time window merging, or session merging.

[0110] For each communication content between the first object and each second object, a corresponding target risk level can be obtained. Based on this, step S103 includes: determining the risk level distribution result corresponding to the multiple communication contents based on the target risk levels corresponding to each of the multiple communication contents; and determining the abnormal communication identification result of the communication contents based on the risk level distribution result corresponding to the multiple communication contents and the preset abnormal communication identification rules.

[0111] In this embodiment, firstly, based on the target risk level of each of the N communication contents, a risk level distribution result is statistically generated. For example, this might involve statistically analyzing the proportion of high-risk communication content, the concentration of medium- and high-risk communication objects, or plotting risk level histograms. This distribution result reflects the overall risk profile of the first object's communication behavior.

[0112] Subsequently, the risk level distribution results are matched with preset abnormal communication identification rules. For example, preset rules may include: "If the proportion of high-risk communication content exceeds 30%, then abnormal communication behavior is determined to exist"; or "If the communication content with three or more different second objects is identified as high-risk, and the communication content contains the same sensitive keywords, then it is determined to be batch abnormal communication", etc.

[0113] Ultimately, based on the matching results, abnormal communication identification results will be generated, such as outputting "abnormal communication behavior exists" or "communication behavior is normal", and further triggering subsequent processing measures such as alarms, manual review or communication blocking.

[0114] Continue reading Figure 5 In some embodiments, since the judgment rules differ for different regions, the target risk level can be adjusted after the abnormal communication identification model outputs the target risk level. For example, the target risk level can be adjusted through manual intervention (see the bold and black text in the figure), and a risk level distribution result can be generated based on the adjusted target risk level.

[0115] As mentioned earlier, the communication content between the first and second objects can be obtained through follow-up technology. In this case, the completeness and richness of the information contained in the follow-up dialogue directly affect the quality of the follow-up dialogue, and thus affect the effect of subsequent comprehensive judgment. In order to fully explore the key information in the communication process and improve the information density and interaction quality of the follow-up dialogue, in some embodiments, when obtaining the communication content between the first and second objects in step S101, it specifically includes: obtaining the follow-up data corresponding to each preset follow-up slot to obtain the communication content between the first and second objects; wherein, each follow-up slot is set based on the follow-up target, and the follow-up target includes at least one of the following: the identity information of the second object, the communication motivation, the communication topic, and the risky operation.

[0116] Figure 6 This is a schematic diagram of the structure of the follow-up system provided in an embodiment of this application. Figure 6 As shown, the follow-up system includes a predefined slot module 201, a target status tracking and management module 202, a dialogue engine module 203, a slot parsing and filling module 204, and a follow-up questioning and process control logic module 205.

[0117] The slot predefined module 201 provides the function of setting the follow-up target. Users can predefine the complete set of information to be collected for this follow-up visit through the slot predefined module 201, thereby setting the follow-up target. This follow-up target includes a series of structured slots, each slot representing a key data item to be acquired, such as the identity information of the first target, communication motivation, communication topic, risky operations, etc. Furthermore, each slot can be configured with data type and necessity level. Data types include text, numbers, dates, or option lists, etc. Necessity level includes required or optional.

[0118] The target status tracking and management module 202 is used to maintain the status of the current callback session in real time. It dynamically tracks the fill status (filled / unfilled) of all preset slots, the obtained slot values, and whether the slot values ​​have passed validation (such as format validation and range validation). This manager is the control center of the entire system.

[0119] Figure 7 A flowchart illustrating the acquisition of communication content provided in an embodiment of this application. For example... Figure 7 As shown, the process of obtaining the feedback data corresponding to each preset feedback slot includes the following steps S701-S703:

[0120] S701. Based on each return slot in each return slot, generate an abnormal communication query corresponding to that return slot.

[0121] This step involves collecting information from each callback slot sequentially to obtain key contextual information for identifying abnormal communication behavior. See also... Figure 6 After filling the information in the i-th feedback slot, the dialogue engine module 203, based on preset scheduling rules and the current filling status of each feedback slot, selects the next feedback slot from the unfilled slots as the (i+1)-th feedback slot. Then, for the selected (i+1)-th feedback slot, the large language model is used to dynamically generate a natural, fluent, and semantically accurate question, combined with the context of the current dialogue. This question accurately focuses on the information requirement corresponding to the feedback slot, guiding the recipient to provide effective feedback related to abnormal communication behavior.

[0122] The preset rules include the priority order of each callback slot and the logical dependencies between them. The priority order of each callback slot ensures that key information is obtained first; while the logical dependencies between them ensure that the questioning process follows a reasonable reasoning path and business logic, thereby improving the efficiency and completeness of information collection.

[0123] S702. Obtain the response content of the first object to the abnormal communication query.

[0124] Continue reading Figure 6 This step can be achieved through the slot parsing and filling module 204. This module can leverage the powerful natural language understanding capabilities of the large language model to perform deep semantic analysis on the response content of the first object, thereby identifying key information related to the revisited target.

[0125] S703. Fill the follow-up slot based on the reply content to obtain the follow-up data corresponding to the follow-up slot.

[0126] Furthermore, the identified key information can be structured through the slot parsing and filling module 204, and then filled into the corresponding return slots to complete the automatic information filling.

[0127] To assess the completeness and effectiveness of the extracted information, the slot parsing and filling module 204 can also perform basic slot value verification on the filling results. This basic slot value verification includes the following states:

[0128] Successful filling: Key information is complete and semantically clear, and has been accurately filled into the corresponding slots;

[0129] Partial completion: Only some valid information was extracted, and further questions are needed to complete the content;

[0130] Invalid slot value: The identified information does not conform to the preset semantic range or logical rationality and is regarded as invalid input;

[0131] Unrecognized: The response does not contain the information required for the target slot, or the information is too vague to be parsed.

[0132] The aforementioned filling status and verification results will be fed back to the target status tracking and management module 202 in real time to update the overall progress status of the follow-up dialogue, guide the adjustment of subsequent dialogue strategies and slot scheduling decisions, thereby ensuring the continuity of the follow-up process and the systematic nature of information collection.

[0133] Furthermore, if the target status tracking and management module 202 receives a "slot value invalid" or "unrecognized" status from the slot parsing and filling module 204, and the currently processed return slot is a required field, then the follow-up questioning and process control logic module 205 will trigger the dialogue engine module 203 to generate targeted follow-up questions. These follow-up questions aim to guide the second party to provide valid information again through semantic clarification, using more explicit expressions, or adjusting the questioning angle, thereby improving the success rate of obtaining key information.

[0134] This mechanism allows for the preset of flexible follow-up strategies, such as limiting the maximum number of follow-ups to the same slot (e.g., no more than 3 times), to balance information collection efficiency and user experience. If effective filling is completed within the specified number of times, the target status tracking and management module 202 will update the status of the slot to "filled" and simultaneously maintain the overall follow-up progress.

[0135] Subsequently, based on the updated target state and combined with preset scheduling rules (such as slot priority, logical dependency, etc.), the dialogue engine module 203 selects the next target slot to be collected from the remaining unfilled slots and generates the corresponding question statement, thus entering a new round of information collection process, that is, repeating steps S501 to S503.

[0136] This cyclical information collection process continues to iterate until any of the following termination conditions are met: all required feedback slots have been successfully filled with valid values; the feedback target has reached the preset completeness standard (e.g., all required fields are completed, and the filling ratio of optional fields reaches a set threshold); the user actively expresses the intention to end the session (e.g., explicitly replying "I will not continue" or remaining unresponsive for a long time); or the system determines, based on preset rules, that the current information is sufficient to support subsequent comprehensive analysis.

[0137] When any termination condition is met, a closing statement will be triggered to terminate this follow-up process.

[0138] This embodiment sets the follow-up target based on at least one of the first object's identity information, communication motivation, communication topic, and risky operation, and constructs various follow-up slots. Each follow-up slot corresponds to an information dimension to be collected, and multiple follow-up slots constitute an information collection framework to ensure comprehensive coverage of key information dimensions.

[0139] Based on this, by leveraging the powerful semantic understanding and natural language generation capabilities of large language models, and guided by the goal of follow-up visits, questions are dynamically generated in a semantically driven manner, without being limited by preset fixed question templates and limited options. It can adaptively construct natural, coherent and semantically accurate question content according to the actual dialogue context.

[0140] Furthermore, the wording and expression of questions can be flexibly adjusted according to the language style, spoken expression habits, and even dialect characteristics of the second respondent to enhance the comprehensibility of the dialogue and thus more effectively guide the first respondent to provide truthful and detailed feedback.

[0141] Therefore, this embodiment can realize the transformation of the follow-up process from "passive response" to "active exploration". It can not only actively track missing information and dynamically adjust the questioning path, but also accurately identify and extract key semantic content in complex expressions, thereby significantly improving the completeness, accuracy and richness of information collection, and providing high-quality data support for subsequent comprehensive risk assessment.

[0142] In summary, the embodiments of this application achieve dynamic management of the follow-up target through a large language model. Leveraging the generalization ability of the large model and its advantages in semantic understanding and context modeling, it significantly enhances the ability to understand the intent behind user responses. Furthermore, by combining the current progress status of the follow-up target, it adaptively generates subsequent questions, achieving intelligent advancement of the dialogue process. Compared to traditional follow-up systems based on fixed templates and explicit state transitions, this embodiment effectively reduces the manual configuration and maintenance costs of question templates, response options, and state transition logic. Simultaneously, due to the large language model's deep understanding of natural language, it possesses stronger semantic coverage and scenario adaptability, significantly improving the quality of the follow-up dialogue.

[0143] Furthermore, by introducing an abnormal communication identification model, which undergoes a two-stage training task, it not only possesses the ability to accurately identify known abnormal communication types, but also enhances its ability to generalize and assess the risks of unknown new abnormal communication types and rapidly evolving variant abnormal communication types.

[0144] Furthermore, during the analysis process, the model simultaneously extracts feature labels from multiple dimensions to form a structured risk feature vector. Combined with flexibly configurable pre-defined abnormal communication identification rules, it allows for flexible handling of abnormal communication numbers. Specifically, the stringency of the analysis logic can be dynamically adjusted according to business scenario requirements. For example, in specific high-risk scenarios (such as "refund-related calls in the education and training industry"), a strong rule strategy can be configured: once a relevant topic is identified and sensitive language is present, a high-risk judgment is triggered, and automatic shutdown is executed. This achieves refined and scenario-based management of risk prevention strategies, significantly improving the response capability and handling efficiency against new and variant forms of fraud while ensuring normal business communication.

[0145] Exemplary device

[0146] Corresponding to the above-mentioned method for identifying abnormal communication, this application also provides an apparatus for identifying abnormal communication. Figure 8This is a schematic diagram of the structure of an abnormal communication identification device provided in an embodiment of this application. Figure 8 As shown, the abnormal communication identification device provided in this application embodiment includes: an acquisition unit 601, a risk level identification unit 602, and an abnormal communication identification unit 603; wherein, the acquisition unit 601 is used to acquire the communication content between a first object and a second object; the communication content is obtained by acquiring the return visit data between the first object and the second object, and the return visit data is obtained by asking questions based on a preset return visit target using a large language model; the risk level identification unit 602 is used to determine the target risk level corresponding to the communication content by using an abnormal communication identification model according to a preset abnormal communication knowledge base and a preset risk level classification system, wherein the abnormal communication knowledge base includes multiple known abnormal communication types and corresponding speech features, and the abnormal communication identification model is obtained by performing a multi-stage training task on the large language model based on communication content samples and the abnormal communication knowledge base, wherein the multi-stage training task includes a training task for identifying known abnormal communication types and a training task for identifying unknown abnormal communication types, and the unknown abnormal communication types are those outside the abnormal communication knowledge base. The communication content samples include a first communication content sample and a second communication content sample. The training task for identifying known abnormal communication types is trained using the first communication content sample to obtain an initial abnormal communication identification model. The training task for identifying unknown abnormal communication types includes: using the initial abnormal communication identification model, based on the second communication content sample and the abnormal communication knowledge base, and employing a thought chain reasoning mechanism to infer from the second communication content sample to the actual abnormal communication identification result, to obtain a predicted abnormal communication identification result; determining a reward signal based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result; updating the parameters of the initial abnormal communication identification model based on the reward signal until training ends, to obtain the abnormal communication identification model; and an abnormal communication identification unit 603 is used to determine the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and preset abnormal communication identification rules. The abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

[0147] In some embodiments, the first communication content sample corresponds to a real abnormal communication identification result, and the training task for the known abnormal communication type identification includes: using the large language model to identify the known abnormal communication type of the first content sample based on the abnormal communication knowledge base, and obtaining a predicted abnormal communication identification result; determining the prediction loss of abnormal communication identification based on the difference between the predicted abnormal communication identification result and the real abnormal communication identification result; and converging the large language model based on the prediction loss of abnormal communication identification to obtain an initial abnormal communication identification model.

[0148] In some embodiments, when the acquisition unit 601 acquires the communication content between the first object and the second object, it specifically includes: acquiring the return data corresponding to each preset return slot to obtain the communication content between the first object and the second object; wherein, each return slot is set based on the return target, and the return target includes at least one of the first object's identity information, communication motivation, communication topic, and risky operation.

[0149] In some embodiments, obtaining the callback data corresponding to each preset callback slot includes: generating an abnormal communication question corresponding to each callback slot based on each callback slot; obtaining the response content of the first object to the abnormal communication question; and filling the callback slot based on the response content to obtain the callback data corresponding to the callback slot.

[0150] In some embodiments, the apparatus further includes: an extraction unit 604, configured to extract multi-dimensional feature tags from the communication content through the abnormal communication identification model, wherein the multi-dimensional feature tags include a content summary of the communication content, a call topic, a risky operation, and at least two of the identity information of the first object and the second object; wherein the abnormal communication identification unit 603 determines the abnormal communication identification result according to the target risk level corresponding to the communication content and a preset abnormal communication identification rule, including: determining the abnormal communication identification result according to the multi-dimensional feature tags, the target risk level corresponding to the communication content, and the preset abnormal communication identification rule.

[0151] In some embodiments, the abnormal communication identification unit 603 determines the abnormal communication identification result based on the feature labels of the multiple dimensions, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules, including: determining the weights of the feature labels of the multiple dimensions based on the mapping relationship between the preset scene type and the weights corresponding to the feature labels of the multiple dimensions; and determining the abnormal communication identification result of the communication content based on the feature labels of the multiple dimensions and their respective weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules.

[0152] In some embodiments, the communication content includes multiple communication contents, which are communication contents corresponding to the first object and multiple second objects respectively. Each of the multiple communication contents corresponds to a target risk level. The first object is an object potentially exhibiting abnormal communication behavior, and the second objects are objects potentially facing security risks. The abnormal communication identification unit 603 determines the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and a preset abnormal communication identification rule, including: determining the risk level distribution result corresponding to the multiple communication contents based on the target risk level corresponding to each of the multiple communication contents; and determining the abnormal communication identification result of the communication content based on the risk level distribution result corresponding to the multiple communication contents and the preset abnormal communication identification rule.

[0153] The abnormal communication identification device provided in this embodiment belongs to the same application concept as the abnormal communication identification method provided in the above embodiments of this application. It can execute the abnormal communication identification method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the abnormal communication identification method. Technical details not described in detail in this embodiment can be found in the specific processing content of the abnormal communication identification method provided in the above embodiments of this application, and will not be repeated here.

[0154] The functions implemented by the acquisition unit 601, risk level identification unit 602, abnormal communication identification unit 603 and extraction unit 604 can be implemented by the same or different processors, and this application embodiment does not limit them.

[0155] It should be understood that the units in the above device can be implemented by a processor calling software. For example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. By designing the hardware circuits, some or all of the unit functions can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files to implement the functions of some or all of the above units. All units in the above device can be implemented entirely by a processor calling software, entirely by hardware circuits, or partially by a processor calling software with the remaining parts implemented by hardware circuits.

[0156] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0157] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0158] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0159] Exemplary electronic devices

[0160] This application provides an electronic device, see [link to relevant documentation] Figure 9 As shown, the electronic device includes:

[0161] Memory 200 and processor 210;

[0162] The memory 200 is connected to the processor 210 and is used to store programs;

[0163] The processor 210 is configured to implement the abnormal communication identification method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0164] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0165] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:

[0166] A bus can include a pathway for transmitting information between various components of a computer system.

[0167] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0168] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.

[0169] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0170] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0171] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0172] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0173] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of any of the abnormal communication identification methods provided in the above embodiments of this application.

[0174] This application also proposes a chip including a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the abnormal communication identification method described in any of the above embodiments. For the specific processing procedure and its beneficial effects, please refer to the embodiments of the abnormal communication identification method described above.

[0175] Exemplary computer program products and storage media

[0176] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the abnormal communication identification methods according to various embodiments of this application as described in any of the foregoing embodiments of this specification.

[0177] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0178] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor through steps in the abnormal communication identification method according to various embodiments of this application described in any of the foregoing embodiments of this specification.

[0179] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0180] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0181] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0182] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0183] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0184] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0185] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0186] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0187] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0188] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0189] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying abnormal communication, characterized in that, include: Retrieve the communication content between the first object and the second object; The communication content is obtained by acquiring the return visit data between the first object and the second object. The return visit data is obtained by asking questions based on a preset return visit target through a large language model. The return visit target includes at least one of the first object's identity information, communication motivation, communication topic, and risky operation, which is used to guide the large language model to ask questions to obtain the information required for the return visit target. The abnormal communication identification model determines the target risk level corresponding to the communication content based on a preset abnormal communication knowledge base and a preset risk level classification system. The abnormal communication knowledge base includes multiple known abnormal communication types and corresponding linguistic features. The abnormal communication identification model is obtained by performing a multi-stage training task on a large language model based on communication content samples and the abnormal communication knowledge base. The multi-stage training task includes a training task for identifying known abnormal communication types and a training task for identifying unknown abnormal communication types. The unknown abnormal communication types are abnormal communication types outside the abnormal communication knowledge base. The communication content samples include a first communication content sample and a second communication content sample. The training task for identifying known abnormal communication types is trained using the first communication content sample to obtain an initial abnormal communication identification model. The training task for identifying unknown abnormal communication types includes: Based on the second communication content sample and the abnormal communication knowledge base, the initial abnormal communication identification model uses a thought chain reasoning mechanism to reason from the second communication content sample to the actual abnormal communication identification result, thereby obtaining the predicted abnormal communication identification result. A reward signal is determined based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result; The parameters of the initial abnormal communication identification model are updated according to the reward signal until the training ends, thus obtaining the abnormal communication identification model. Based on the target risk level corresponding to the communication content and the preset abnormal communication identification rules, the abnormal communication identification result of the communication content is determined, and the abnormal communication identification result is used to indicate whether the communication content is abnormal communication.

2. The method according to claim 1, characterized in that, The first communication content sample corresponds to a real abnormal communication identification result, and the training task for identifying the known abnormal communication type includes: The large language model identifies the abnormal communication type of the first communication content sample based on the abnormal communication knowledge base, and the predicted abnormal communication identification result is obtained. The prediction loss for abnormal communication identification is determined based on the difference between the predicted abnormal communication identification result and the actual abnormal communication identification result. Based on the prediction loss of the abnormal communication identification, the large language model is converged to obtain the initial abnormal communication identification model.

3. The method according to claim 1 or 2, characterized in that, The step of obtaining the communication content between the first object and the second object includes: Obtain the corresponding callback data for each preset callback slot to obtain the communication content between the first object and the second object; The various return visit slots are set based on the return visit target.

4. The method according to claim 3, characterized in that, The step of obtaining the feedback data corresponding to each preset feedback slot includes: Based on each of the aforementioned return slots, generate an abnormal communication question corresponding to that return slot. Obtain the response content of the first object to the question about the abnormal communication; The response content is used to fill the response slot to obtain the response data corresponding to the response slot.

5. The method according to claim 1 or 2, characterized in that, The method further includes: The abnormal communication identification model extracts multi-dimensional feature tags from the communication content. The multi-dimensional feature tags include a content summary of the communication content, the call topic, the risky operation, and at least two of the identity information of the first object and the second object. The step of determining the abnormal communication identification result based on the target risk level corresponding to the communication content and the preset abnormal communication identification rules includes: The abnormal communication identification result is determined based on the feature labels of the multiple dimensions, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules.

6. The method according to claim 5, characterized in that, The step of determining the abnormal communication identification result based on the feature labels of the multiple dimensions, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules includes: The weights of the feature labels in the multiple dimensions are determined based on the mapping relationship between the preset scene type and the weights corresponding to the feature labels in the multiple dimensions. Based on the feature labels of the multiple dimensions and their corresponding weights, the target risk level corresponding to the communication content, and the preset abnormal communication identification rules, the abnormal communication identification result of the communication content is determined.

7. The method according to claim 1 or 2, characterized in that, The communication content includes multiple communication contents, which are the communication contents corresponding to the first object and multiple second objects respectively. The multiple communication contents are respectively corresponding to target risk levels. The first object is an object that may have abnormal communication behavior, and the second object is an object that may face security risks. The step of determining the abnormal communication identification result of the communication content based on the target risk level corresponding to the communication content and the preset abnormal communication identification rules includes: Based on the target risk level corresponding to each of the multiple communication contents, the risk level distribution result corresponding to the multiple communication contents is determined; Based on the risk level distribution results corresponding to the multiple communication contents and the preset abnormal communication identification rules, the abnormal communication identification result of the communication contents is determined.

8. An electronic device, characterized in that, Including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method as described in any one of claims 1 to 7 by running a program in the memory.

9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, It includes computer program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Number processing method and device and storage medium

    CN117221896A

  • Question and answer model training method oriented to specific field and intelligent question and answer method and device

    CN120654813A