Harassment call identification method and device, electronic equipment and storage medium

By initiating multiple rounds of scenario conversations and analyzing response information, and utilizing a pre-trained language model to identify harassing calls, the problem of low identification reliability in existing technologies is solved, and efficient identification of harassing calls from telephone robots is achieved.

CN116489274BActive Publication Date: 2025-11-07APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310532391.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2025-11-07
Estimated Expiration
2043-05-11

AI Technical Summary

Technical Problem

In existing technologies, the reliability of identifying nuisance calls is low, especially for call numbers that have not been marked as nuisance calls.

Method used

By initiating at least two rounds of scenario conversations, receiving the caller's response information, and based on the similarity of the response information and the response speed, using a pre-trained language model to identify harassing calls, including an autoregressive generative model with a Transformer architecture, performing natural language understanding and text matching, and determining whether the call request is a harassing call initiated by a telephone robot.

Benefits of technology

It improves the reliability and coverage of nuisance call identification, can accurately identify nuisance call requests initiated by telephone robots, reduces the false positive rate, and enhances the flexibility and intelligence of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489274B_ABST
    Figure CN116489274B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of artificial intelligence, in particular to communication security and speech recognition technology, and specifically to a nuisance call identification method and device, an electronic device and a storage medium. The specific implementation scheme is as follows: in response to a call request initiated by a calling party, at least two rounds of scenario conversations are initiated; for each round of scenario conversation initiated, response information made by the calling party to the scenario conversation is received to obtain at least two pieces of response information; and based on the at least two pieces of response information, a nuisance identification result is obtained; wherein the nuisance identification result is used to represent whether the call request is a nuisance call request initiated by a telephone robot. The present disclosure can improve the reliability of nuisance call identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to communication security and voice recognition technology, and specifically to a nuisance call identification method and device, an electronic device, and a storage medium. BACKGROUND

[0002] Currently, marketing calls are disturbing, and malicious phone harassment is increasingly prominent, seriously affecting the normal life of the people, and illegal and criminal harassment calls are more likely to cause great property losses to the people. Therefore, identifying harassment calls has a wide range of applications in social life. At present, the main method is to identify whether the calling number used by the calling party when initiating a call request has been marked as a harassment call to obtain a harassment identification result. This method has low reliability because it cannot identify calling numbers that have not been marked as harassment calls. SUMMARY

[0003] The present disclosure provides a nuisance call identification method, device, electronic device, and storage medium.

[0004] According to an aspect of the present disclosure, a nuisance call identification method is provided, comprising:

[0005] initiating at least two rounds of scenario conversations in response to a call request initiated by a calling party;

[0006] For each round of scenario conversations initiated, receiving response information made by the calling party for the scenario conversations to obtain at least two pieces of response information;

[0007] obtaining a harassment identification result based on the at least two pieces of response information; wherein the harassment identification result is used to represent whether the call request is a harassment call request initiated by a telephone robot.

[0008] According to another aspect of the present disclosure, a nuisance call identification device is provided, comprising:

[0009] a conversation initiation unit configured to initiate at least two rounds of scenario conversations in response to a call request initiated by a calling party;

[0010] a response receiving unit configured to, for each round of scenario conversations initiated, receive response information made by the calling party for the scenario conversations to obtain at least two pieces of response information;

[0011] a harassment identification unit configured to obtain a harassment identification result based on the at least two pieces of response information; wherein the harassment identification result is used to represent whether the call request is a harassment call request initiated by a telephone robot.

[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0013] at least one processor;

[0014] a memory connected with the at least one processor in communication;

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.

[0018] The present disclosure can improve the reliability of spam call identification.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0021] Figure 1 A flowchart of a spam call identification method provided by an embodiment of the present disclosure;

[0022] Figure 2 A complete flowchart of a spam call identification method provided by an embodiment of the present disclosure;

[0023] Figure 3 A scene diagram of a spam call identification method provided by an embodiment of the present disclosure;

[0024] Figure 4 A schematic structural block diagram of a spam call identification device provided by an embodiment of the present disclosure;

[0025] Figure 5 A schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are meant to be exemplary in nature, and include various details intended to facilitate understanding of the present disclosure. However, those skilled in the art will recognize that the exemplary embodiments described herein can be practiced with variation of these details, and that such variations are considered to be within the scope of the present disclosure. Thus, it is to be understood that the present disclosure is not to be limited to the specific examples described herein.

[0027] The present disclosure provides a spam call identification method, which can be applied to an electronic device. Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings. Figure 1 A flowchart of the present disclosure is shown in the flowchart, and the present disclosure provides a spam call identification method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in other orders.

[0028] Step S101, in response to the call request initiated by the caller, at least two rounds of scenario conversation are initiated;

[0029] Step S102, after initiating a round of scenario conversation, the caller's response information for the scenario conversation is received to obtain at least two pieces of response information;

[0030] Step S103, based on the at least two pieces of response information, a spam identification result is obtained; wherein the spam identification result is used to represent whether the call request is a spam call request initiated by a telephone robot.

[0031] Wherein, the call request can be a regular telephone request or a network telephone request, and the present disclosure does not limit this.

[0032] In addition, it can be understood that in the present disclosure, after receiving the call request, the call request can be intercepted first, and the step of initiating at least two rounds of scenario conversation can be performed, rather than directly generating a call prompt for reminding the host to answer the call. In a specific example, after receiving the call request, at least two rounds of scenario conversation can be initiated through a pre-trained language model, and the selection of the scenario conversation can have randomness. Wherein, the language model is pre-trained, has general language knowledge, world knowledge, and professional knowledge in various fields, and is stored in the model in the form of parameters, and has strong language interaction function. In a specific example, the language model can be a self-regressive generation model of Transformer architecture.

[0033] The scenario conversation can be a common-sense conversation. In a specific example, the scenario conversation can be a common-sense conversation related to a date or a common-sense conversation related to weather. In addition, in the embodiments of the present disclosure, the scenario conversation can be a declarative sentence, an interrogative sentence, an exclamatory sentence, or the like. For example, in the embodiments of the present disclosure, the scenario conversation can be “What day is today?”, “How is the weather today?”, “The weather is fine today”, “What a fine weather today!” or the like.

[0034] In the embodiments of the present disclosure, at least two pieces of response information are obtained by receiving response information made by the calling party for each round of scenario conversation after each round of scenario conversation is initiated. Then, the harassment identification result is obtained based on the at least two pieces of response information. For example, the harassment identification result can be obtained according to the similarity and / or response speed of the at least two pieces of response information. The harassment identification result is used to represent whether the call request is a harassment call request initiated by a telephone robot. The telephone robot can be a hardware device or an application software installed on a hardware device, which is not limited here.

[0035] In the harassment call identification method provided by the embodiments of the present disclosure, after receiving the call request, instead of directly identifying whether the call number used by the calling party to initiate the call request has been marked as a harassment call from the network side as in the prior art, the call request is first intercepted, at least two rounds of scenario conversation are initiated, and at least two pieces of response information are obtained by receiving response information made by the calling party for each round of scenario conversation. Then, the harassment identification result is obtained based on the at least two pieces of response information. On the one hand, in the embodiments of the present disclosure, whether the call number used by the calling party to initiate the call request has been marked as a harassment call or not, the harassment call identification is performed, which has a wider coverage, and thus the reliability of the harassment call identification can be improved. On the other hand, in the embodiments of the present disclosure, the harassment identification result is obtained based on the response information made by the calling party for the scenario conversation, which has a higher intelligence and more flexibility compared with the scheme of simply relying on historical marking for identification, and can accurately identify the harassment call request initiated by the telephone robot, and thus the reliability of the harassment call identification can be further improved.

[0036] As described above, in the embodiments of the present disclosure, obtaining the harassment identification result based on the at least two pieces of response information includes the following steps:

[0037] The harassment identification result is obtained according to the similarity and / or response speed of the at least two pieces of response information.

[0038] Generally, the harassing call is dialed by a telephone robot, and the telephone robot has the characteristics of poor flexibility and fast response speed. The flexibility can be represented by the similarity of the at least two response information, and specifically, the flexibility is negatively correlated with the similarity of the at least two response information. The response speed can be represented by the response speed of the at least two response information, and specifically, the response speed is positively correlated with the response speed of the at least two response information.

[0039] Based on this, in the embodiments of the present disclosure, after obtaining the at least two response information, the similarity of the at least two response information can be obtained, and then whether the calling party is a telephone robot is determined according to the similarity of the at least two response information, and the harassment identification result is obtained accordingly; or, after obtaining the at least two response information, the response speed of the at least two response information can be obtained, and then whether the calling party is a telephone robot is determined according to the response speed of the at least two response information, and the harassment identification result is obtained accordingly; or, after obtaining the at least two response information, the similarity of the at least two response information and the response speed of the at least two response information can be obtained, and then whether the calling party is a telephone robot is determined according to the similarity and the response speed of the at least two response information, and the harassment identification result is obtained accordingly.

[0040] Since the telephone robot has the characteristics of poor flexibility and fast response speed, the flexibility can be represented by the similarity of the at least two response information, and the response speed can be represented by the response speed of the at least two response information. Therefore, in the above steps, the influencing variable of the harassment identification result is located to the similarity and / or the response speed of the at least two response information, which can improve the relevance between the influencing variable and the harassment identification result, so as to further improve the reliability of the harassing call identification.

[0041] As mentioned above, in the embodiments of the present disclosure, the flexibility is negatively correlated with the similarity of the at least two response information, and the response speed is positively correlated with the response speed of the at least two response information. Based on this, in some optional embodiments, "obtaining the harassment identification result according to the similarity and / or the response speed of the at least two response information" includes one of the following steps:

[0042] In the case that the similarity of the at least two response information is greater than a similarity threshold, obtaining a harassment identification result for representing that the call request initiated by the calling party is a harassing call request;

[0043] In the case that the response speed of the at least two response information is less than a speed threshold, obtaining a harassment identification result for representing that the call request initiated by the calling party is a harassing call request;

[0044] In a case where the similarity degree of the at least two pieces of reply information is greater than the similarity threshold value and the reply speed of the at least two pieces of reply information is less than the speed threshold value, a nuisance identification result for characterizing the call request initiated by the calling party as a nuisance call request is obtained.

[0045] In the embodiments of the present disclosure, the similarity threshold value can be preset and can be set according to actual application requirements, for example, can be set to 90%, which is not limited herein. Similarly, in the embodiments of the present disclosure, the speed threshold value can be preset and can be set according to actual application requirements, for example, can be set to 200 milliseconds (ms), which is not limited herein.

[0046] In a specific example, in a case where the similarity degree of the at least two pieces of reply information is greater than the similarity threshold value, a nuisance identification result for characterizing the call request as a nuisance call request initiated by a telephone robot is obtained, that is, it is determined that the calling party is a telephone robot and that the calling number used when the calling party initiates the call request is a nuisance call; otherwise, a nuisance identification result for characterizing the call request as a non-nuisance call request is obtained, that is, it is determined that the calling party is not a telephone robot and that the calling number used when the calling party initiates the call request is not a nuisance call.

[0047] In another specific example, in a case where the reply speed of the at least two pieces of reply information is less than the speed threshold value, a nuisance identification result for characterizing the call request as a nuisance call request initiated by a telephone robot is obtained, that is, it is determined that the calling party is a telephone robot and that the calling number used when the calling party initiates the call request is a nuisance call; otherwise, a nuisance identification result for characterizing the call request as a non-nuisance call request initiated by a telephone robot is obtained, that is, it is determined that the calling party is not a telephone robot and that the calling number used when the calling party initiates the call request is not a nuisance call.

[0048] In yet another specific example, in a case where the reply speed of the at least two pieces of reply information is less than the speed threshold value and the reply speed of the at least two pieces of reply information is less than the speed threshold value, a nuisance identification result for characterizing the call request as a nuisance call request initiated by a telephone robot is obtained, that is, it is determined that the calling party is a telephone robot and that the calling number used when the calling party initiates the call request is a nuisance call; otherwise, a nuisance identification result for characterizing the call request as a non-nuisance call request initiated by a telephone robot is obtained, that is, it is determined that the calling party is not a telephone robot and that the calling number used when the calling party initiates the call request is not a nuisance call.

[0049] Through the above steps, in the embodiments of the present disclosure, the similarity threshold can be set in advance, and the similarity of the at least two pieces of response information can be used to obtain the harassment identification result; or the speed threshold can be set in advance, and the response speed of the two pieces of response information can be used to obtain the harassment identification result; or the similarity threshold and the speed threshold can be set in advance, and the similarity and the response speed of the at least two pieces of response information can be used to obtain the harassment identification result. On the one hand, the calculation logic of the harassment call identification can be simplified, and on the other hand, the flexibility of the harassment call identification method can be increased.

[0050] In some optional embodiments, the harassment call identification method further includes the following steps:

[0051] The at least two pieces of response information are identified by using the pre-trained language model to obtain at least two pieces of response identification results; wherein each piece of response identification result corresponds to one piece of response information;

[0052] The content matching result and the length matching result of the at least two pieces of response identification results are obtained;

[0053] The similarity of the at least two pieces of response information is obtained based on the content matching result and the length matching result.

[0054] As described above, in the embodiments of the present disclosure, the language model is pre-trained.

[0055] In the embodiments of the present disclosure, the natural language understanding (NLU) function of the language model can be used to identify the at least two pieces of response information to obtain at least two pieces of response identification results, and then the content matching result and the length matching result of the at least two pieces of response identification results are obtained by using the character matching algorithm, and the similarity of the at least two pieces of response information is obtained based on the content matching result and the length matching result. Among them, the content matching result is used to represent the content similarity of the at least two pieces of response identification results, and the length matching result is used to represent the length similarity of the at least two pieces of response identification results. Based on this, in the embodiments of the present disclosure, the similarity of the at least two pieces of response information can be determined by the similarity calculation logic S = AS1 + BS2. Among them, S is the similarity of the at least two pieces of response information, A is the first weight, S1 is the content matching result of the at least two pieces of response identification results, B is the second weight, and S2 is the length matching result of the at least two pieces of response identification results. Among them, the first weight and the second weight can be set according to the actual application requirements, for example, the first weight can be set to 0.6, and the second weight can be set to 0.4, which is not limited here.

[0056] In the following, the above steps included in the harassment call identification method will be described in combination with specific examples.

[0057] Assuming that a three-round scenario conversation is initiated in response to a call request initiated by a calling party, a total of three pieces of response information will be obtained, and each piece of response information corresponds to a scenario conversation.

[0058] Among them, the first round of scenario conversation is "What day is today?", After obtaining the first response information corresponding to the first round of scenario conversation, the NLU function of the language model is used to identify the first response information, and the obtained first response identification result is "We are XX Company here, our product is YY, are you interested in learning about it?"; The second round of scenario conversation is "What is the weather like today?", After obtaining the second response information corresponding to the second round of scenario conversation, the NLU function of the language model is used to identify the second response information, and the obtained second response identification result is "We are XX Company here, our product is YY, are you interested in learning about it?"; The third round of scenario conversation is "Today's weather is good.", After obtaining the third response information corresponding to the third round of scenario conversation, the NLU function of the language model is used to identify the third response information, and the obtained third response identification result is "We are XX Company here, our product is YY, are you interested in learning about it".

[0059] Thereafter, the three pieces of response identification results can be matched in content using a text matching algorithm to obtain a content matching result, and the content matching result indicates that the content similarity of the three pieces of response identification results is 100%; At the same time, the length matching of the three pieces of response identification results is carried out using the text matching algorithm to obtain a length matching result, and the length matching result indicates that the length similarity of the three pieces of response identification results is 100%.

[0060] After obtaining the content matching result and the length matching result of the three pieces of response identification result, the similarity degree of the three pieces of response information can be obtained based on the above similarity calculation logic, specifically S = A + B.

[0061] Through the above steps, in the embodiment of the disclosure, at least two pieces of response information can be identified using a pre-trained language model to obtain at least two pieces of response identification result, and the content matching result and the length matching result of the at least two pieces of response identification result are obtained, and the similarity degree of the at least two pieces of response information is obtained based on the content matching result and the length matching result. Since the language model is pre-trained, it has general language knowledge, world knowledge, and professional knowledge in various fields, etc., therefore, by using the pre-trained language model to identify at least two pieces of response information, not only can the identification efficiency of the response information be improved to improve the efficiency of the nuisance call identification, but also the identification accuracy of the response information can be improved to further improve the reliability of the nuisance call identification.

[0062] In some optional embodiments, the method for identifying the harassing call further comprises the following steps:

[0063] For each piece of response information, obtaining a response time of the response information, and an initiation time of the scenario session corresponding to the response information;

[0064] Calculating a time difference between the response time and the initiation time;

[0065] According to the time difference corresponding to each piece of response information, obtaining a response speed of at least two pieces of response information.

[0066] In a specific example, after obtaining the time difference corresponding to each piece of response information to obtain at least two time differences, the mean of the two time differences can be calculated as the response speed of at least two pieces of response information.

[0067] Continuing to combine the above example. Assuming that the time difference corresponding to the first response information is T1, the time difference corresponding to the second response information is T2, and the time difference corresponding to the third response information is T3, then the response speed of the three pieces of response information is T=(T1+T2+T3) / 3.

[0068] Through the above steps, in the embodiments of the present disclosure, for each piece of response information, the response time of the response information and the initiation time of the scenario session corresponding to the response information can be obtained, and then the time difference between the response time and the initiation time is calculated, and according to the time difference corresponding to each piece of response information, the response speed of at least two pieces of response information is obtained. Since the response speed of at least two pieces of response information is related to the time difference corresponding to each piece of response information, the reliability of the response speed can be improved to further improve the reliability of the harassing call identification.

[0069] In addition, as described above, in a specific example, after receiving the call request, at least two rounds of scenario sessions can be initiated through the pre-trained language model. Based on this, in some optional embodiments, in response to the call request initiated by the calling party, initiating at least two rounds of scenario sessions comprises the following steps:

[0070] Generating a session initiation instruction in the case of monitoring the call request initiated by the calling party;

[0071] Sending the session initiation instruction to the pre-trained language model;

[0072] Initiating at least two rounds of scenario sessions through the language model.

[0073] The language model can pre-store a plurality of candidate scenario conversations, and any two of the plurality of candidate scenario conversations have different conversation contents. In addition, the plurality of candidate scenario conversations stored in the language model can be updated regularly. For example, the plurality of candidate scenario conversations stored in the language model can be updated once every week.

[0074] After monitoring the call request initiated by the calling party, the call request can be intercepted, a session initiation instruction is generated again, and the session initiation instruction is sent to the pre-trained language model to initiate at least two rounds of scenario conversations through the language model.

[0075] In a specific example, the language model can randomly select a target scenario conversation from the plurality of pre-stored candidate scenario conversations. Thereafter, at least two rounds of scenario conversations are initiated based on the target scenario conversation. For example, a target scenario conversation randomly selected from the plurality of pre-stored candidate scenario conversations is “What day is today?”, and three rounds of scenario conversations are initiated through the language model. The first round of scenario conversation, the first round of scenario conversation, and the third round of scenario conversation can all be “What day is today?”, that is, the three rounds of scenario conversations have similar conversation contents.

[0076] In another specific example, the language model can randomly select at least two target scenario conversations from the plurality of pre-stored candidate scenario conversations. Thereafter, at least two rounds of scenario conversations are initiated based on the at least two target scenario conversations. For example, three target scenario conversations are randomly selected from the plurality of pre-stored candidate scenario conversations, which are “What day is today?”, “How is the weather today?”, and “The weather is fine today.”, and three rounds of scenario conversations are initiated through the language model. The first round of scenario conversation can be “What day is today?”, the second round of scenario conversation can be “How is the weather today?”, and the third round of scenario conversation can be “The weather is fine today.”, that is, the three rounds of scenario conversations have different conversation contents.

[0077] Through the above steps, in the embodiment of the disclosure, the session initiation instruction can be generated and sent to the pre-trained language model when the call request initiated by the calling party is monitored, and at least two rounds of scenario conversations can be initiated through the language model. Since the language model is pre-trained and has general language knowledge, world knowledge, and professional knowledge in various fields, and is stored in the model in the form of parameters and has strong language interaction function, at least two rounds of scenario conversations can be initiated through the pre-trained language model, which can reduce the development difficulty of the nuisance call identification method and reduce the development cost.

[0078] In addition, based on the above description, it can be understood that, in the embodiments of the present disclosure, the at least two rounds of scenario conversations can have similar conversation contents or can have dissimilar conversation contents.

[0079] In actual applications, for a telephone robot with a high degree of intelligence, different response information can be made for scenario conversations with dissimilar conversation contents, and the same response information can be made for scenario conversations with similar conversation contents. For example, for the scenario conversation "What is the weather like today?", the response information made by the telephone robot with a high degree of intelligence can be "It is sunny today, suitable for outing", for the scenario conversation "What is the weather like today?" with dissimilar conversation contents, the response information made by the telephone robot with a high degree of intelligence can be "It is Thursday today", and for the scenario conversation "What is the weather like today?" with similar conversation contents, the response information made by the telephone robot with a high degree of intelligence can be "It is sunny today, suitable for outing". In this way, when the harassment identification result is obtained according to the similarity degree and / or the response speed of the at least two rounds of response information, the harassment telephone call request initiated by this type of telephone robot can be misjudged. Based on this, in the embodiments of the present disclosure, the at least two rounds of scenario conversations can be preferably set to have similar conversation contents, so as to further improve the reliability of the harassment telephone call identification.

[0080] For example, "What is the weather like today?" and "What is the weather like today?" have the same semantics, and therefore have similar conversation contents. For another example, "What is the day today?" and "What is the weather like today?" have different semantics, and therefore have dissimilar conversation contents.

[0081] In some optional embodiments, "initiating at least two rounds of scenario conversations in response to the call request initiated by the calling party" includes the following steps:

[0082] initiating at least one host information question in response to the call request initiated by the calling party;

[0083] receiving reply information made by the calling party for the host information question to obtain at least one reply information each time a host information question is initiated;

[0084] obtaining the familiarity degree of the calling party and the host based on the at least one reply information;

[0085] initiating at least two rounds of scenario conversations in the case where the familiarity degree is greater than a preset familiarity degree.

[0086] It can be understood that, in the embodiments of the present disclosure, after receiving the call request, first, the call request can be intercepted, and the step of initiating at least one machine owner information question can be performed, instead of directly generating a call prompt for reminding the machine owner to answer the phone. In a specific example, after receiving the call request, at least one machine owner information question can be initiated through a pre-trained language model, and the selection of the machine owner information question can have randomness.

[0087] In a specific example, the machine owner information question can be a question related to machine owner personal information. The machine owner personal information can be the machine owner's name, gender, age, etc. Based on this, in the embodiments of the present disclosure, the machine owner information question can be "What is the machine owner's name?", "What is the machine owner's gender?", "What is the machine owner's age?", etc.

[0088] In addition, in the embodiments of the present disclosure, at least one reply information is obtained for each machine owner information question initiated by the calling party, to obtain at least one reply information. Thereafter, based on the at least one reply information, the familiarity degree between the calling party and the machine owner is obtained, and at least two rounds of scene conversations are initiated in the case that the familiarity degree is greater than a preset familiarity degree; in the case that the familiarity degree is less than or equal to the preset familiarity degree, the call request is rejected, and the calling number used by the calling party when initiating the call request is marked as a nuisance call. The familiarity degree threshold can be set according to actual application requirements, for example, it can be set to 70%, which is not limited here.

[0089] Through the above steps, in the embodiments of the present disclosure, the familiarity degree between the calling party and the machine owner can be predicted, at least two rounds of scene conversations are initiated in the case that the familiarity degree is greater than a preset familiarity degree, and in the case that the familiarity degree is less than or equal to the preset familiarity degree, the call request can be rejected, and the calling number used by the calling party when initiating the call request is marked as a nuisance call, thereby avoiding that in the case that the familiarity degree is less than or equal to the preset familiarity degree, the subsequent steps of the nuisance call identification method are also continued to be performed, so as to improve the efficiency of nuisance call identification.

[0090] In addition, as described above, in a specific example, at least one machine owner information question can be initiated through a pre-trained language model after receiving the call request. Based on this, in some optional embodiments, "initiating at least one machine owner information question in response to the call request initiated by the calling party" includes the following steps:

[0091] generating a question initiation instruction in the case that the calling party initiates the call request;

[0092] sending the question initiation instruction to the pre-trained language model;

[0093] initiating at least one machine owner information question through the language model.

[0094] The language model can pre-store a plurality of candidate host information questions, and any two of the plurality of candidate host information questions have different question contents. In addition, the plurality of host information questions stored in the language model can be updated regularly. For example, the plurality of host information questions stored in the language model can be updated once every week.

[0095] After monitoring the call request initiated by the calling party, the call request can be intercepted, a question initiation instruction is generated again, and the question initiation instruction is sent to the pre-trained language model to initiate at least one host information question through the language model.

[0096] In a specific example, the language model can randomly select at least one host information question from the plurality of pre-stored host information questions, and initiate the at least one host information question. For example, three host information questions are randomly selected from the plurality of pre-stored host information questions, which are “What is the host's name?”, “What is the host's gender?”, and “What is the host's age?”, and then the three host information questions can be issued in turn.

[0097] Through the above steps, in the embodiment of the disclosure, the question initiation instruction can be generated when the calling party initiates the call request, the question initiation instruction is sent to the pre-trained language model, and at least one host information question is initiated through the language model. Since the language model is pre-trained and has general language knowledge, world knowledge, and professional knowledge in various fields, and is stored in the model in the form of parameters and has strong language interaction function, at least one host information question can be initiated through the pre-trained language model, which can reduce the development difficulty of the nuisance call identification method and reduce the development cost.

[0098] In some optional embodiments, “obtaining the familiarity between the calling party and the host based on the at least one reply information” includes the following steps:

[0099] The pre-trained language model is used to identify the at least one reply information respectively to obtain at least one reply identification result; wherein each reply identification result corresponds to one reply information;

[0100] For each reply identification result, a reference reply corresponding to the reply identification result is determined;

[0101] The reply matching result of the reply identification result and the reference reply is obtained;

[0102] According to the reply matching result corresponding to each reply identification result, the familiarity between the calling party and the host is obtained.

[0103] As mentioned above, in the embodiments of the present disclosure, the language model is pre-trained.

[0104] In the embodiments of the present disclosure, the NLU function of the language model can be used to identify at least one piece of reply information, and obtain at least one piece of reply identification result. Thereafter, for each piece of reply identification result, a reference reply corresponding to the reply identification result is determined, and a keyword matching algorithm is used to obtain a reply matching result of the reply identification result and the reference reply, so as to represent whether the reply identification result and the reference reply are successfully matched. Wherein, for the reply identification result that can be successfully matched with the corresponding reference reply, it can be taken as a valid identification result. Finally, according to the matching result corresponding to each piece of reply identification result, the familiarity between the caller and the machine owner is obtained. In a specific example, the first total number of reply identification results and the second total number of valid identification results can be obtained, and the ratio of the second total number to the first total number is calculated as the familiarity between the caller and the machine owner.

[0105] In addition, based on the above steps, it can be understood that in the embodiments of the present disclosure, the language model stores a plurality of information question and answer pairs, each information question and answer pair includes a machine owner information question and a reference reply corresponding to the machine owner information question. The reference reply will also be taken as the reply information made by the caller to the machine owner information question, and the pre-trained language model is used to identify the reply information to obtain the reply identification result, and the reference reply corresponding to the reply identification result.

[0106] Wherein, the plurality of information question and answer pairs can be stored in the language model by pre-training, or can be stored in the language model by pre-setting, which is not limited here. In addition, in the embodiments of the present disclosure, after the plurality of information question and answer pairs are stored in the language model, they can also be synchronized to the cloud as backup data.

[0107] In the following, the above steps included in "obtaining the familiarity between the caller and the machine owner based on at least one piece of reply information" will be described in combination with specific examples.

[0108] Suppose that in response to the call request initiated by the caller, three machine owner information questions are initiated, then three pieces of reply information will be obtained in total, and each piece of reply information corresponds to a machine owner information question.

[0109] The first machine owner information question is "What is the machine owner's name?", after obtaining the first reply information corresponding to the first machine owner information question, the NLU function of the language model is used to identify the first reply information, and the obtained first reply identification result is "the machine owner's name is Zhang San"; the second machine owner information question is "What is the machine owner's gender?", after obtaining the second reply information corresponding to the second machine owner information question, the NLU function of the language model is used to identify the second reply information, and the obtained second reply identification result is "the machine owner's gender is male"; the third machine owner information question is "What is the machine owner's age?", after obtaining the third reply information corresponding to the third machine owner information question, the NLU function of the language model is used to identify the third reply information, and the obtained third reply identification result is "the machine owner's age is 28 years old".

[0110] Thereafter, the first reply identification result and the first reply matching result of the corresponding reference reply (assuming "Li Si") can be obtained by using the keyword matching algorithm, and the first reply matching result indicates that the first reply identification result and the corresponding reference reply are not successfully matched; at the same time, the second reply identification result and the second reply matching result of the corresponding reference reply (assuming "male") are obtained by using the keyword matching algorithm, and the second reply matching result indicates that the second reply identification result and the corresponding reference reply are successfully matched, so the second reply identification result is taken as the valid identification result; the third reply identification result and the third reply matching result of the corresponding reference reply (assuming "32 years old") are obtained by using the keyword matching algorithm, and the third reply matching result indicates that the third reply identification result and the corresponding reference reply are not successfully matched.

[0111] Finally, the first total number of reply identification results, which is 3, and the second total number of valid identification results, which is 1, can be obtained, and the ratio of the second total number to the first total number, which is 33%, is calculated as the familiarity degree of the calling party and the machine owner.

[0112] Through the above steps, in the embodiment of the disclosure, at least one piece of reply information can be identified by using the pre-trained language model to obtain at least one piece of reply identification result, thereafter, for each piece of reply information, a reference reply corresponding to the reply information is determined, and then the matching result of the reply information and the reference reply is obtained, and the familiarity degree of the calling party and the machine owner is obtained according to the matching result corresponding to each piece of reply information. Since the language model is pre-trained and has general language knowledge, world knowledge, and professional knowledge in various fields, etc., by using the pre-trained language model to identify at least one piece of reply information, not only can the identification efficiency of the reply information be improved to improve the efficiency of the nuisance call identification, but also the identification accuracy of the reply information can be improved to further improve the reliability of the nuisance call identification.

[0113] In some optional embodiments, the step of "initiating at least one host information query in response to the call request initiated by the caller" comprises the following steps:

[0114] In response to the call request initiated by the caller, obtaining the call number used by the caller when initiating the call request;

[0115] In the case that the call number is not stored in the host's contact list and there is no valid call record associated with the call number in the host's call record, initiating at least one host information query.

[0116] The valid call record can be a call record with a total call duration greater than a duration threshold and / or a call frequency greater than a frequency threshold. The duration threshold and the frequency threshold can be set according to actual application requirements, for example, the duration threshold can be set to 4 minutes (min) and the frequency threshold can be set to 2 times, which are not limited herein.

[0117] In the case that the call number is not stored in the host's contact list and there is no valid call record associated with the call number in the host's call record, it indicates that the caller is an untrusted caller, and therefore, the subsequent steps of the spam call identification method can be continued to execute; otherwise, it indicates that the caller is a trusted caller, and therefore, it is unnecessary to continue to execute the subsequent steps of the spam call identification method.

[0118] Through the above steps, in the embodiments of the present disclosure, the call number used by the caller when initiating the call request can be obtained in response to the call request initiated by the caller, and in the case that the call number is not stored in the host's contact list and there is no valid call record associated with the call number in the host's call record, it indicates that the caller is an untrusted caller, and therefore, the subsequent steps of the spam call identification method can be continued to execute, that is, at least one host information query is initiated, thereby ensuring the coverage rate of the spam call identification for untrusted callers, while maintaining the original call request efficiency of trusted callers, to improve the usability of the spam call identification method.

[0119] In some optional embodiments, the spam call identification method further comprises the following steps:

[0120] In the case that the spam identification result indicates that the call request is a spam call request initiated by a telephone robot, rejecting the call request;

[0121] And / or, in the case that the spam identification result indicates that the call request is not a spam call request initiated by a telephone robot, generating a call prompt for reminding the host to answer the call.

[0122] In a case where the harassment identification result represents that the call request is a harassment call request initiated by a telephone robot, the call request can be rejected, and a call number used by the caller to initiate the call request can be marked as a harassment call.

[0123] The call receiving prompt can be a vibration prompt, an audio prompt, etc., and is not limited herein.

[0124] In a case where the harassment identification result represents that the call request initiated by the caller is a harassment call request, the call request can be rejected in the method according to the embodiments of the present disclosure; and in a case where the harassment identification result represents that the call request initiated by the caller is not a harassment call request, a call receiving prompt for prompting the host to receive the call can be generated, so as to improve the automation degree of the harassment call identification method.

[0125] In the following, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Figure 2 The integrity process of the harassment call identification method according to the embodiments of the present disclosure will be described.

[0126] At least one host information question is initiated in response to the call request initiated by the caller.

[0127] For each host information question initiated, reply information made by the caller to the host information question is received, at least one piece of reply information is obtained, and the familiarity between the caller and the host is obtained based on the at least one piece of reply information.

[0128] In a case where the familiarity is greater than a preset familiarity, at least two rounds of scenario conversations are initiated; otherwise, the call request is rejected, and a call number used by the caller to initiate the call request is marked as a harassment call.

[0129] For each round of scenario conversation initiated, response information made by the caller to the scenario conversation is received, and at least two pieces of response information are obtained.

[0130] In a case where the similarity between the at least two pieces of response information is greater than a similarity threshold, and the response speed of the at least two pieces of response information is less than a speed threshold, a harassment identification result representing that the call request is a harassment call request initiated by a telephone robot is obtained, and the call request is rejected, and a call number used by the caller to initiate the call request is marked as a harassment call; otherwise, a call receiving prompt for prompting the host to receive the call is generated.

[0131] Please refer to Figure 3 The scenario schematic diagram of the harassment call identification method according to the embodiments of the present disclosure.

[0132] As described above, the method for identifying a harassing call provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of communication devices, such as a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices.

[0133] The electronic device can be used for:

[0134] initiating at least two rounds of scenario conversations in response to a call request initiated by a calling party;

[0135] receiving, for each round of scenario conversation initiated, response information made by the calling party to the scenario conversation, to obtain at least two pieces of response information;

[0136] obtaining a harassing identification result based on the at least two pieces of response information, wherein the harassing identification result is used to represent whether the call request is a harassing call request initiated by a telephone robot.

[0137] The calling party can be various forms of communication devices or a telephone robot. The telephone robot can be a hardware device or an application software installed on a hardware device, which is not limited here. The hardware device can be various forms of digital computers, such as a laptop computer, a desktop computer, a workstation, a personal digital processor, a server, a blade server, a mainframe computer, and other suitable computers. The hardware device can also represent various forms of mobile devices, such as a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices.

[0138] It should be noted that in the embodiments of the present disclosure, Figure 3 The scenario schematic diagram shown is only schematic and not limited, and those skilled in the art can make various obvious changes and / or replacements based on the examples Figure 3 The technical solutions obtained still belong to the disclosure range of the embodiments of the present disclosure.

[0139] In order to better implement the method for identifying a harassing call, the embodiments of the present disclosure further provide a device for identifying a harassing call, which can be integrated in an electronic device. Hereinafter, a device 400 for identifying a harassing call provided by the embodiments of the present disclosure will be described in conjunction with the structure schematic diagram shown. Figure 4

[0140] The device 400 for identifying a harassing call comprises:

[0141] A conversation initiation unit 401 is configured to initiate at least two rounds of scenario conversations in response to a call request initiated by a calling party;

[0142] A response receiving unit 402 is configured to, for each round of scenario conversation initiated, receive response information made by the calling party to the scenario conversation, to obtain at least two pieces of response information.​

[0143] The harassment identification unit 403 is configured to obtain a harassment identification result based on the at least two pieces of response information, wherein the harassment identification result is used to represent whether the call request is a harassment call request initiated by a telephone robot.

[0144] In some optional embodiments, the harassment identification unit 403 is configured to:

[0145] obtain the harassment identification result according to the similarity degree and / or the response speed of the at least two pieces of response information.

[0146] In some optional embodiments, the harassment identification unit 403 is configured to one of:

[0147] obtain the harassment identification result used to represent that the call request initiated by the caller is a harassment call request in a case where the similarity degree of the at least two pieces of response information is greater than a similarity threshold;

[0148] obtain the harassment identification result used to represent that the call request initiated by the caller is a harassment call request in a case where the response speed of the at least two pieces of response information is less than a speed threshold;

[0149] obtain the harassment identification result used to represent that the call request initiated by the caller is a harassment call request in a case where the similarity degree of the at least two pieces of response information is greater than a similarity threshold and the response speed of the at least two pieces of response information is less than a speed threshold.

[0150] In some optional embodiments, the harassment call identification apparatus 400 further comprises a similarity obtaining unit configured to:

[0151] identify the at least two pieces of response information respectively by using a pre-trained language model to obtain at least two pieces of response identification results, wherein each piece of response identification result corresponds to one piece of response information;

[0152] obtain a content matching result and a length matching result of the at least two pieces of response identification results;

[0153] obtain the similarity degree of the at least two pieces of response information based on the content matching result and the length matching result.

[0154] In some optional embodiments, the harassment call identification apparatus 400 further comprises a response speed obtaining unit configured to:

[0155] obtain, for each piece of response information, a response time of the response information and an initiation time of a scene conversation corresponding to the response information;

[0156] calculate a time difference between the response time and the initiation time;

[0157] According to the time difference corresponding to each piece of response information, the response speed of at least two pieces of response information is obtained.

[0158] In some optional embodiments, the session initiation unit 401 is configured to:

[0159] generate a session initiation instruction in response to the call request initiated by the caller;

[0160] send the session initiation instruction to the pre-trained language model;

[0161] initiate at least two rounds of scenario conversations through the language model.

[0162] In some optional embodiments, the at least two rounds of scenario conversations have similar conversation contents.

[0163] In some optional embodiments, the session initiation unit 401 is configured to:

[0164] initiate at least one piece of host information question in response to the call request initiated by the caller;

[0165] receive reply information made by the caller in response to the at least one piece of host information question to obtain at least one piece of reply information;

[0166] obtain the familiarity between the caller and the host based on the at least one piece of reply information;

[0167] initiate at least two rounds of scenario conversations in response to the familiarity being greater than a preset familiarity.

[0168] In some optional embodiments, the session initiation unit 401 is configured to:

[0169] identify the at least one piece of reply information respectively through the pre-trained language model to obtain at least one piece of reply identification result, wherein each piece of reply identification result corresponds to one piece of reply information;

[0170] determine a reference reply corresponding to each piece of reply identification result;

[0171] obtain a reply matching result of the reply identification result and the reference reply;

[0172] obtain the familiarity between the caller and the host according to the reply matching result corresponding to each piece of reply identification result.

[0173] In some optional embodiments, the session initiation unit 401 is configured to:

[0174] obtain a call number used by the caller when initiating the call request in response to the call request initiated by the caller;

[0175] In a case where the call number is not stored in the contact list of the owner and there is no valid call record related to the call number in the call record of the owner, at least one piece of owner information query is initiated.

[0176] In some optional embodiments, the spam call identification apparatus 400 further comprises a call control unit, configured to:

[0177] In a case where the spam identification result indicates that the call request is a spam call request initiated by a telephone robot, the call request is rejected.

[0178] In a case where the spam identification result indicates that the call request is not a spam call request initiated by a telephone robot, a call prompt for prompting the owner to answer the call is generated.

[0179] The specific functions and examples of the units of the spam call identification apparatus 400 according to the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here.

[0180] In the technical solutions of the present disclosure, the acquisition, storage and application of the user personal information (for example, the owner personal information) involved all comply with the relevant legal regulations and do not violate public order and good customs.

[0181] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0182] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0183] As Figure 5As shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a Read-Only Memory (ROM) 502 or a computer program loaded from a storage unit 508 into a Random Access Memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.

[0184] Various components in the device 500 are connected to the I / O interface 505, including an input unit 506 such as a keyboard, a mouse, etc., an output unit 507 such as various types of displays, speakers, etc., a storage unit 508 such as a magnetic disk, an optical disk, etc., and a communication unit 509 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0185] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various special-purpose Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Process (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, e.g., the spam call identification method. For example, in some embodiments, the spam call identification method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, e.g., the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the spam call identification method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the spam call identification method by any other appropriate means, e.g., by means of firmware.

[0186] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a system on chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0187] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as part of a separate software package, and partially on a remote machine or server.

[0188] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electronic storage, a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) monitor or a Liquid Crystal Display (LCD)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0190] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0191] The computer system can include clients and servers. This relationship can be

[0192] The embodiments of the present disclosure further provide a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to make the computer execute the method for identifying harassment calls.

[0193] The embodiments of the present disclosure further provide a computer program product, comprising a computer program which, when executed by a processor, implements the method for identifying harassment calls.

[0194] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein. In addition, in the present disclosure, relationship terms such as "first", "second", "third" and the like are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. In addition, in the present disclosure, "a plurality of" can be understood as at least two.

[0195] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A method for identifying a spam call, comprising: initiating at least two rounds of scenario conversations in response to a call request initiated by a caller; receiving response information made by the caller for the scenario conversations to obtain at least two pieces of response information, each time a round of the scenario conversations is initiated; obtaining a spam identification result based on the at least two pieces of response information, wherein the spam identification result is used to represent whether the call request is a spam call request initiated by a telephone robot; wherein the obtaining of the spam identification result based on the at least two pieces of response information comprises obtaining the spam identification result according to a similarity degree and / or a response speed of the at least two pieces of response information.

2. The method of claim 1, wherein, The obtaining of the spam identification result according to the similarity degree and / or the response speed of the at least two pieces of response information comprises one of the following: in a case where the similarity degree of the at least two pieces of response information is greater than a similarity threshold, obtaining a spam identification result used to represent that the call request initiated by the caller is a spam call request; in a case where the response speed of the at least two pieces of response information is less than a speed threshold, obtaining a spam identification result used to represent that the call request initiated by the caller is a spam call request; in a case where the similarity degree of the at least two pieces of response information is greater than the similarity threshold and the response speed of the at least two pieces of response information is less than the speed threshold, obtaining a spam identification result used to represent that the call request initiated by the caller is a spam call request.

3. The method of claim 1, further comprising: identifying the at least two pieces of response information respectively by using a pre-trained language model to obtain at least two pieces of response identification result, wherein each of the response identification results corresponds to one of the response information; obtaining a content matching result and a length matching result of the at least two pieces of response identification result; obtaining the similarity degree of the at least two pieces of response information based on the content matching result and the length matching result.

4. The method of claim 1, further comprising: for each of the response information, obtaining a response time of the response information and an initiation time of a scenario conversation corresponding to the response information; calculating a time difference value between the response time and the initiation time; obtaining a response speed of the at least two pieces of response information according to the time difference value corresponding to each of the response information.

5. The method of claim 1, wherein, The initiating of the at least two rounds of scenario conversations in response to the call request initiated by the caller comprises: generating a conversation initiation instruction in a case where a call request initiated by a caller is monitored; sending the conversation initiation instruction to a pre-trained language model; initiating the at least two rounds of scenario conversations by the language model.

6. The method according to any one of claims 1 to 5, wherein, The at least two rounds of scenario conversations have similar conversation contents.

7. The method according to any one of claims 1 to 5, wherein, The initiating of the at least two rounds of scenario conversations in response to the call request initiated by the caller comprises: initiating at least one machine information question in response to the call request initiated by the caller; receiving reply information made by the caller for the machine information question to obtain at least one piece of reply information, each time a round of the machine information question is initiated; obtain familiarity between the caller and the machine owner based on the at least one reply information; in a case where the familiarity is greater than a preset familiarity, initiate the at least two rounds of scenario conversations.

8. The method of claim 7, wherein, The obtaining of the familiarity between the caller and the machine owner based on the at least one reply information comprises: identifying the at least one piece of reply information respectively by using a pre-trained language model to obtain at least one reply identification result; wherein each reply identification result corresponds to one piece of reply information; for each reply identification result, determining a reference reply corresponding to the reply identification result; obtaining a reply matching result of the reply identification result and the reference reply; obtaining the familiarity between the caller and the machine owner according to the reply matching result corresponding to each reply identification result.

9. The method of claim 7, wherein, The initiation of the at least one machine owner information query in response to the call request initiated by the caller comprises: in response to the call request initiated by the caller, obtaining a call number used by the caller when initiating the call request; in a case where the call number is not stored in the machine owner's address book and there is no valid call record related to the call number in the machine owner's call record, initiating the at least one machine owner information query.

10. The method of claim 1, further comprising: in a case where the harassment identification result represents that the call request is a harassment call request initiated by a telephone robot, rejecting the call request; and / or, in a case where the harassment identification result represents that the call request is not a harassment call request initiated by a telephone robot, generating a call prompt for reminding the machine owner to answer the call.

11. A harassment call identification device, comprising: a conversation initiation unit configured to initiate at least two rounds of scenario conversations in response to a call request initiated by a caller; a reply receiving unit configured to receive reply information made by the caller for the scenario conversations to obtain at least two pieces of reply information after each round of the scenario conversations is initiated; a harassment identification unit configured to obtain a harassment identification result based on the at least two pieces of reply information; wherein the harassment identification result is used to represent whether the call request is a harassment call request initiated by a telephone robot; wherein the harassment identification unit is configured to obtain the harassment identification result according to the similarity and / or reply speed of the at least two pieces of reply information.

12. The apparatus of claim 11, wherein, The harassment identification unit is configured to perform one of the following: in a case where the similarity of the at least two pieces of reply information is greater than a similarity threshold, obtain a harassment identification result representing that the call request initiated by the caller is a harassment call request; in a case where the reply speed of the at least two pieces of reply information is less than a speed threshold, obtain a harassment identification result representing that the call request initiated by the caller is a harassment call request; in a case where the similarity of the at least two pieces of reply information is greater than a similarity threshold and the reply speed of the at least two pieces of reply information is less than a speed threshold, obtain a harassment identification result representing that the call request initiated by the caller is a harassment call request.

13. The apparatus of claim 11, further comprising a similarity obtaining unit, configured to: obtain, for each of the response identification results, a piece of the response information; obtain a content matching result and a length matching result of the at least two response identification results; and obtain a similarity degree of the at least two pieces of response information based on the content matching result and the length matching result. The pre-trained language model is used to respectively identify the at least two pieces of response information, and at least two pieces of response identification results are obtained.

14. The apparatus of claim 11, further comprising a response speed obtaining unit, configured to: obtain, for each of the response information, a response time of the response information and an initiation time of a scenario session corresponding to the response information; calculate a time difference value between the response time and the initiation time; and obtain a response speed of the at least two pieces of response information according to the time difference value corresponding to each of the response information. The session initiation unit is configured to: in a case where a calling party initiates a call request, generate a session initiation instruction; send the session initiation instruction to a pre-trained language model; and initiate the at least two rounds of scenario sessions through the language model. The at least two rounds of scenario sessions have similar session contents. The session initiation unit is configured to: in response to the call request initiated by the calling party, initiate at least one host information question; receive reply information made by the calling party for the host information question to obtain at least one piece of reply information for each host information question initiated; obtain a familiarity degree of the calling party with a host based on the at least one piece of reply information; and in a case where the familiarity degree is greater than a preset familiarity degree, initiate the at least two rounds of scenario sessions. The session initiation unit is configured to: identify the at least one piece of reply information respectively by using the pre-trained language model to obtain at least one reply identification result, wherein each of the reply identification results corresponds to a piece of the reply information; determine, for each of the reply identification results, a reference reply corresponding to the reply identification result; obtain a reply matching result of the reply identification result and the reference reply; and obtain the familiarity degree of the calling party with the host according to the reply matching result corresponding to each of the reply identification results. The session initiation unit is configured to: in response to the call request initiated by the calling party, obtain a calling number used by the calling party when initiating the call request; and in a case where the calling number is not stored in a contact list of the host and there is no valid call record related to the calling number in a call record of the host, initiate the at least one host information question.

20. The apparatus of claim 11, further comprising a call control unit, configured to: in a case where the harassment identification result indicates that the call request is a harassment call request initiated by a telephone robot, reject the call request; and / or in a case where the harassment identification result indicates that the call request is not a harassment call request initiated by a telephone robot, generate a call receiving prompt for reminding a host to receive a call.

15. The apparatus of claim 11, wherein, 21. An electronic device, comprising: at least one processor; a memory connected to the at least one processor in communication; ​ ​ ​ 16. The apparatus of any one of claims 11-15, wherein, ​ 17. The apparatus of any one of claims 11-15, wherein, ​ ​ ​ ​ ​ 18. The apparatus of claim 17, wherein, ​ ​ ​ ​ ​ 19. The apparatus of claim 17, wherein, ​ ​ ​ ​ ​ ​ ​ ​ ​ The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

22. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method of any one of claims 1-10.

23. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Incoming call processing method and mobile terminal

    CN106534535A