Call warning method, device, server, storage medium and product

By analyzing the audio data and voiceprint characteristics of the call, determining the hidden danger level of the call and conducting targeted early warnings, the problem of low warning effectiveness caused by users ignoring incoming call marks in the prior art is solved, and the effectiveness of the early warning is improved.

CN114125153BActive Publication Date: 2025-05-16SOUNDAI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111305193.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-05-16
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

When prior art warns a call, users often ignore the incoming call number marking, resulting in low warning effectiveness.

Method used

By acquiring the first audio data of the call, the first semantic information and the first voiceprint characteristics are determined, the hidden danger level of the call is determined based on these information, and targeted early warning is made according to the hidden danger level.

Benefits of technology

It improves the effectiveness of early warning of calls that cause hidden dangers, ensures that users can receive early warning information in a timely manner, and avoid potential property losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114125153B_ABST
    Figure CN114125153B_ABST
Patent Text Reader

Abstract

The present application provides a call warning method, device, server, storage medium and product, belonging to the field of mobile communication technology. The method includes: if the call is a target type call, obtaining the first audio data of the call, the first audio data includes the audio data of the first call object of the call, the target type is a call type with hidden dangers, and the first call object is the call object that causes hidden dangers; determining the first semantic information corresponding to the first audio data and the first voiceprint feature of the first call object; based on the first semantic information and the first voiceprint feature, determining the hidden danger level corresponding to the call; based on the early warning measures information corresponding to the hidden danger level, warning the call. Since the method first determines the hidden danger level of the call, and then warns the call based on the early warning measures information corresponding to the hidden danger level, it can achieve targeted early warning for the call, thereby improving the effectiveness of early warning for calls that cause hidden dangers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of mobile communication technology, and in particular to a call early warning method, device, server, storage medium and product. Background Art

[0002] With the rapid development of mobile communication technology, telephones have become a necessity in users' lives. While telephones bring convenience to users, they also bring hidden dangers to users' lives. For example, some people defraud users of their money through telephone communications, and due to the lack of awareness of users, it is easy to cause property losses. Therefore, it is very important to warn of the hidden dangers caused by these calls.

[0003] In the related art, unsafe phone numbers are marked and stored in a database, so that when an unsafe phone number calls, the phone number will carry the mark to remind the user who answers the call that the phone number is an unsafe phone number. Since users often ignore the mark of the incoming phone number when answering the call, marking the phone number cannot effectively warn the user, and thus the effectiveness of this method in warning of hidden dangers caused by calls is low. Summary of the invention

[0004] The embodiments of the present application provide a call warning method, device, server, storage medium and product, which can improve the effectiveness of warning calls that cause hidden dangers. The technical solution is as follows:

[0005] In one aspect, a call early warning method is provided, the method comprising:

[0006] If the call is a target type call, obtaining first audio data of the call, the first audio data including audio data of a first call object of the call, the target type is a call type with hidden dangers, and the first call object is a call object causing the hidden dangers;

[0007] Determining first semantic information corresponding to the first audio data and a first voiceprint feature of the first call object;

[0008] Determining a risk level corresponding to the call based on the first semantic information and the first voiceprint feature;

[0009] Based on the early warning measure information corresponding to the hidden danger level, an early warning is issued for the call.

[0010] In a possible implementation manner, determining the risk level corresponding to the call based on the first semantic information and the first voiceprint feature includes:

[0011] If the first voiceprint feature matches the target voiceprint feature, and the first semantic information matches the target semantic information, it is determined that the risk level is the first level, and the target voiceprint feature and the target semantic information are the voiceprint feature and semantic information corresponding to the call of the target type, respectively;

[0012] If the first voiceprint feature does not match the target voiceprint feature, but the first semantic information matches the target semantic information, determining that the risk level is the second level;

[0013] If the first voiceprint feature matches the target voiceprint feature, but the first semantic information does not match the target semantic information, it is determined that the hidden danger level is the third level.

[0014] In a possible implementation, the issuing of an early warning for the call based on the early warning measure information corresponding to the risk level includes:

[0015] Determining a potential risk type of the call based on the first semantic information;

[0016] Obtaining behavior information and response information corresponding to the hidden danger type, wherein the behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the method of responding to the hidden danger;

[0017] Based on the hidden danger level, the hidden danger type, the behavior information and the response information are output to the second call partner of the call.

[0018] In a possible implementation, the outputting the hidden danger type, the behavior information, and the response information to the second call partner of the call based on the hidden danger level includes:

[0019] If the hidden danger level is the first level, controlling to cut off the call, establishing an early warning call with the first terminal used by the second call partner of the call, and outputting the hidden danger type, the behavior information and the response information to the second call partner through the early warning call;

[0020] If the hidden danger level is the second level, after the call ends, a warning call is established with the first terminal used by the second call partner of the call, and the hidden danger type, the behavior information and the response information are output to the second call partner through the warning call;

[0021] If the hidden danger level is the third level, a warning message is sent to the first terminal, where the warning message includes the hidden danger type, the behavior information, and the response information.

[0022] In a possible implementation, the method further includes:

[0023] Obtaining the behavior intention of the second call party of the call;

[0024] If the behavior is intended to cause hidden danger to the second call object, a prompt message is sent to the second terminal, where the second terminal is a terminal used by a service personnel of the hidden danger event, and the prompt message is used to prompt the service personnel to dissuade the second call object.

[0025] In a possible implementation, the method further includes:

[0026] After determining the hidden danger type of the call, storing the telephone number corresponding to the first call object in a telephone number database, the telephone number database is used to store the telephone number causing the hidden danger;

[0027] Based on the risk type, the telephone number is marked in the telephone number database.

[0028] In a possible implementation, the process of determining the target type of call includes:

[0029] If the phone number corresponding to the first call object is marked as a potential risk phone number, determining that the call is a target type of call; or

[0030] If the telephone number corresponding to the first call object is not a real-name telephone number, determining that the call is a target type of call; or

[0031] Obtain second audio data of the call, determine second semantic information corresponding to the second audio data and a first voiceprint feature of the first call object, and if at least one of the second semantic information and the first voiceprint feature meets a preset condition, determine that the call is a target type of call, and the preset condition is a condition that causes hidden dangers.

[0032] In a possible implementation, the process of determining whether the first voiceprint feature meets a preset condition includes:

[0033] respectively determining similarities between the first voiceprint feature and a plurality of second voiceprint features in a voiceprint database, where the second voiceprint features are voiceprint features of an object causing hidden dangers;

[0034] Determine a target number of similarities with the largest values ​​among multiple similarities;

[0035] If the target number of similarities are all greater than a preset threshold, it is determined that the first voiceprint feature meets a preset condition.

[0036] In a possible implementation, the respectively determining similarities between the first voiceprint feature and a plurality of second voiceprint features in a voiceprint database includes:

[0037] For each second voiceprint feature, based on the first voiceprint feature, perform a vector nearest neighbor search on the second voiceprint feature to obtain a cosine distance between the first voiceprint feature and the second voiceprint feature;

[0038] The cosine distance is normalized to obtain the similarity.

[0039] In a possible implementation, before acquiring the first audio data of the call, the method further includes:

[0040] If the number of calls between the first end and the second end of the call is not greater than a preset number; or if the phone number corresponding to the first call object is not in the contact list of the second call object of the call, execute the step of obtaining the first audio data of the call.

[0041] In a possible implementation manner, the process of determining the first voiceprint feature includes:

[0042] The first audio data is input into a voiceprint recognition model, a first feature vector corresponding to the first audio data is output, and the first feature vector is determined as a first voiceprint feature of the first call object. The voiceprint recognition model is used to extract the first feature vector of the audio data.

[0043] On the other hand, a call warning device is provided, the device comprising:

[0044] A first acquisition module is used to acquire first audio data of the call if the call is a target type call, the first audio data including audio data of a first call object of the call, the target type is a call type with hidden dangers, and the first call object is a call object causing the hidden dangers;

[0045] A first determining module, configured to determine first semantic information corresponding to the first audio data and a first voiceprint feature of the first call object;

[0046] A second determination module, configured to determine a risk level corresponding to the call based on the first semantic information and the first voiceprint feature;

[0047] The early warning module is used to issue an early warning for the call based on the early warning measure information corresponding to the hidden danger level.

[0048] In a possible implementation manner, the second determining module is configured to:

[0049] If the first voiceprint feature matches the target voiceprint feature, and the first semantic information matches the target semantic information, it is determined that the risk level is the first level, and the target voiceprint feature and the target semantic information are the voiceprint feature and semantic information corresponding to the call of the target type, respectively;

[0050] If the first voiceprint feature does not match the target voiceprint feature, but the first semantic information matches the target semantic information, determining that the risk level is the second level;

[0051] If the first voiceprint feature matches the target voiceprint feature, but the first semantic information does not match the target semantic information, it is determined that the hidden danger level is the third level.

[0052] In a possible implementation, the early warning module includes:

[0053] A first determining unit, configured to determine a potential risk type of the call based on the first semantic information;

[0054] A first acquisition unit is used to acquire behavior information and response information corresponding to the hidden danger type, wherein the behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the way to respond to the hidden danger;

[0055] An output unit is used to output the hidden danger type, the behavior information and the response information to the second call partner of the call based on the hidden danger level.

[0056] In a possible implementation, the output unit is used to:

[0057] If the hidden danger level is the first level, controlling to cut off the call, establishing an early warning call with the first terminal used by the second call partner of the call, and outputting the hidden danger type, the behavior information and the response information to the second call partner through the early warning call;

[0058] If the hidden danger level is the second level, after the call ends, a warning call is established with the first terminal used by the second call partner of the call, and the hidden danger type, the behavior information and the response information are output to the second call partner through the warning call;

[0059] If the hidden danger level is the third level, a warning message is sent to the first terminal, where the warning message includes the hidden danger type, the behavior information, and the response information.

[0060] In a possible implementation manner, the device further includes:

[0061] A second acquisition module, used to acquire the behavior intention of the second call object of the call;

[0062] The sending module is used to send a prompt message to the second terminal if the behavior intends to cause hidden dangers to the second call object. The second terminal is a terminal used by the service personnel of the hidden danger event. The prompt message is used to prompt the service personnel to dissuade the second call object.

[0063] In a possible implementation manner, the device further includes:

[0064] A storage module, configured to store the telephone number corresponding to the first call object in a telephone number database after determining the hidden danger type of the call, wherein the telephone number database is used to store the telephone number causing the hidden danger;

[0065] A marking module is used to mark the telephone number in the telephone number database based on the risk type.

[0066] In a possible implementation, the first acquisition module includes:

[0067] a second determining unit, configured to determine that the call is a call of a target type if the phone number corresponding to the first call object is marked as a potential risk phone number;

[0068] a third determining unit, configured to determine that the call is a target type of call if the telephone number corresponding to the first call object is not a real-name telephone number;

[0069] The second acquisition unit is used to acquire second audio data of the call, determine second semantic information corresponding to the second audio data and a first voiceprint feature of the first call object, and if at least one of the second semantic information and the first voiceprint feature meets a preset condition, determine that the call is a target type of call, and the preset condition is a condition that causes hidden dangers.

[0070] In a possible implementation, the second acquiring unit includes:

[0071] A first determining subunit is used to respectively determine the similarity between the first voiceprint feature and a plurality of second voiceprint features in a voiceprint database, where the second voiceprint features are voiceprint features of an object causing hidden dangers;

[0072] A second determination subunit is used to determine a target number of similarities with the largest values ​​among the multiple similarities;

[0073] The third determining subunit is configured to determine that the first voiceprint feature satisfies a preset condition if the target number of similarities are all greater than a preset threshold.

[0074] In a possible implementation manner, the first determining subunit is configured to:

[0075] For each second voiceprint feature, based on the first voiceprint feature, perform a vector nearest neighbor search on the second voiceprint feature to obtain a cosine distance between the first voiceprint feature and the second voiceprint feature;

[0076] The cosine distance is normalized to obtain the similarity.

[0077] In a possible implementation manner, the device further includes:

[0078] An execution module is used to execute the step of obtaining the first audio data of the call if the number of calls between the first end and the second end of the call is not greater than a preset number; or if the telephone number corresponding to the first call object is not in the contact list of the second call object of the call.

[0079] In a possible implementation manner, the first determining module is configured to:

[0080] The first audio data is input into a voiceprint recognition model, a first feature vector corresponding to the first audio data is output, and the first feature vector is determined as a first voiceprint feature of the first call object. The voiceprint recognition model is used to extract the first feature vector of the audio data.

[0081] On the other hand, a server is provided, which includes one or more processors and one or more memories, wherein at least one instruction is stored in the one or more memories, and the at least one instruction is loaded and executed by the one or more processors to implement the operations performed by the call warning method described in any of the above-mentioned implementation methods.

[0082] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the operations performed by the call warning method described in any of the above implementations.

[0083] On the other hand, a computer program product is provided, which includes at least one instruction, and the at least one instruction is loaded and executed by a server to implement the operations performed by the call warning method described in any of the above implementations.

[0084] The beneficial effects of the technical solution provided by the embodiments of the present application include at least:

[0085] An embodiment of the present application provides a method for warning of a call. The method can determine the potential risk level of the call based on the first audio data and the first voiceprint feature of the call party that causes the potential risk in the call when there is a potential risk in the call. In this way, since the potential risk level of the call is first determined, and then the call is warned based on the warning measure information corresponding to the potential risk level, a targeted warning of the call can be achieved, thereby improving the effectiveness of warning of calls that cause potential risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0087] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0088] Figure 2 is a flow chart of a method for determining a target type of call provided by an embodiment of the present application;

[0089] Figure 3 It is a flow chart of a call warning method provided by an embodiment of the present application;

[0090] Figure 4 It is a block diagram of a call warning device provided in an embodiment of the present application;

[0091] Figure 5 is a block diagram of a terminal provided in an embodiment of the present application;

[0092] Figure 6 This is a block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0093] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0094] The terms "first", "second", "third" and "fourth" etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0095] The present application embodiment provides an implementation environment of a call warning method, see Figure 1 , the implementation environment includes a terminal 10 and a server 20. In some embodiments, the terminal 10 is a device for making a call; the terminal 10 can be a receiving end and a dialing end of a call, and can be at least one of a mobile phone, a landline, a smart watch, and other devices capable of making calls; and the terminal 10 is used to send call information of the current call to the server 20.

[0096] In some embodiments, the server 20 is a server 20 of an operator or a relevant department, and the server 20 may be at least one of a single server 20, a server cluster consisting of multiple servers 20, a cloud server, a cloud computing platform, and a virtualization center. The server 20 is used to obtain call information of the current call, and to determine whether the call is a call with hidden dangers, and if it is determined that the call is a call with hidden dangers, an early warning is issued for the call.

[0097] The present application embodiment provides a method for determining a target type of call, see Figure 2 , methods include:

[0098] Step 201: The server obtains the second audio data of the call.

[0099] The second audio data includes audio data of the second call object and audio data of the first call object, the first call object is the call object that causes the hidden danger, and the second call object is the call object to which the hidden danger is caused.

[0100] Step 202: The server determines second semantic information corresponding to the second audio data and a first voiceprint feature of the first call partner.

[0101] Among them, the process of determining the second semantic information includes: the server identifies the text information corresponding to the second audio data through the audio recognition module; the server segments the text information through the natural language processing library configured in the semantic understanding module to obtain the second semantic information including multiple words.

[0102] The process of determining the first voiceprint feature includes: the server inputs the first audio data into a voiceprint recognition model, outputs a first feature vector corresponding to the first audio data, and determines the first feature vector as the first voiceprint feature of the first call object. The voiceprint recognition model is used to extract the first feature vector of the audio data, and the first audio data includes the audio data of the first call object.

[0103] Optionally, the first eigenvector is an i-vector (a type of eigenvector); it should be noted that the traditional joint factor analysis for vector feature extraction is mainly based on two different spaces, the speaker space defined by the eigenvoice space matrix and the channel space defined by the eigenvoice space, when establishing the voiceprint recognition model. Inspired by the theory of joint factor analysis, Dehak extracted a more compact vector from the mean supervector of the Gaussian mixture model, called the i-vector vector, where i represents the meaning of identity and the i-vector vector is equivalent to the speaker's identifier. The i-vector vector uses one space to replace two spaces. This new space can become a global difference space, which contains both the voiceprint differences between speakers and the differences between channels. Therefore, the modeling process of the i-vector vector does not strictly distinguish between the influence of the speaker and the influence of the channel in the Gaussian mixture model. The modeling method is derived from Dehak's research theory that "the channel factor after JFA (Jointfactor analysis) modeling not only contains the channel effect but also contains the speaker's information." Among them, the i-vector vector is obtained through factor analysis based on Gaussian supervector. The space of the i-vector vector obtained by the cross-channel algorithm based on a single channel includes both the information of the speaker space and the channel space information, which is equivalent to projecting the audio data from a high-dimensional space to a low-dimensional space using the factor analysis method. In general, the dimension of the i-vector vector is between 400-600. The i-vector vector can represent the identity of the speaker, has strong distinguishability, and has a relatively low dimension, which can greatly reduce the amount of calculation, thereby improving the efficiency of obtaining the first voiceprint feature.

[0104] Step 203: If at least one of the second semantic information and the first voiceprint feature meets a preset condition, the server determines that the call is a target type of call.

[0105] Among them, the preset conditions are the conditions that cause hidden dangers, and the target type is the call type that has hidden dangers.

[0106] The process of determining whether the second semantic information meets the preset condition includes: the server compares the second semantic information with a plurality of preset keywords, where the plurality of keywords are words with hidden dangers; if the number of keywords included in the second semantic information exceeds a preset value, it is determined that the second semantic information meets the preset condition. Optionally, the plurality of keywords are "cheated", "loss", "I believe", "remittance", "transfer money", etc.

[0107] In one implementation, the process of determining whether the first voiceprint feature meets the preset condition includes the following steps (1)-(3):

[0108] (1) The server determines the similarity between the first voiceprint feature and a plurality of second voiceprint features in the voiceprint database, respectively, where the second voiceprint features are voiceprint features of the object causing the hidden danger.

[0109] In one implementation, the server performs a vector nearest neighbor search on each second voiceprint feature based on the first voiceprint feature to obtain a cosine distance between the first voiceprint feature and the second voiceprint feature; the server normalizes the cosine distance to obtain a similarity. In one implementation, the server performs a vector nearest neighbor search on the second voiceprint feature based on Faiss (an indexing tool).

[0110] (2) The server determines a target number of similarities with the largest values ​​among the multiple similarities.

[0111] The target quantity can be set and changed as needed and is not specifically limited here.

[0112] (3) If the target number of similarities are all greater than the preset threshold, the server determines that the first voiceprint feature meets the preset condition.

[0113] The preset threshold can be set and changed as needed, and is not specifically limited here. In the embodiment of the present application, the similarity between the first voiceprint feature and the second voiceprint feature is determined by the vector nearest neighbor search method. Since the vector nearest neighbor search method is simple in concept, mature in theory, and has low training complexity and high accuracy, it improves the efficiency of determining whether the first voiceprint feature meets the preset conditions.

[0114] An embodiment of the present application provides a method for determining a target type of call. The method determines whether a call is a target type of call through semantic information and voiceprint features. Since the semantic information and voiceprint features can be used to represent the semantics of whether it is a hidden danger and the voiceprint features of whether it is an object causing the hidden danger, respectively, determining whether a call is a target type of call through semantic information and voiceprint features can improve the accuracy of determining the target type of call.

[0115] The present application embodiment provides a call warning method, see Figure 3 , methods include:

[0116] Step 301: The server obtains call information of the current call, and determines whether the call is a target type call based on the call information.

[0117] In some embodiments, the receiving end of the call sends the call information of the call to the server, and starts sending the call information of the call to the server when the receiving end receives the incoming call of the call. The call information includes the telephone numbers of the receiving end and the dialing end of the call, the mark of the telephone number, the audio data during the call, etc. Among them, the mark of the telephone number includes the location mark, the contact mark, the call type mark, etc. For example, they can be respectively marked as "** Province ** City", "Dad", "Express Delivery", "Nuisance Call", "Hidden Danger Phone Number", "Advertising Promotion", etc. In some embodiments, the dialing end of the call sends the call information of the call to the server, and starts sending the call information of the call to the server when the dialing end makes a call.

[0118] In one implementation, if the phone number corresponding to the first call object is marked as a potential risk phone number, the server determines that the call is a target type call.

[0119] The first call object is the call object that causes hidden dangers. Since the mark of the phone number can be directly displayed when a call comes in, the server can directly obtain the mark of the phone number and determine whether the call is a target type of call based on the mark, thereby improving the efficiency of the server in determining whether the call is a target type of call.

[0120] In another implementation, if the telephone number corresponding to the first call object is not a real-name telephone number, the server determines that the call is a target type call.

[0121] It should be noted that the real-name telephone number includes a variety of fixed telephone number formats; for example, a mobile phone number is an 11-digit telephone number starting with 1, a landline number is a 12-digit telephone number starting with an area code, and a bank number is a 5-digit telephone number starting with 952. Optionally, some non-real-name telephone numbers are virtual telephone numbers starting with 00, 8-digit telephone numbers starting with 952, etc. In this way, whether the call is a target type of call is determined by the format of the telephone number, which improves the efficiency of the server in determining whether the call is a target type of call.

[0122] In another implementation, the server determines the target type of call through steps 201-203, which are not described in detail here. It should be noted that there is no execution order for the above implementations, and the server can determine the target type of call through any implementation.

[0123] It should be noted that if the server determines that the call is a target type call, step 302 is executed; if the server determines that the call is not a target type call, the warning is terminated.

[0124] Step 302: If the call is a target type call, the server obtains first audio data of the call.

[0125] The first audio data includes audio data of the first call object of the call, the target type is a call type with hidden dangers, and the first call object is the call object that causes the hidden dangers.

[0126] In one implementation, before the server obtains the first audio data of a call, if the number of calls between the first end and the second end of the call is not greater than a preset number, the server executes the step of obtaining the first audio data of the call.

[0127] The first end and the second end are the receiving end and the dialing end of the call respectively; the preset number can be set and changed as needed; optionally, the preset number is 1. In this implementation, since the first call object that causes hidden dangers will not frequently call the second call object, only the first audio data of the call with a small number of calls is obtained, avoiding the waste of resources caused by obtaining the first audio data of each call.

[0128] In another implementation, before the server obtains the first audio data of a call, if the phone number corresponding to the first call party is not in the contact list of the second call party of the call, the server executes the step of obtaining the first audio data of the call.

[0129] The contact list includes a phone number list that names the phone numbers; for example, if a phone number is marked as "Dad", then the phone number is in the contact list. In this implementation, since the contact list stores the phone numbers of known call partners, the probability that it is the phone number of the first call partner is small; thus, only the first audio data of calls whose phone numbers are not in the contact list is obtained, avoiding the waste of resources caused by obtaining the first audio data of each call.

[0130] Step 303: The server determines the first semantic information corresponding to the first audio data and the first voiceprint feature of the first call partner.

[0131] It should be noted that the first voiceprint feature of the first call object in this step is the same as the first voiceprint feature of the first call object determined in step 202, and in this step the server can directly obtain the first voiceprint feature determined in step 202. The process in which the server determines the first semantic information corresponding to the first audio data in this step is the same as the process in which the server determines the second semantic information corresponding to the second audio data in step 202, and will not be repeated here.

[0132] Step 304: The server determines the risk level corresponding to the call based on the first semantic information and the first voiceprint feature.

[0133] In this step, if the first voiceprint feature matches the target voiceprint feature, and the first semantic information matches the target semantic information, the server determines the hidden danger level to be level 1. If the first voiceprint feature does not match the target voiceprint feature, but the first semantic information matches the target semantic information, the server determines the hidden danger level to be level 2. If the first voiceprint feature matches the target voiceprint feature, but the first semantic information does not match the target semantic information, the server determines the hidden danger level to be level 3.

[0134] The target voiceprint feature and the target semantic information are the voiceprint feature and semantic information corresponding to the target type of call, respectively. The target voiceprint feature includes the voiceprint feature of the object causing the hidden danger in the voiceprint database stored in advance. In one implementation, the process of determining whether the first voiceprint feature matches the target voiceprint feature is the same as the process of determining whether the first voiceprint feature meets the preset condition in step 203. In one implementation, if the server determines that the first voiceprint feature meets the preset condition in step 203, it is determined that the first voiceprint feature matches the target voiceprint feature. If the server determines that the first voiceprint feature does not meet the preset condition in step 203, it is determined that the first voiceprint feature does not match the target voiceprint feature.

[0135] The target semantic information includes a plurality of keywords stored in advance; in one implementation, the process of determining whether the first semantic information matches the target semantic information includes: the server compares the first semantic information with a plurality of preset keywords, wherein the plurality of keywords are words with hidden dangers; if the number of keywords included in the first semantic information exceeds a preset value, it is determined that the first semantic information matches the target semantic information. Optionally, the plurality of keywords are "cheated", "loss", "I believe", "remittance", "transfer money", etc.

[0136] In the embodiment of the present application, the risk level of the call is determined by the first voiceprint feature and the first semantic information, so the degree of danger that the call will cause to the call party who is causing the risk can be determined, and then different early warning measures can be taken for the call based on the risk type of the call, so that the early warning is targeted and purposeful, thereby improving the effectiveness of the early warning.

[0137] Step 305: The server issues an early warning for the call based on the early warning measure information corresponding to the risk level.

[0138] The warning measures include at least one of the following: the server disconnects the call, dials the warning call, sends the warning information, etc. This step includes the following steps (1)-(3):

[0139] (1) The server determines the potential risk type of the call based on the first semantic information.

[0140] Among them, the risk types include low-price shopping risks, credit card application risks, card consumption risks, and induced transfer risks. Optionally, when the first semantic information includes "credit card", the server determines that the risk type of the call is a credit card application risk. When the first semantic information includes "transfer", the server determines that the risk type of the call is an induced transfer risk.

[0141] In some embodiments, after determining the hidden danger type of the call, the server stores the phone number corresponding to the first call object in a phone number database, and the phone number database is used to store phone numbers that cause hidden dangers. The server marks the phone number in the phone number database based on the hidden danger type. In this way, when the phone number calls, it can be directly determined based on the mark that the call corresponding to the phone number is a call of the target type, and the hidden danger type of the call can be directly determined, so that the call can be effectively warned.

[0142] (2) The server obtains behavior information and response information corresponding to the hidden danger type. The behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the method of responding to the hidden danger.

[0143] In one implementation, the server obtains behavior information and response information corresponding to multiple hidden danger types in advance; and associates and stores each hidden danger type and the behavior information and response information corresponding to the hidden danger type, so that the server can directly obtain the corresponding behavior information and response information based on the hidden danger type.

[0144] Optionally, for the hidden danger type of inducing money transfer, the corresponding behavioral information includes impersonating a bank employee, providing a so-called safe account on the grounds of upgrading the bank card, inducing the transfer of funds to a designated account, etc.; the corresponding response information includes hanging up the call, adding the phone number of the call to the blacklist, reporting the phone number of the call, etc.

[0145] (3) Based on the risk level, the server outputs the risk type, behavior information, and response information to the second call partner of the call.

[0146] This step includes the following implementation methods:

[0147] A1: If the hidden danger level is the first level, the server controls the call to be disconnected, and establishes a warning call with the first terminal used by the second call partner. The server outputs the hidden danger type, behavior information and response information to the second call partner through the warning call.

[0148] It should be noted that, if the server is the operator's server, the server directly controls the call disconnection. If the server is not the operator's server, the server sends an alarm message to the third terminal used by the operator through the early warning module, and the alarm message is used to prompt the operator to disconnect the call.

[0149] Among them, the server establishes a warning call with the first terminal through the outbound call module; after the server controls to cut off the call, it triggers the outbound call module to automatically call the first terminal used by the second call object, and outputs the hidden danger type, behavior information and response information to the second call object through the pre-trained intelligent voice. In this implementation, since the hidden danger level is the first level, the hidden danger level is relatively high, so by controlling the early warning measures of cutting off the call, the hidden danger can be avoided in time, and by establishing an early warning call to output the hidden danger type, behavior information and response information to the second call object, the second call object can prevent the hidden danger in time, further avoiding the situation of causing hidden dangers, thereby improving the effectiveness of the early warning.

[0150] A2: If the hidden danger level is the second level, then after the call ends, an early warning call is established with the first terminal used by the second call partner, and the hidden danger type, behavior information and response information are output to the second call partner through the early warning call.

[0151] In this implementation, after the server receives the call end information, it triggers the outbound call module to automatically call the first terminal used by the second call object, and outputs the hidden danger type, behavior information and response information to the second call object through the pre-trained intelligent voice. In this implementation, since the hidden danger level is the second level, the hidden danger level is also high. After the call ends, the hidden danger type, behavior information and response information are output to the second call object by establishing an early warning call, so that the second call object can prevent hidden dangers in time and avoid the situation that causes hidden dangers, thereby improving the effectiveness of the early warning.

[0152] A3: If the hidden danger level is the third level, a warning message is sent to the first terminal, where the warning message includes the hidden danger type, behavior information, and response information.

[0153] In one implementation, after determining that the hidden danger level is the third level, the server sends a warning message to the first terminal. In another implementation, the server sends a warning message to the first terminal after the call ends. Optionally, the warning message is sent in the form of a text message. In this implementation, since the hidden danger level is the third level, the hidden danger level is relatively low. In this way, only by sending the warning message to the first terminal, the inconvenience caused by cutting off the call to the second call party is avoided, and by sending the warning message to the first terminal, the second call party can prevent the hidden danger in time and avoid the situation of causing the hidden danger, thereby improving the effectiveness of the warning.

[0154] Step 306: The server obtains the behavior intention of the second call partner of the call; if the behavior intention poses a hidden danger to the second call partner, the server sends a prompt message to the second terminal.

[0155] The second terminal is a terminal used by a service personnel of a potential danger event, and the prompt information is used to prompt the service personnel to dissuade the second call object. In one implementation, the service personnel establishes a dissuading call with the first terminal used by the second call object through the second terminal, thereby facilitating the service personnel to dissuade the second call object.

[0156] In one implementation, the server determines the behavior intention of the second call object based on the second semantic information of the second call object. For example, if the second semantic information includes information such as "I trust you" and "I will transfer money immediately", the server determines that the behavior intention of the second call object is to transfer money, that is, the behavior intention is that the second call object will cause hidden dangers.

[0157] In another implementation, the server determines the behavioral intention of the second call object based on the third semantic information of the second service object in the warning call. For example, if the third semantic information includes information such as "I don't believe you" and "Don't call", the server determines that the behavioral intention of the second call object is to believe the first call object and not believe the warning call, that is, the behavioral intention is that the second call object will be exposed to hidden dangers.

[0158] In the embodiment of the present application, when the behavior of the second call party intends to cause hidden dangers to the second call party, the service personnel dissuade the second call party, thereby further strengthening the early warning of the call, which can effectively avoid hidden dangers to the second call party, thereby improving the effectiveness of the early warning of the call.

[0159] An embodiment of the present application provides a method for warning of a call. The method can determine the potential risk level of the call based on the first audio data and the first voiceprint feature of the call party that causes the potential risk in the call when there is a potential risk in the call. In this way, since the potential risk level of the call is first determined, and then the call is warned based on the warning measure information corresponding to the potential risk level, a targeted warning of the call can be achieved, thereby improving the effectiveness of warning of calls that cause potential risks.

[0160] The present application also provides a call warning device, see Figure 4 , the device comprises:

[0161] A first acquisition module 401 is used to acquire first audio data of the call if the call is a target type call, the first audio data including audio data of a first call object of the call, the target type is a call type with hidden dangers, and the first call object is a call object causing the hidden dangers;

[0162] A first determination module 402, configured to determine first semantic information corresponding to the first audio data and a first voiceprint feature of a first call partner;

[0163] A second determination module 403, configured to determine a risk level corresponding to the call based on the first semantic information and the first voiceprint feature;

[0164] The warning module 404 is used to issue a warning for the call based on the warning measure information corresponding to the potential danger level.

[0165] In a possible implementation, the second determining module 403 is configured to:

[0166] If the first voiceprint feature matches the target voiceprint feature, and the first semantic information matches the target semantic information, it is determined that the risk level is the first level, and the target voiceprint feature and the target semantic information are the voiceprint feature and semantic information corresponding to the target type of call, respectively;

[0167] If the first voiceprint feature does not match the target voiceprint feature, but the first semantic information matches the target semantic information, the hidden danger level is determined to be the second level;

[0168] If the first voiceprint feature matches the target voiceprint feature, but the first semantic information does not match the target semantic information, the hidden danger level is determined to be the third level.

[0169] In a possible implementation, the early warning module 404 includes:

[0170] A first determining unit, configured to determine a potential risk type of the call based on the first semantic information;

[0171] A first acquisition unit is used to acquire behavior information and response information corresponding to the hidden danger type, where the behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the way to respond to the hidden danger;

[0172] The output unit is used to output the hidden danger type, behavior information and response information to the second call party of the call based on the hidden danger level.

[0173] In a possible implementation, the output unit is used to:

[0174] If the hidden danger level is the first level, the call is cut off, and a warning call is established with the first terminal used by the second call partner of the call, and the hidden danger type, behavior information and response information are output to the second call partner through the warning call;

[0175] If the hidden danger level is the second level, after the call ends, a warning call is established with the first terminal used by the second call partner of the call, and the hidden danger type, behavior information and response information are output to the second call partner through the warning call;

[0176] If the hidden danger level is the third level, a warning message is sent to the first terminal, and the warning message includes the hidden danger type, behavior information and response information.

[0177] In a possible implementation, the device further includes:

[0178] A second acquisition module is used to acquire the behavior intention of the second call object of the call;

[0179] The sending module is used to send a prompt message to the second terminal if the behavioral intention causes hidden dangers to the second call object. The second terminal is a terminal used by the service personnel of the hidden danger event. The prompt message is used to prompt the service personnel to dissuade the second call object.

[0180] In a possible implementation, the device further includes:

[0181] A storage module, configured to store the telephone number corresponding to the first call object in a telephone number database after determining the hidden danger type of the call, wherein the telephone number database is used to store the telephone number causing the hidden danger;

[0182] The marking module is used to mark the telephone number in the telephone number database based on the hidden danger type.

[0183] In a possible implementation, the first acquisition module 401 includes:

[0184] A second determining unit is configured to determine that the call is a call of a target type if the phone number corresponding to the first call object is marked as a potential risk phone number;

[0185] A third determining unit, configured to determine that the call is a call of a target type if the phone number corresponding to the first call object is not a real-name phone number;

[0186] The second acquisition unit is used to acquire the second audio data of the call, determine the second semantic information corresponding to the second audio data and the first voiceprint feature of the first call object, and if at least one of the second semantic information and the first voiceprint feature meets the preset conditions, determine that the call is a target type of call, and the preset conditions are conditions that cause hidden dangers.

[0187] In a possible implementation, the second acquisition unit includes:

[0188] A first determining subunit is used to respectively determine the similarity between the first voiceprint feature and a plurality of second voiceprint features in the voiceprint database, where the second voiceprint features are voiceprint features of the object causing the hidden danger;

[0189] A second determination subunit is used to determine a target number of similarities with the largest values ​​among the multiple similarities;

[0190] The third determination subunit is used to determine that the first voiceprint feature meets a preset condition if the target number of similarities are all greater than a preset threshold.

[0191] In a possible implementation manner, the first determining subunit is configured to:

[0192] For each second voiceprint feature, based on the first voiceprint feature, perform a vector nearest neighbor search on the second voiceprint feature to obtain a cosine distance between the first voiceprint feature and the second voiceprint feature;

[0193] The cosine distance is normalized to obtain the similarity.

[0194] In a possible implementation, the device further includes:

[0195] The execution module is used to execute the step of obtaining the first audio data of the call if the number of calls between the first end and the second end of the call is not greater than a preset number; or if the telephone number corresponding to the first call object is not in the contact list of the second call object of the call.

[0196] In a possible implementation, the first determining module 402 is configured to:

[0197] The first audio data is input into a voiceprint recognition model, a first feature vector corresponding to the first audio data is output, and the first feature vector is determined as a first voiceprint feature of the first call object. The voiceprint recognition model is used to extract the first feature vector of the audio data.

[0198] Figure 5The structure block diagram of a terminal 500 provided by an exemplary embodiment of the present application is shown. The terminal 500 may be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 500 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.

[0199] Typically, the terminal 500 includes a processor 501 and a memory 502 .

[0200] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0201] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one instruction, which is used to be executed by the processor 501 to implement the call warning method provided in the method embodiment of the present application.

[0202] In some embodiments, the terminal 500 may also optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, the memory 502 and the peripheral device interface 503 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 503 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, a positioning assembly 508 and a power supply 509.

[0203] The peripheral device interface 503 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.

[0204] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 504 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 504 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0205] The display screen 505 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 505 is a touch display screen, the display screen 505 also has the ability to collect touch signals on the surface or above the surface of the display screen 505. The touch signal can be input to the processor 501 as a control signal for processing. At this time, the display screen 505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 505 can be one, arranged on the front panel of the terminal 500; in other embodiments, the display screen 505 can be at least two, respectively arranged on different surfaces of the terminal 500 or in a folding design; in other embodiments, the display screen 505 can be a flexible display screen, arranged on a curved surface or a folding surface of the terminal 500. Even, the display screen 505 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 505 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode, organic light-emitting diode).

[0206] The camera assembly 506 is used to capture images or videos. Optionally, the camera assembly 506 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 506 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0207] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 501 for processing, or input them into the radio frequency circuit 504 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 500. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0208] The positioning component 508 is used to locate the current geographical location of the terminal 500 to implement navigation or LBS (Location Based Service). The positioning component 508 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system or Russia's Galileo system.

[0209] The power supply 509 is used to power various components in the terminal 500. The power supply 509 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0210] In some embodiments, the terminal 500 further includes one or more sensors 510 , including but not limited to: an acceleration sensor 511 , a gyroscope sensor 512 , a pressure sensor 513 , a fingerprint sensor 514 , an optical sensor 515 , and a proximity sensor 516 .

[0211] The acceleration sensor 511 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 500. For example, the acceleration sensor 511 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 501 can control the display screen 505 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 511. The acceleration sensor 511 can also be used to collect game or user motion data.

[0212] The gyro sensor 512 can detect the body direction and rotation angle of the terminal 500, and the gyro sensor 512 can cooperate with the acceleration sensor 511 to collect the user's 3D actions on the terminal 500. The processor 501 can implement the following functions based on the data collected by the gyro sensor 512: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0213] The pressure sensor 513 can be set on the side frame of the terminal 500 and / or the lower layer of the display screen 505. When the pressure sensor 513 is set on the side frame of the terminal 500, it can detect the user's holding signal of the terminal 500, and the processor 501 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 513. When the pressure sensor 513 is set on the lower layer of the display screen 505, the processor 501 controls the operability controls on the UI interface according to the user's pressure operation on the display screen 505. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0214] The fingerprint sensor 514 is used to collect the user's fingerprint, and the processor 501 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 514, or the fingerprint sensor 514 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 501 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, paying, and changing settings. The fingerprint sensor 514 can be set on the front, back, or side of the terminal 500. When a physical button or a manufacturer logo is set on the terminal 500, the fingerprint sensor 514 can be integrated with the physical button or the manufacturer logo.

[0215] The optical sensor 515 is used to collect the ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 according to the ambient light intensity collected by the optical sensor 515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is reduced. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera component 506 according to the ambient light intensity collected by the optical sensor 515.

[0216] The proximity sensor 516, also called a distance sensor, is usually disposed on the front panel of the terminal 500. The proximity sensor 516 is used to collect the distance between the user and the front of the terminal 500. In one embodiment, when the proximity sensor 516 detects that the distance between the user and the front of the terminal 500 is gradually decreasing, the processor 501 controls the display screen 505 to switch from the screen-on state to the screen-off state; when the proximity sensor 516 detects that the distance between the user and the front of the terminal 500 is gradually increasing, the processor 501 controls the display screen 505 to switch from the screen-off state to the screen-on state.

[0217] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the terminal 500, and the terminal 500 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.

[0218] Figure 6 It is a block diagram of a server provided by an embodiment of the present disclosure. The server 600 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 601 and one or more memories 602, wherein the memory 602 is used to store executable instructions, and the processor 601 is configured to execute the above executable instructions to implement the call warning method provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0219] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 602 including instructions, and the instructions can be executed by a processor 601 of a server 600 to complete the method of the above-mentioned service request. Optionally, the storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0220] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the call warning method of any of the above-mentioned implementation modes.

[0221] An embodiment of the present application also provides a computer program product, which includes at least one instruction, and the at least one instruction is loaded and executed by a server to implement the operations performed by the call warning method of any of the above-mentioned implementation modes.

[0222] In some embodiments, the computer program involved in the embodiments of the present application may be deployed and executed on one server, or on multiple servers located in one location, or on multiple servers distributed in multiple locations and interconnected by a communication network. Multiple servers distributed in multiple locations and interconnected by a communication network may constitute a blockchain system.

[0223] An embodiment of the present application provides a method for warning of a call. The method can determine the potential risk level of the call based on the first audio data and the first voiceprint feature of the call party that causes the potential risk in the call when there is a potential risk in the call. In this way, since the potential risk level of the call is first determined, and then the call is warned based on the warning measure information corresponding to the potential risk level, a targeted warning of the call can be achieved, thereby improving the effectiveness of warning of calls that cause potential risks.

[0224] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A call warning method, characterized in that: Executed by an operator server, the method includes: Acquire second audio data of a call, where the second audio data includes audio data of a first call object and audio data of a second call object of the call, where the first call object is the call object causing the hidden danger, and the second call object is the call object caused the hidden danger; Determine second semantic information corresponding to the second audio data and a first voiceprint feature of the first call object; If at least one of the second semantic information and the first voiceprint feature meets a preset condition, determining that the call is a target type of call, if the number of calls between the first end and the second end of the call is not greater than a preset number or if the phone number corresponding to the first call object is not in the contact list of the second call object of the call, obtaining first audio data of the call, the first audio data including audio data of the first call object, and the target type is a call type with hidden dangers; Determining first semantic information corresponding to the first audio data; If the first voiceprint feature matches the target voiceprint feature, and the first semantic information matches the target semantic information, it is determined that the hidden danger level is the first level, and the target voiceprint feature and the target semantic information are respectively the voiceprint feature and the semantic information corresponding to the call of the target type; if the first voiceprint feature does not match the target voiceprint feature, but the first semantic information matches the target semantic information, it is determined that the hidden danger level is the second level; if the first voiceprint feature matches the target voiceprint feature, but the first semantic information does not match the target semantic information, it is determined that the hidden danger level is the third level; Determining a potential risk type of the call based on the first semantic information; Obtaining behavior information and response information corresponding to the hidden danger type, wherein the behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the method of responding to the hidden danger; If the hidden danger level is the first level, the call is controlled to be cut off, and an early warning call is established with the first terminal used by the second call object, and the hidden danger type, the behavior information and the response information are output to the second call object through the early warning call; if the hidden danger level is the second level, after the call ends, an early warning call is established with the first terminal used by the second call object, and the hidden danger type, the behavior information and the response information are output to the second call object through the early warning call; if the hidden danger level is the third level, early warning information is sent to the first terminal, and the early warning information includes the hidden danger type, the behavior information and the response information; Acquire the behavior intention of the second call object based on the second semantic information; wherein, when the hidden danger level is the first level or the second level, acquire the behavior intention of the second call object based on the third semantic information of the second call object in the warning call; If one of the behavioral intentions causes hidden danger to the second call object, a prompt message is sent to the second terminal, where the second terminal is a terminal used by a service personnel of the hidden danger event, and the prompt message is used to prompt the service personnel to dissuade the second call object.

2. The method according to claim 1, characterized in that The method further comprises: After determining the hidden danger type of the call, storing the telephone number corresponding to the first call object in a telephone number database, the telephone number database is used to store the telephone number causing the hidden danger; Based on the risk type, the telephone number is marked in the telephone number database.

3. The method according to claim 1, characterized in that: The process of determining the target type of call also includes: If the phone number corresponding to the first call object is marked as a potential risk phone number, determining that the call is a target type of call; or If the telephone number corresponding to the first call object is not a real-name telephone number, the call is determined to be a target type of call.

4. The method according to claim 1, characterized in that: The process of determining whether the first voiceprint feature meets the preset condition includes: respectively determining similarities between the first voiceprint feature and a plurality of second voiceprint features in a voiceprint database, where the second voiceprint features are voiceprint features of an object causing hidden dangers; Determine a target number of similarities with the largest values ​​among multiple similarities; If the target number of similarities are all greater than a preset threshold, it is determined that the first voiceprint feature meets a preset condition.

5. The method according to claim 4, characterized in that The respectively determining the similarity between the first voiceprint feature and a plurality of second voiceprint features in the voiceprint database includes: For each second voiceprint feature, based on the first voiceprint feature, perform a vector nearest neighbor search on the second voiceprint feature to obtain a cosine distance between the first voiceprint feature and the second voiceprint feature; The cosine distance is normalized to obtain the similarity.

6. The method according to claim 1, characterized in that The process of determining the first voiceprint feature includes: The first audio data is input into a voiceprint recognition model, a first feature vector corresponding to the first audio data is output, and the first feature vector is determined as a first voiceprint feature of the first call object. The voiceprint recognition model is used to extract the first feature vector of the audio data.

7. A call warning device, characterized in that: The device comprises: A first acquisition module is configured to acquire second audio data of a call, wherein the second audio data includes audio data of a first call object and audio data of a second call object of the call, wherein the first call object is a call object causing hidden dangers, and the second call object is a call object caused by hidden dangers; determine second semantic information corresponding to the second audio data and a first voiceprint feature of the first call object; if at least one of the second semantic information and the first voiceprint feature meets a preset condition, determine that the call is a target type of call, and if the number of calls between the first end and the second end of the call is not greater than a preset number or if the phone number corresponding to the first call object is not in the contact list of the second call object of the call, acquire the first audio data of the call, wherein the first audio data includes the audio data of the first call object, and the target type is a call type with hidden dangers; A first determining module, used to determine first semantic information corresponding to the first audio data; a second determination module, configured to determine that the hidden danger level is the first level if the first voiceprint feature matches the target voiceprint feature and the first semantic information matches the target semantic information, wherein the target voiceprint feature and the target semantic information are respectively the voiceprint feature and the semantic information corresponding to the call of the target type; if the first voiceprint feature does not match the target voiceprint feature but the first semantic information matches the target semantic information, the hidden danger level is determined to be the second level; if the first voiceprint feature matches the target voiceprint feature but the first semantic information does not match the target semantic information, the hidden danger level is determined to be the third level; and determine the hidden danger type of the call based on the first semantic information; an early warning module, for determining the hidden danger type of the call based on the first semantic information; obtaining the behavior information and response information corresponding to the hidden danger type, wherein the behavior information is used to indicate the process of creating the hidden danger, and the response information is used to indicate the way of responding to the hidden danger; if the hidden danger level is the first level, then controlling the disconnection of the call, establishing an early warning call with the first terminal used by the second call object, and outputting the hidden danger type, the behavior information and the response information to the second call object through the early warning call; if the hidden danger level is the second level, then establishing an early warning call with the first terminal used by the second call object after the call ends, and outputting the hidden danger type, the behavior information and the response information to the second call object through the early warning call; if the hidden danger level is the third level, sending early warning information to the first terminal, wherein the early warning information includes the hidden danger type, the behavior information and the response information; A second acquisition module is used to acquire the behavior intention of the second call object based on the second semantic information; wherein, when the hidden danger level is the first level or the second level, the behavior intention of the second call object is also acquired based on the third semantic information of the second call object in the warning call; The sending module is used to send a prompt message to the second terminal if one of the behavioral intentions causes a hidden danger to the second call object. The second terminal is a terminal used by the service personnel of the hidden danger event. The prompt message is used to prompt the service personnel to dissuade the second call object.

8. A server, characterized in that: The server includes one or more processors and one or more memories, wherein at least one instruction is stored in the one or more memories, and the at least one instruction is loaded and executed by the one or more processors to implement the operations performed by the call warning method as described in any one of claims 1 to claim 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the operation performed by the call warning method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by the server to implement the operation performed by the call early warning method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for identifying bad conversation

    CN105516989A

  • Telephone fraud detection method, storage medium and electronic device

    CN107197463A

  • Method and system for recognizing fraud information, electronic equipment and server

    CN107360576A

  • Method and system for protecting account security

    CN110636505A

  • Voiceprint serial-parallel recognition method, individual soldier system and storage medium

    CN112802482A