Method and apparatus for determining call unconnected state, storage medium and electronic device

By performing feature extraction and phased detection of ringback audio data, combined with specific frequency bands and keyword recognition, the problem of low accuracy in the phone's missed state analysis is solved, and more accurate missed reason recognition and computing resources are achieved.

CN115766943BActive Publication Date: 2025-07-29HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211643486.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-07-29
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

In the prior art, the analysis results of the telephone missed status are relatively low, and it is impossible to accurately distinguish the reasons for missed, resulting in limited improvement in the effectiveness of outbound call services.

Method used

By extracting the ringback audio data, target audio detection and keyword detection of a specific frequency band are used, and the unconnected state is determined based on the results of the two, including frame-by-frame analysis of audio features and keyword recognition.

Benefits of technology

It improves the accuracy and efficiency of the detection status of the phone's missed state, can refine the identification of the reason for the missed, saves computing resources, and avoids repeated detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766943B_ABST
    Figure CN115766943B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and apparatus for determining a call-unconnected state, a storage medium, and an electronic device, belonging to the technical field of data processing. The method includes: extracting features of the ringback tone audio data of a target call to obtain corresponding audio features; performing target audio detection on the audio features according to the time sequence of the ringback tone audio data until a detection end condition is reached to obtain a first detection result; in response to the current existence of unprocessed audio features, performing keyword detection on the unprocessed audio features to obtain a second detection result; and determining the call-unconnected state of the target call according to the first detection result and the second detection result. The present disclosure can improve the detection accuracy of the call-unconnected state.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure described in the claims. The description herein is not admitted to be prior art merely by including it in this section.

[0003] Outbound calls have become an important means for enterprises to conduct business. In various outbound call scenarios, it is very common that the call is not connected. And differentiating the status of various unconnected calls is an important means to improve the effect of outbound call services.

[0004] In the related art, most of them roughly distinguish the reasons for the call not being connected through voice prompts according to the unconnected code returned by the telephone line provider interface, and cannot accurately analyze the reasons for the call not being connected, resulting in a low accuracy of the analysis results of the reasons for the call not being connected. Summary of the Invention

[0005] Embodiments of the present disclosure provide a method and apparatus for determining the unconnected state of a call, a storage medium, and an electronic device.

[0006] In the first aspect of the embodiments of the present disclosure, a method for determining the unconnected state of a call is provided. The method includes: extracting features from the ringback tone audio data of a target call to obtain corresponding audio features; performing target audio detection on the audio features according to the time sequence of the ringback tone audio data until a detection end condition is reached to obtain a first detection result; the target audio detection is audio detection for a specific frequency band; in response to the current existence of unprocessed audio features, performing keyword detection on the unprocessed audio features to obtain a second detection result; and determining the unconnected state of the target call according to the first detection result and the second detection result.

[0007] Optionally, the performing target audio detection on the audio features includes: determining frame by frame whether the audio features are valid audio; in response to the current frame being valid audio, determining the energy ratio of the signal of the current frame in the specific frequency band in the full frequency band; in response to the energy ratio satisfying a preset first threshold range, determining the current frame as candidate audio; in response to the detection end condition not being reached and the ratio of the number of frames of the candidate audio in the number of frames of the valid audio being greater than a preset second threshold, determining the candidate audio as target audio; the detection end condition includes that the number of frames of the candidate audio is less than a preset third threshold.

[0008] Optionally, the target audio includes a first target audio, and the first detection result includes a first pending status corresponding to the first target audio. The method further includes: in response to the number of frames of the candidate audio satisfying a preset fourth threshold range, determining that the first target audio is detected; in response to the first target audio not being the target audio detected for the first time, determining that the first pending status is detected.

[0009] Optionally, the target audio includes a second target audio, and the first detection result includes that the called party rejects the call. The method further includes: in response to the number of frames of the candidate audio not satisfying the fourth threshold range, determining that the second target audio is detected; in response to detecting the second target audio and the first pending status, determining that the unconnected status of the target call is that the called party rejects the call.

[0010] Optionally, the first detection result includes a second pending status corresponding to the second target audio. The method further includes: in response to detecting the second target audio, not detecting the first pending status, and there being unprocessed audio features currently, determining that the second pending status is detected.

[0011] Optionally, the method further includes any one of the following: in response to determining the unconnected status of the target call in the target audio detection, determining that the detection end condition is reached; in response to the first target audio being the target audio detected for the first time, determining that the detection end condition is reached; in response to detecting the second target audio, not detecting the first pending status, and there being no unprocessed audio features currently, determining that the detection end condition is reached.

[0012] Optionally, after obtaining the first detection result, the method further includes: in response to there being no unprocessed audio features currently, determining the unconnected status of the target call according to the first detection result.

[0013] Optionally, the keyword detection of the unprocessed audio features includes: performing sliding sampling on the unprocessed audio features by using a sliding window to obtain sampling data; inputting the sampling data into a keyword detection model to output the probability value of the current sampling data on each label; the keyword detection model is obtained by training with historical ringback tone audio data; performing smoothing processing on the output results of consecutive multiple sampling data; in response to the probability value of the output result after smoothing processing on any label being greater than or equal to a preset fifth threshold, determining that the label is a candidate keyword of the current sampling data; in response to the consecutive cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determining that the candidate keyword is the second detection result.

[0014] Optionally, the second detection result includes that the call is in progress. Determining the unanswered status of the target call according to the first detection result and the second detection result includes: in response to the second detection result being that the call is in progress and the first detection result including the second pending status, determining that the unanswered status of the target call is that the called party rejects the call; in response to the second detection result being that the call is in progress and the first detection result not including the second pending status, determining that the unanswered status of the target call is that the called party is on a call.

[0015] Optionally, the method further includes: in response to the second detection result not being that the call is in progress, determining the unanswered status of the target call according to the keywords included in the second detection result.

[0016] In a second aspect of the embodiments of the present invention, a device for determining the unanswered status of a call is provided. The device includes: a feature extraction module, a first detection module, a second detection module, and a first status determination module. Among them, the feature extraction module is used to extract features from the ringback tone audio data of the target call to obtain corresponding audio features; the first detection module is used to perform target audio detection on the audio features according to the time sequence of the ringback tone audio data until the detection end condition is reached to obtain a first detection result; the target audio detection is audio detection for a specific frequency band; the second detection module is used to, in response to the current existence of unprocessed audio features, perform keyword detection on the unprocessed audio features to obtain a second detection result; the first status determination module is used to determine the unanswered status of the target call according to the first detection result and the second detection result.

[0017] Optionally, the first detection module is further used to: determine frame by frame whether the audio feature is valid audio; in response to the current frame being valid audio, determine the energy ratio of the signal energy of the current frame in the specific frequency band in the full frequency band; in response to the energy ratio meeting a preset first threshold range, determine the current frame as candidate audio; in response to the detection end condition not being reached and the ratio of the number of frames of the candidate audio in the number of frames of the valid audio being greater than a preset second threshold, determine the candidate audio as the target audio; the detection end condition includes that the number of frames of the candidate audio is less than a preset third threshold.

[0018] Optionally, the target audio includes a first target audio, the first detection result includes a first pending status corresponding to the first target audio, and the first detection module includes a first status detection module. The first status detection module is used to: in response to the number of frames of the candidate audio meeting a preset fourth threshold range, determine that the first target audio is detected; in response to the first target audio not being the target audio detected for the first time, determine that the first pending status is detected.

[0019] Optionally, the target audio includes a second target audio, the first detection result includes that the called party rejects the call, and the first detection module includes a second status detection module, where the second status detection module is configured to: in response to the number of frames of the candidate audio not meeting the fourth threshold range, determine that the second target audio is detected; in response to detecting the second target audio and the first pending status, determine that the unconnected status of the target call is that the called party rejects the call.

[0020] Optionally, the first detection result includes a second pending status corresponding to the second target audio, and the second status detection module is further configured to: in response to detecting the second target audio, not detecting the first pending status, and there being unprocessed audio features currently, determine that the second pending status is detected.

[0021] Optionally, the first detection module further includes an end-of-detection determination module, where the end-of-detection determination module is configured to perform any one of the following: in response to determining the unconnected status of the target call during the target audio detection, determine that the end-of-detection condition is reached; in response to the first target audio being the first detected target audio, determine that the end-of-detection condition is reached; in response to detecting the second target audio, not detecting the first pending status, and there being no unprocessed audio features currently, determine that the end-of-detection condition is reached.

[0022] Optionally, the apparatus further includes: a second status determination module, configured to, after obtaining the first detection result, in response to there being no unprocessed audio features currently, determine the unconnected status of the target call according to the first detection result.

[0023] Optionally, the second detection module includes: a sliding sampling module, a keyword detection module, a smoothing module, a first keyword determination module, and a second keyword determination module, where: the sliding sampling module is configured to perform sliding sampling on the unprocessed audio features by using a sliding window to obtain sampling data; the keyword detection module is configured to input the sampling data into a keyword detection model and output the probability value of the current sampling data on each label; the keyword detection model is trained by using historical ringback tone audio data; the smoothing module is configured to perform smoothing processing on the output results of consecutive multiple sampling data; the first keyword determination module is configured to, in response to the probability value of the output result after smoothing processing on any label being greater than or equal to a preset fifth threshold, determine that the label is a candidate keyword of the current sampling data; the second keyword determination module is configured to, in response to the consecutive cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determine that the candidate keyword is the second detection result.

[0024] Optionally, the second detection result includes that the call is in progress. The first status determination module is further configured to: in response to the second detection result indicating that the call is in progress and the first detection result including the second pending status, determine that the unanswered status of the target call is that the called party rejects the call; in response to the second detection result indicating that the call is in progress and the first detection result not including the second pending status, determine that the unanswered status of the target call is that the called party is on the phone.

[0025] Optionally, the apparatus further includes: a third status determination module, configured to, in response to the second detection result not indicating that the call is in progress, determine the unanswered status of the target call according to the keywords included in the second detection result.

[0026] In the third aspect of the embodiments of the present invention, a storage medium is provided, on which a program is stored, and when the program is executed by a processor, the method for determining the unanswered status of a call as described in the above embodiments is implemented.

[0027] In the fourth aspect of the embodiments of the present invention, an electronic device is provided, including: a processor and a memory, where the memory stores executable instructions, and the processor is configured to call the executable instructions stored in the memory to execute the method for determining the unanswered status of a call as described in the above embodiments.

[0028] According to the method for determining the unanswered status of a call provided by the embodiments of the present disclosure, on the one hand, in view of the characteristics of the ringback tone audio data, through the target audio detection for a specific frequency band, a first detection result is obtained, and then keyword detection is performed on the remaining unprocessed audio features to obtain a second detection result. By combining the first detection result of the specific frequency band and the second detection result of the keyword detection, the unanswered status of the call is comprehensively determined. The present disclosure can fully explore the audio features of each stage of the ringback tone by performing phased detection on the voice features of the unanswered call. The unanswered status of the call obtained by the embodiments of the present disclosure is more refined, and at the same time, the detection accuracy can be improved. On the other hand, by performing phased detection according to the time sequence of the ringback tone audio data, it is possible to avoid repeated detection of different parts of the ringback tone audio, save computing resources, and improve the detection efficiency. Description of the Drawings

[0029] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, where:

[0030] Figure 1 Schematically shows one of the schematic diagrams of the method flow for determining the unanswered status of a call according to an embodiment of the present disclosure.

[0031] Figure 2 Schematically shows a schematic diagram of a target audio detection process according to an embodiment of the present disclosure.

[0032] Figure 3 Schematically shows a schematic diagram of a first detection result determination process according to an embodiment of the present disclosure.

[0033] Figure 4 Schematically shows a schematic diagram of a keyword detection process according to an embodiment of the present disclosure.

[0034] Figure 5 Schematically shows a second flowchart of an implementation of a method for determining a call not connected state according to an embodiment of the present disclosure.

[0035] Figure 6 Schematically shows a block diagram of a device for determining a call not connected state according to an embodiment of the present disclosure.

[0036] Figure 7 Schematically shows a schematic diagram of a structure of an electronic device suitable for implementing an embodiment of the present invention.

[0037] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0038] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0039] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0040] According to an embodiment of the present disclosure, a method for determining a call not connected state, a device for determining a call not connected state, a storage medium, and an electronic device are provided. Summary of the Invention

[0042] A ringback tone refers to the sound heard by the calling party before a call is connected, which is audio data transmitted from the called party to the calling party. Exemplarily, the call tone heard when a call is connected is usually a long tone (such as a long beep); while a busy tone (such as a short beep) is heard when the other party is busy, and the sound is short; sometimes, there will be a voice prompt after a period of busy tone or long tone.

[0043] In the related art, in the outbound call service scenario, it is necessary to determine whether the current call is connected. If it is not connected, specific reasons need to be given. And the specific reasons for non-connection are mostly identified by error codes returned by communication providers, and this method cannot accurately identify the non-connection status of the call, such as the other party is turned off, the other party is busy, the other party rejects the call, no one answers, the service is suspended, or the mobile phone number is incorrect (such as a non-existent number), etc. In some special cases, manual marking is sometimes carried out by humans, and this method is inefficient and there are subjective human factors.

[0044] In view of the above situation, a method and device for determining the non-connection status of a call in the present disclosure are designed, which can be applied to at least one of electronic devices including but not limited to a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the method for determining the non-connection status of a call can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The present disclosure is described by taking a server as an example.

[0045] Exemplary Method

[0046] The preferred embodiments of the present disclosure are described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for the purpose of illustration and explanation of the present invention, and are not used to limit the present invention. And without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0047] The following refers to Figure 1 to describe a method for determining the non-connection status of a call according to an exemplary embodiment of the present disclosure, which can be applied to the real-time ringback tone detection process and can include steps S110 - S140.

[0048] Step S110: Extract features from the ringback tone audio data of the target call to obtain corresponding audio features.

[0049] Step S120: Perform target audio detection on the audio features according to the time sequence of the ringback tone audio data until the detection end condition is reached, and obtain a first detection result.

[0050] Step S130: In response to the existence of unprocessed audio features currently, perform keyword detection on the unprocessed audio features to obtain a second detection result.

[0051] Step S140: Determine the unanswered status of the target call according to the first detection result and the second detection result.

[0052] For the method for determining the unanswered status of a call provided by the embodiments of the present disclosure, on the one hand, in view of the characteristics of the ringback tone audio data, through target audio detection for a specific frequency band, a first detection result is obtained, and then keyword detection is performed on the remaining unprocessed audio features to obtain a second detection result. By combining the first detection result of the specific frequency band and the second detection result of the keyword detection, the unanswered status of the call is comprehensively determined. By performing staged detection on the voice features of the unanswered call, the audio features of each stage of the ringback tone can be fully exploited. The unanswered status of the call obtained through the implementation manner of the present disclosure is more refined, and at the same time, the detection accuracy can be improved. On the other hand, by performing staged detection according to the time sequence of the ringback tone audio data, repeated detection of different parts of the ringback tone audio can be avoided, saving computing resources and improving detection efficiency.

[0053] The following elaborates on each of the above steps in detail.

[0054] In step S110, feature extraction is performed on the ringback tone audio data of the target call to obtain corresponding audio features.

[0055] In the exemplary implementation manner of this example, the ringback tone audio data may be the audio data returned in real time by the called party in the target call. The real-time returned audio data can be processed online in real time to determine its unanswered status. One or more of the energy features, time-domain features (such as zero-crossing rate, autocorrelation, etc.), frequency-domain features (such as spectral centroid, Mel-frequency cepstral coefficients MFCC, filter bank-based features FBank, etc.), and perceptual features (loudness, sharpness, etc.) of the ringback tone audio data can be extracted. This example does not make any limitations in this regard.

[0056] Exemplarily, since the human ear's response to the sound spectrum is non-linear, FBank can process the audio in a manner similar to the human ear, improving the performance of speech recognition. Therefore, taking the extraction of FBank features as an example to illustrate the feature extraction process. The ringback tone audio data can be preprocessed first to maximize the audio features of the speech signal. For example, the preprocessing may include one or more of pre-emphasis, framing, and windowing processing. This example does not make any limitations in this regard. Perform a fast Fourier transform on the preprocessed data, then pass it through the Mel filter bank, and perform a logarithmic operation to obtain the FBank features. The feature dimension can be determined according to the actual situation, such as 40 dimensions.

[0057] In step S120, target audio detection is performed on the audio features according to the timing of the ringback tone audio data until the detection end condition is reached, and a first detection result is obtained.

[0058] In the present exemplary embodiment, target audio detection refers to audio detection for a specific frequency band. The specific frequency band refers to the frequency band for long response tones (such as long beeps) or short response tones (such as short beeps) in the ringback tone, such as 450 Hz or (450 ± X) Hz, where X can be set according to specific circumstances. For example, X can be set larger when the noise is large. Exemplarily, the value of X can be 25 - 100, such as X = 50.

[0059] In the present exemplary embodiment, the detection end condition can be that the call not connected state has been determined or the target audio detection ends, such as reaching a certain detection time threshold or / and frame number threshold. This is not limited in this example. In this example, real-time detection can be performed according to the timing of the ringback tone audio data.

[0060] Exemplarily, referring to Figure 2 , the target audio detection process can include the following steps.

[0061] Step S210, determine frame by frame whether the audio feature is valid audio. If so, proceed to step S220; otherwise, proceed to the judgment of the next frame.

[0062] In the present exemplary embodiment, the VAD (Voice Activity Detection) algorithm can be used to determine whether the current frame is valid audio. The function of the VAD algorithm is to detect speech and give the start and end times of the speech simultaneously, which is widely used in various speech-related tasks. For example, the VAD based on energy detection can be used to remove the silent segments of the audio, or the VAD based on neural network can be used to accurately give the start and end times of the audio features. The computational complexity of the target audio detection can be reduced through the VAD algorithm.

[0063] Step S220, determine the energy ratio of the signal energy of the current frame in the specific frequency band in the full frequency band.

[0064] In the present exemplary embodiment, an effective audio count S1 can be set and initialized to S1 = 0. When it is determined that the current frame is valid audio, the effective audio count S1 is incremented by 1. Exemplarily, the ratio of the energy of the current frame in a specific frequency band (such as the [450 - 50 Hz, 450 + 50 Hz] frequency band of the beep) in the full frequency band energy can be calculated. The specific frequency band can be fine-tuned around 450 ± 25 Hz according to the actual situation. This is not limited in this example.

[0065] Step S230: Determine whether the energy ratio in step S220 meets the preset first threshold range. If so, go to step S240; otherwise, go to step S210 to enter the judgment of the next frame.

[0066] In this exemplary embodiment, the first threshold range can be set according to the detection target and experience. For example, when the target audio is a beep sound, the first threshold range can be set to [0.1, 0.111]. Of course, this range can also be fine-tuned according to the actual situation, and this example does not make special limitations on this.

[0067] Step S240: Determine that the current frame is a candidate audio.

[0068] In this exemplary embodiment, the candidate audio determined through the above steps is a possible target audio. A candidate audio count S2 can be set and initialized to 0. When it is determined that the current frame is a candidate audio, S2 is incremented by 1, and during the whole process, S2 is gradually accumulated.

[0069] Step S250: Determine whether the number of frames of the candidate audio is greater than the preset third threshold. If so, go to step S280; otherwise, go to step S260.

[0070] In this exemplary embodiment, the third threshold refers to the maximum number of frames of the candidate audio that may actually exist, and this value can be set according to the actual situation. For example, the third threshold is set to 120. When the number of frames S2 of the candidate audio is greater than the preset third threshold, it is determined that the target audio will not appear subsequently, and the subsequent keyword detection link can be entered.

[0071] Step S260: Determine whether the ratio of the number of frames of the candidate audio in the number of frames of the valid audio is greater than the preset second threshold. If so, go to step S270; otherwise, go to step S210 to enter the judgment of the next frame.

[0072] In this exemplary embodiment, the second threshold can be set according to experience. For example, the second threshold can be set to 0.9, that is, when S2 / S1 > 0.9, it is determined that the target audio is detected. That is to say, when the ratio of the number of frames of the candidate audio in the valid audio is relatively large, it is determined that the target audio is detected.

[0073] Step S270: Determine that the candidate audio is the target audio.

[0074] In this exemplary embodiment, the target audio refers to a beep sound. Whether there is a beep sound in the ringback tone is determined through the above steps. The beep sound can include a long beep sound and a short and rapid short beep sound, which respectively indicate that the phone can be connected and the phone is busy. The beep sound is a signal with a bandwidth of 450 ± 25 Hz. Generally, the long beep sound rings for 1 second and is silent for 4 seconds, and the short beep sound rings for 0.35 seconds and is silent for 0.35 seconds.

[0075] Step S280, the target audio detection ends.

[0076] After the target audio detection in the above embodiments ends, it is also necessary to further distinguish the types and occurrence order of the target audio, that is, whether the target audio is a long beep or a short beep and its occurrence order.

[0077] In some embodiments, referring to Figure 3 , the type of the target audio can be further determined through the following steps, and then the first detection result can be determined.

[0078] Step S301, determine whether the number of frames of the candidate audio satisfies a preset fourth threshold range. If so, go to step S302; otherwise, go to step S303.

[0079] Step S302, determine that the first target audio is detected, and go to step S304.

[0080] Step S303, determine that the second target audio is detected, and go to step S306.

[0081] Step S340, determine whether the first target audio is the target audio detected for the first time. If so, go to step S310; otherwise, go to step S305.

[0082] Step S305, determine that the first pending state is detected.

[0083] Step S306, determine whether the second target audio and the first pending state are detected. If so, go to step S307; otherwise, go to S308.

[0084] Step S307, determine that the unanswered state of the target call is that the called party rejects the call.

[0085] In the present exemplary embodiment, when the first target audio and the second target audio are detected simultaneously, it is determined that the unanswered state of the target call is that the called party rejects the call. This situation corresponds to the ringback tone audio starting with a "long beep", repeating multiple times, and then starting to appear a "short beep", then this situation is considered to be that the called party rejects the call.

[0086] Step S308, determine whether there are still unprocessed audio features currently. If so, go to step S309; otherwise, go to step S310.

[0087] Step S309, determine that the second pending state is detected.

[0088] Step S310, determine that the unanswered state of the target call is other states.

[0089] In this exemplary embodiment, the fourth threshold range can be determined according to the characteristics of the first target audio, where the first target audio is a short beep and the second target audio is a long beep. Exemplarily, the fourth threshold range can be set to [30, 40]. When the number of frames of the target audio meets this threshold range, it is determined that a short beep appears; otherwise, it is determined to be a long beep. In this example, it is possible to determine whether the target audio is detected for the first time by tagging the target audio.

[0090] In this exemplary embodiment, the first detection result can include a first pending status, a second pending status, the called party rejecting the call, and other statuses. The first pending status means that a short beep is determined to be detected in this part, and the unconnected status still needs to be further determined according to the keyword detection result. The second pending status means that a long beep is determined to be detected in this part, and the unconnected status still needs to be further determined according to the keyword detection result. Other statuses refer to other special statuses other than the regular unconnected status. The regular unconnected status can include shutdown, invalid number, out of service, suspended service, unable to connect, missed call reminder, call again later, user is busy, no answer, call in progress, the called party rejecting the call, etc. The specific reason for the unconnected status of other statuses can be determined according to the error code (unconnected code) returned by the call operator, or it can be determined in combination with manual operation. This example does not make any limitations in this regard.

[0091] In some embodiments, it is also necessary to determine whether the target audio ends, that is, whether the detection end condition is reached. Whether the detection end condition is reached can be determined by any of the following:

[0092] In response to determining the unconnected status of the target call during the target audio detection, it is determined that the detection end condition is reached.

[0093] In this exemplary embodiment, as long as the unconnected status of the target call is determined during the entire detection process, the entire detection ends, that is, the target audio detection ends and there is no need to perform subsequent keyword detection processes.

[0094] In response to the first target audio being the target audio detected for the first time, it is determined that the detection end condition is reached.

[0095] In this exemplary embodiment, when the target audio detected for the first time is the first target audio (short beep), it is determined to be in other status and the detection ends.

[0096] In response to detecting the second target audio, not detecting the first pending status, and there being no unprocessed audio features currently, it is determined that the detection end condition is reached.

[0097] In this exemplary embodiment, when the second target audio (long beep) is detected, the first pending status is not detected, and there are no unprocessed audio features, it is determined to be in other status and the detection ends.

[0098] In the above embodiments, after the target audio is detected, if there are no unprocessed audio features, the subsequent keyword detection process may not be performed, and the unconnected state of the target call can be directly determined according to the first detection result. For example, the other states determined above and the called party rejection state can be used as the final detection result of the unconnected state.

[0099] In step S130, in response to the existence of unprocessed audio features, keyword detection is performed on the unprocessed audio features to obtain a second detection result.

[0100] In the exemplary implementation manner of this example, after the target audio is detected, the ringback tone audio data may have been processed or there may still be unprocessed ringback tone audio data. In the case where there is still unprocessed ringback tone audio data, keyword detection is performed on the remaining unprocessed audio features. Keyword detection can detect all occurrence positions of a specified word in the audio signal. Keyword detection is used to detect specific keywords that appear in the ringback tone audio data. Each keyword tag can correspond to one of the above-mentioned conventional unconnected states, such as power off, invalid number, out of service, suspended service, unable to connect, missed call reminder, call later, user busy, no answer, call in progress, etc.

[0101] In the exemplary implementation manner of this example, keyword detection can adopt methods such as neural networks and similarity calculation methods based on the Dynamic Time Warping (DTW) algorithm. This example does not make any limitations in this regard. The neural network can adopt convolutional networks, feedforward neural networks, or recurrent neural networks, etc. This example does not make any limitations in this regard.

[0102] Exemplarily, with reference to Figure 4 , keyword detection of the processed audio features may include the following steps S410 - S450.

[0103] Step S410, perform sliding sampling on the unprocessed audio features using a sliding window to obtain sampling data.

[0104] In the exemplary implementation manner of this example, the input data is input into the model in units of windows through the sliding sampling method. The window size can be set according to the actual situation. For example, the window size can be set to 100 frames. Then, each time the sliding window slides, 100 frames of sampling data are obtained. In this way, the data in the model is processed in units of windows. Compared with the conventional processing process based on single-frame data, the model's recognition in units of windows can accelerate the network's calculation process and can combine longer contexts, which can improve the keyword detection accuracy.

[0105] Step S420: Input the sampled data into the keyword detection model, and output the probability value of the current sampled data on each label; the keyword detection model is obtained by training with historical ringback tone audio data.

[0106] In the exemplary embodiment, the keyword detection model can be a neural network model. Here, a convolutional neural network model is taken as an example for illustration. Convolutional Neural Networks (CNN) is a class of feedforward neural networks with convolutional calculations and a deep structure. The convolutional neural network has the ability of feature learning and can perform shift-invariant classification on the input information according to its hierarchical structure.

[0107] In the exemplary embodiment, the keyword detection model based on the convolutional neural network can be obtained by training with historical ringback tone audio data. The keyword detection model can include three convolutional layers, one pooling layer, and two fully connected layers. The sampled data is sequentially subjected to convolution processing, pooling, and full connection, and finally the posterior probability of the current sampled data on a certain label is output in the softmax layer.

[0108] Step S430: Smooth the output results of continuous multiple sampled data.

[0109] In the exemplary embodiment, median filtering can be used to smooth the continuous keyword probability values, so as to avoid the problem that the continuity of the keyword detection results is poor and the accuracy of the detection results is affected due to the richness of audio data changes (such as noise fluctuations, etc.).

[0110] Exemplarily, the smoothing process is as follows: A smoothing window with an odd length N can be set. After arranging the posterior probability values of continuous N sampled data in the smoothing window in ascending order, the probability with the middle position index of (N + 1) / 2 is taken as the output probability of the current window. For example, for a smoothing window with a length of 5, the probability sequence of 5 consecutive frames is [0.97, 0.96, 0.8, 0.93, 0.95]. After sorting, the probability at the middle position is taken as the output of the current smoothing window, that is, 0.95.

[0111] Step S440: In response to the probability value of the output result after smoothing on any label being greater than or equal to a preset fifth threshold, determine that the label is the candidate keyword of the current sampled data.

[0112] In the present exemplary embodiment, for any keyword tag, it can be set that when the probability value of the current sampled data on the tag is greater than a fifth threshold (such as 0.9 or 0.92, etc.), it is determined that the window sampled data may belong to the keyword, that is, it is a candidate keyword.

[0113] Step S450, in response to the continuous cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determine the candidate keyword as the second detection result.

[0114] In the present exemplary embodiment, the sixth threshold can be set according to experience, and can be set to several or dozens, such as 20, and the present example does not limit this. It can be set that after smoothing processing, if the continuous cumulative number of frames (such as the number of consecutive windows) of a certain keyword tag exceeds the sixth threshold, it is determined that the keyword is detected.

[0115] In step S140, according to the first detection result and the second detection result, determine the unconnected state of the target call.

[0116] In the present exemplary embodiment, the unconnected state can be determined by combining the audio detection result of a specific frequency band and the keyword detection result. Exemplarily, when the first detection result is a long tone and the second detection result includes that the call is in progress, it is determined that the unconnected state of the target call is that the called party rejects the call. The usage strategy of the outbound call line can be adjusted according to the determined unconnected state to improve the outbound call effect.

[0117] Exemplarily, in response to the second detection result being that the call is in progress and the first detection result including the second pending state, determine that the unconnected state of the target call is that the called party rejects the call.

[0118] In the present exemplary embodiment, when the second detection result is that the call is in progress and the first detection result includes the second pending state, it can correspond to the situation where the ringback tone audio data starts with a "long beeping sound" and after repeating multiple times, a Chinese prompt sound "the call is in progress" appears, and the unconnected state of this situation is determined to be that the called party rejects the call.

[0119] In response to the second detection result being that the call is in progress and the first detection result not including the second pending state, determine that the unconnected state of the target call is that the called party is on the phone.

[0120] In the present exemplary embodiment, when the second detection result is that the call is in progress and the first detection result does not include the second pending state, it can correspond to the situation where a Chinese prompt sound "the call is in progress" appears in the ringback tone audio data, and the unconnected state of this situation is determined to be that the called party is on the phone.

[0121] In response to the second detection result not being that the call is in progress, determine the unconnected state of the target call according to the keyword included in the second detection result.

[0122] In the present exemplary embodiment, when the second detection result does not include a call in progress, the unconnected state can be determined according to the keywords in the keyword detection result. For example, if the keyword in the second detection result is "the call you dialed is powered off", the determined unconnected state is that the called party is powered off; when the keyword in the second detection result is "transferring you to the voicemail", the determined unconnected state is missed call reminder.

[0123] The following Figure 5 introduces the specific process of a method for determining the unconnected state of a call in an embodiment of the present application.

[0124] Referring to Figure 5 , the specific process of the method for determining the unconnected state of a call may include the following steps S501 - S509.

[0125] Step S501: Receive the ringback tone audio data of the target call.

[0126] Step S502: Extract features from the ringback tone audio data of the target call to obtain the corresponding audio features.

[0127] Step S503: Perform target audio detection on the audio features in the order of the ringback tone audio data until the detection end condition is reached, and obtain the first detection result.

[0128] Step S504: Determine whether there is unprocessed audio feature currently. If yes, go to step S505; otherwise, go to step S508.

[0129] Step S505: Determine whether the first detection result includes the first target audio and the second target audio. If yes, go to step S506; otherwise, go to S507.

[0130] Step S506: Determine the unconnected state as the called party rejects the call.

[0131] Step S507: Determine the unconnected state as other state.

[0132] Step S508: Perform keyword detection on the unprocessed audio features to obtain the second detection result.

[0133] Step S509: Determine whether the second detection result is a call in progress. If yes, go to step S510; otherwise, go to step S511.

[0134] Step S510: Determine whether the first detection result includes the second target audio. If yes, go to step S506; otherwise, go to step S512.

[0135] Step S511: Determine the unconnected state of the target call according to the keywords included in the second detection result.

[0136] Step S512 determines that the unanswered status of the target call is that the called party is on a call.

[0137] The specific details of the above embodiments have been described in detail in the foregoing method for determining the unanswered call status, and thus will not be elaborated herein.

[0138] On the one hand, the method for determining the unanswered call status according to the embodiments of the present disclosure can meet the real-time detection process of the unanswered status. Although a neural network is used, by inputting and processing data in units of windows, the consumption of computing resources is reduced and the processing speed is increased; through smoothing processing, the accuracy of keyword detection is improved, and the phenomenon of missed detection is avoided. On the other hand, through the combination of target audio detection and keyword detection processes, the status that the called party rejects the call can be accurately identified. The called party's rejection means that the called party actively hangs up. The present disclosure designs a detection scheme for this status for two typical forms of rejection. One is that the ringback tone audio starts with a long beep, repeats once or multiple times, and then a short beep appears; the other is that the ringback tone audio starts with a long beep, repeats once or multiple times, and then a Chinese prompt tone appears, and the keywords of the Chinese prompt tone include "on a call". By accurately identifying the above two typical rejection statuses, it provides an important basis for formulating subsequent outbound call strategies and improves the outbound call effect.

[0139] The present disclosure can accurately identify various unanswered statuses of unanswered calls through target audio detection and keyword detection, refine the unanswered status, and improve the detection accuracy; and the phased detection can improve the detection efficiency. In addition, the present disclosure can be compatible with the detection of unanswered statuses of various Chinese prompt tones and does not depend on the language of a specific operator.

[0140] The method of the present disclosure has low resource consumption and high detection accuracy, and can detect the status of unanswered calls in large quantities at an extremely low computing cost, realizing accurate and efficient identification of the quality of outbound calls.

[0141] Exemplary Apparatus

[0142] It should be noted that for the method for determining the unanswered call status provided by the embodiments of the present disclosure, the execution subject may be a device for determining the unanswered call status, or a control module in the device for determining the unanswered call status that executes the method for determining the unanswered call status. In the embodiments of the present disclosure, the method for determining the unanswered call status is taken as an example in which the device for determining the unanswered call status executes the method for determining the unanswered call status to illustrate the method for determining the unanswered call status provided by the embodiments of the present disclosure. Next, refer to Figure 6 A device for determining the unanswered call status according to an exemplary embodiment of the present disclosure will be described.

[0143] Figure 6A block diagram of a determination device for a call not connected state according to an embodiment of the present invention is schematically shown.

[0144] Referring to Figure 6 As shown, a determination device 600 for a call not connected state according to an embodiment of the present invention, the device 600 may include: a feature extraction module 610, a first detection module 620, a second detection module 630, and a first state determination module 640; wherein, the feature extraction module 610 is configured to extract features from the ringback tone audio data of the target call to obtain corresponding audio features; the first detection module 620 is configured to perform target audio detection on the audio features according to the time sequence of the ringback tone audio data until a detection end condition is reached to obtain a first detection result; the target audio detection is audio detection for a specific frequency band; the second detection module 630 is configured to, in response to the existence of unprocessed audio features currently, perform keyword detection on the unprocessed audio features to obtain a second detection result; the first state determination module 640 is configured to determine the unconnected state of the target call according to the first detection result and the second detection result.

[0145] In some embodiments of the present disclosure, based on the foregoing solution, the first detection module 620 is further configured to: determine frame by frame whether the audio feature is valid audio; in response to the current frame being valid audio, determine the energy ratio of the signal energy of the current frame in the specific frequency band in the full frequency band; in response to the energy ratio satisfying a preset first threshold range, determine the current frame as candidate audio; in response to not reaching the detection end condition and the ratio of the number of frames of the candidate audio in the number of frames of the valid audio being greater than a preset second threshold, determine the candidate audio as the target audio; the detection end condition includes that the number of frames of the candidate audio is less than a preset third threshold.

[0146] In some embodiments of the present disclosure, based on the foregoing solution, the target audio includes a first target audio, the first detection result includes a first undetermined state corresponding to the first target audio, and the first detection module 620 includes a first state detection module, and the first state detection module is configured to: in response to the number of frames of the candidate audio satisfying a preset fourth threshold range, determine that the first target audio is detected; in response to the first target audio not being the target audio detected for the first time, determine that the first undetermined state is detected.

[0147] In some embodiments of the present disclosure, based on the foregoing solution, the target audio includes a second target audio, the first detection result includes that the called party rejects the call, and the first detection module 620 includes a second state detection module, and the second state detection module is configured to: in response to the number of frames of the candidate audio not satisfying the fourth threshold range, determine that the second target audio is detected; in response to detecting the second target audio and the first undetermined state, determine that the unconnected state of the target call is that the called party rejects the call.

[0148] In some embodiments of the present disclosure, based on the foregoing solution, the first detection result includes a second pending state corresponding to the second target audio, and the second state detection module 630 is further configured to: in response to detecting the second target audio, not detecting the first pending state, and there being unprocessed audio features currently, determine that the second pending state is detected.

[0149] In some embodiments of the present disclosure, based on the foregoing solution, the first detection module 620 further includes an end-of-detection determination module, and the end-of-detection determination module is configured to perform any one of the following: in response to determining the unconnected state of the target call in the target audio detection, determine that the end-of-detection condition is reached; in response to the first target audio being the first detected target audio, determine that the end-of-detection condition is reached; in response to detecting the second target audio, not detecting the first pending state, and there being no unprocessed audio features currently, determine that the end-of-detection condition is reached.

[0150] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus 600 further includes: a second state determination module, configured to, after obtaining the first detection result, in response to there being no unprocessed audio features currently, determine the unconnected state of the target call according to the first detection result.

[0151] In some embodiments of the present disclosure, based on the foregoing solution, the second detection module 630 includes: a sliding sampling module, a keyword detection module, a smoothing module, a first keyword determination module, and a second keyword determination module, where: the sliding sampling module is configured to perform sliding sampling on the unprocessed audio features by using a sliding window to obtain sampling data; the keyword detection module is configured to input the sampling data into a keyword detection model and output the probability value of the current sampling data on each label; the keyword detection model is obtained by training with historical ringback tone audio data; the smoothing module is configured to perform smoothing processing on the output results of continuous multiple sampling data; the first keyword determination module is configured to, in response to the probability value of the output result after smoothing processing on any label being greater than or equal to a preset fifth threshold, determine that the label is a candidate keyword of the current sampling data; the second keyword determination module is configured to, in response to the continuous cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determine that the candidate keyword is the second detection result.

[0152] In some embodiments of the present disclosure, based on the foregoing solution, the second detection result includes being in a call, and the first state determination module 640 is further configured to: in response to the second detection result being in a call and the first detection result including the second pending state, determine that the unconnected state of the target call is that the called party rejects the call; in response to the second detection result being in a call and the first detection result not including the second pending state, determine that the unconnected state of the target call is that the called party is in a call.

[0153] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus 600 further includes: a third status determination module, configured to determine the unanswered status of the target call according to the keywords included in the second detection result in response to the second detection result not being in a call.

[0154] The specific details of each module or unit in the above apparatus for determining the unanswered status of a call have been described in detail in the corresponding method for determining the unanswered status of a call, and thus will not be elaborated herein.

[0155] Exemplary Medium

[0156] After introducing the method of the exemplary embodiments of the present invention, next, the medium of the exemplary embodiments of the present invention will be described.

[0157] In some possible embodiments, each aspect of the present invention may also be implemented as a storage medium, on which program code is stored, and when the program code is executed by a processor of a device, it is used to implement the steps in the method for determining the unanswered status of a call according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0158] Specifically, when the processor of the device executes the program code, it is used to implement the following steps:

[0159] Extract features from the ringback tone audio data of the target call to obtain corresponding audio features; perform target audio detection on the audio features according to the time sequence of the ringback tone audio data until the detection end condition is reached to obtain a first detection result; the target audio detection is an audio detection for a specific frequency band; in response to the current existence of unprocessed audio features, perform keyword detection on the unprocessed audio features to obtain a second detection result; determine the unanswered status of the target call according to the first detection result and the second detection result.

[0160] The above is a schematic solution of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above method for determining the unanswered status of a call belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference may be made to the description of the technical solution of the above method for determining the unanswered status of a call.

[0161] It should be noted that: the above storage medium may be a readable storage medium. The readable storage medium may, for example, be but is not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0162] The program code contained on the readable storage medium can be transmitted with any appropriate medium, including but not limited to: wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0163] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's electronic device, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external electronic device (for example, by connecting through an Internet service provider via the Internet).

[0164] Exemplary Electronic Device

[0165] After introducing the methods, media, and devices of the exemplary embodiments of the present disclosure, next, an electronic device according to another exemplary embodiment of the present disclosure will be introduced.

[0166] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".

[0167] The following refers to Figure 7 to describe the electronic device 700 according to this embodiment of the present invention. Figure 7 The illustrated electronic device 700 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0168] As Figure 7 shown, the electronic device 700 is presented in the form of a general electronic device. The components of the electronic device 700 may include, but are not limited to: at least one of the above processing units 710, at least one of the above storage units 720, and a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710).

[0169] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section of the present specification above.

[0170] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202, and may further include a read-only storage unit (ROM) 7203.

[0171] The storage unit 720 may further include a program / utility 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0172] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0173] The electronic device 700 can also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication can be carried out through the display unit 740 and the input / output (I / O) interface 750 connected to the display unit 740. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0174] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable an electronic device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0175] The above is a schematic solution of an electronic device 700 in this embodiment. It should be noted that the technical solution of the electronic device 700 and the technical solution of the method for determining the call-unconnected state described above belong to the same concept. For the details not described in detail in the technical solution of the electronic device, reference can be made to the description of the technical solution of the method for determining the call-unconnected state described above.

[0176] It should be noted that although several modules or sub-modules of the device for determining the call-unconnected state are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0177] Moreover, although the operations of the method of the present invention are depicted in the drawings in a particular order, this is not a requirement or implication that these operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0178] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division of various aspects does not imply that the features in these aspects cannot be combined for benefit. Such division is only for the convenience of description. The present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for determining the call not connected state, characterized in that, The method includes: Extracting features from the ringback tone audio data of the target call to obtain corresponding audio features; Performing target audio detection on the audio features according to the time sequence of the ringback tone audio data until the detection end condition is reached, to obtain a first detection result; the target audio detection is audio detection for a specific frequency band; wherein, the performing target audio detection on the audio features includes: determining frame by frame whether the audio features are valid audio; in response to the current frame being valid audio, determining the energy ratio of the signal energy of the current frame in the specific frequency band in the full frequency band; in response to the energy ratio satisfying a preset first threshold range, determining the current frame as candidate audio; in response to the detection end condition not being reached and the ratio of the number of frames of the candidate audio in the number of frames of the valid audio being greater than a preset second threshold, determining the candidate audio as the target audio; the detection end condition includes that the number of frames of the candidate audio is greater than a preset third threshold; the target audio includes a first target audio, the first detection result includes a first pending status corresponding to the first target audio, and the method further includes: in response to the number of frames of the candidate audio satisfying a preset fourth threshold range, determining that the first target audio is detected; in response to the first target audio not being the first detected target audio, determining that the first pending status is detected; In response to there being unprocessed audio features currently, performing keyword detection on the unprocessed audio features to obtain a second detection result; Determining the unanswered status of the target call according to the first detection result and the second detection result.

2. The method for determining the call unconnected state according to claim 1, wherein The target audio includes a second target audio, the first detection result includes that the called party rejects the call, and the method further includes: In response to the number of frames of the candidate audio not satisfying the fourth threshold range, determining that the second target audio is detected; In response to detecting the second target audio and the first pending status, determining that the unanswered status of the target call is that the called party rejects the call.

3. The method for determining the call not connected state according to claim 2, wherein The first detection result includes a second pending status corresponding to the second target audio, and the method further includes: In response to detecting the second target audio, not detecting the first pending status, and there being unprocessed audio features currently, determining that the second pending status is detected.

4. The method for determining the call unconnected state according to claim 3, wherein The method further includes any one of the following: In response to determining the unanswered status of the target call in the target audio detection, determining that the detection end condition is reached; In response to the first target audio being the first detected target audio, determining that the detection end condition is reached; In response to detecting the second target audio, not detecting the first pending status, and there being no unprocessed audio features currently, determining that the detection end condition is reached.

5. The method for determining the call unconnected state according to claim 1, wherein After obtaining the first detection result, the method further includes: In response to there being no unprocessed audio features currently, determining the unanswered status of the target call according to the first detection result.

6. The method for determining the call-unconnected state according to claim 1, wherein The performing keyword detection on the unprocessed audio features includes: Using a sliding window to perform sliding sampling on the unprocessed audio features to obtain sampling data; Input the sampled data into a keyword detection model to output the probability values of the current sampled data on each label; the keyword detection model is obtained by training with historical ringback tone audio data; Smooth the output results of consecutive multiple sampled data; In response to the probability value of the smoothed output result on any label being greater than or equal to a preset fifth threshold, determine that label as the candidate keyword of the current sampled data; In response to the consecutive cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determine the candidate keyword as the second detection result.

7. The method for determining the call unconnected state according to claim 3, wherein The second detection result includes being in a call. Based on the first detection result and the second detection result, determining the unanswered state of the target call includes: In response to the second detection result being in a call and the first detection result including the second pending state, determine the unanswered state of the target call as the called party rejecting the call; In response to the second detection result being in a call and the first detection result not including the second pending state, determine the unanswered state of the target call as the called party being in a call.

8. The method for determining the call-unconnected state according to claim 7, wherein The method further includes: In response to the second detection result not being in a call, determine the unanswered state of the target call according to the keyword included in the second detection result.

9. A determination device for a call not connected state, characterized in that, The apparatus includes: A feature extraction module, configured to extract features from the ringback tone audio data of the target call to obtain corresponding audio features; A first detection module, configured to perform target audio detection on the audio features in the time sequence of the ringback tone audio data until a detection end condition is reached to obtain a first detection result; the target audio detection is audio detection for a specific frequency band; the first detection module is further configured to: determine frame by frame whether the audio feature is valid audio; in response to the current frame being valid audio, determine the energy ratio of the signal energy of the current frame in the specific frequency band in the full frequency band; in response to the energy ratio satisfying a preset first threshold range, determine the current frame as candidate audio; in response to the detection end condition not being reached and the ratio of the number of frames of the candidate audio in the number of frames of the valid audio being greater than a preset second threshold, determine the candidate audio as target audio; the detection end condition includes that the number of frames of the candidate audio is less than a preset third threshold; the target audio includes a first target audio, the first detection result includes a first pending state corresponding to the first target audio, and the first detection module includes a first state detection module, and the first state detection module is configured to: in response to the number of frames of the candidate audio satisfying a preset fourth threshold range, determine that the first target audio is detected; in response to the first target audio not being the first detected target audio, determine that the first pending state is detected; A second detection module, configured to perform keyword detection on the unprocessed audio features in response to the existence of unprocessed audio features to obtain a second detection result; A first state determination module, configured to determine the unanswered state of the target call according to the first detection result and the second detection result.

10. The determining device for the call-unconnected state according to claim 9, wherein The target audio includes a second target audio, the first detection result includes that the called party rejects the call, the first detection module includes a second status detection module, and the second status detection module is used for: In response to the number of frames of the candidate audio not meeting the fourth threshold range, determining that the second target audio is detected; In response to detecting the second target audio and the first pending status, determining that the unconnected status of the target call is that the called party rejects the call.

11. The determining device for the call-unconnected state according to claim 10, wherein The first detection result includes a second pending status corresponding to the second target audio, and the second status detection module is further used for: In response to detecting the second target audio, not detecting the first pending status, and there being unprocessed audio features currently, determining that the second pending status is detected.

12. The determination device for the call not connected state according to claim 11, wherein The first detection module further includes a detection end determination module, and the detection end determination module is used to perform any one of the following: In response to determining the unconnected status of the target call in the target audio detection, determining that the detection end condition is reached; In response to the first target audio being the first detected target audio, determining that the detection end condition is reached; In response to detecting the second target audio, not detecting the first pending status, and there being no unprocessed audio features currently, determining that the detection end condition is reached.

13. The determination device for the call-unconnected state according to claim 9, wherein The device further includes: A second status determination module, configured to, after obtaining the first detection result, in response to there being no unprocessed audio features currently, determine the unconnected status of the target call according to the first detection result.

14. The determination device for the call not connected state according to claim 9, wherein The second detection module includes: A sliding sampling module, configured to perform sliding sampling on the unprocessed audio features by using a sliding window to obtain sampling data; A keyword detection module, configured to input the sampling data into a keyword detection model and output the probability value of the current sampling data on each label; the keyword detection model is obtained by training with historical ringback tone audio data; A smoothing module, configured to perform smoothing processing on the output results of consecutive multiple sampling data; A first keyword determination module, configured to, in response to the probability value of the output result after smoothing processing on any label being greater than or equal to a preset fifth threshold, determine that label as the candidate keyword of the current sampling data; A second keyword determination module, configured to, in response to the consecutive cumulative number of frames of the candidate keyword being greater than or equal to a preset sixth threshold, determine the candidate keyword as the second detection result.

15. The determination device for the call-unconnected state according to claim 11, wherein The second detection result includes being in a call, and the first status determination module is further used for: In response to the second detection result being in a call and the first detection result including the second pending status, determining that the unconnected status of the target call is that the called party rejects the call; In response to the second detection result being in a call and the first detection result not including the second pending status, determining that the unconnected status of the target call is that the called party is in a call.

16. The determination device for the call not connected state according to claim 15, wherein The device further includes: A third status determination module, configured to, in response to the second detection result not being in a call, determine the unconnected status of the target call according to the keyword included in the second detection result.

17. A storage medium having a program stored thereon, the program, when executed by a processor, implementing the method according to any one of claims 1 to 8.

18. An electronic device, comprising: A processor and a memory, the memory storing executable instructions, the processor being configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Phone-call recording access failure reason recognizing method

    CN109658939A

  • Audio processing method and device, storage medium and electronic equipment

    CN113808591A