A call state detection method, device and computer readable storage medium
By analyzing and calculating the cross-correlation value of voice frame signals, the problem of insufficient accuracy in call status detection in existing technologies is solved, enabling accurate detection of sound playback in instant messaging and improving the accuracy of call status detection.
Patent Information
- Application Number
- CN202111291918.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-11-02
AI Technical Summary
Existing call status detection methods determine call status by calculating the RMS of the collected signal, which cannot accurately represent whether the listener has actually heard the sound, resulting in insufficient detection accuracy.
By receiving the current call information sent by the first terminal, the voice frame signal is parsed and the cross-correlation value of the recorded frame signal and the voice frame signal to be played is calculated to determine the call status.
It improves the accuracy of call status detection, enabling it to detect whether sound is played successfully throughout the entire instant messaging process, thus enhancing call quality.
Smart Images

Figure CN116074440B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication technology, in particular to a call state detection method, device and computer readable storage medium. BACKGROUND
[0002] In recent years, with the rapid development of Internet technology, instant messaging applications have become more and more. When making audio and video calls through instant messaging applications, the sounder cannot confirm whether the other party can hear his voice because the caller is in different places, so it is necessary to detect the current call state. The existing call state detection method mainly calculates the rms (root mean square) of the collected signal to detect the size of the signal volume, and then determines the current call state.
[0003] In the research and practice process of the prior art, the inventor of the present application found that the call state determined by calculating the rms of the collected signal only represents whether the local sound collection function is normal, and cannot represent whether the listener really hears. In instant messaging, the collected sound needs to be processed, transmitted and played, etc. Any fault in any link will cause the listener to hear no real sound, thus resulting in insufficient accuracy of call state detection. SUMMARY
[0004] The embodiment of the present application provides a call state detection method, device and computer readable storage medium, which can improve the accuracy of call state detection.
[0005] A call state detection method, comprising:
[0006] Receiving current call information sent by a first terminal, the current call information comprising a voice frame signal and attribute information of the voice frame signal;
[0007] According to the attribute information, the voice frame signal is analyzed to obtain a to-be-played voice frame signal and a current voice frame number of the voice frame signal;
[0008] Playing the to-be-played voice frame signal, and collecting the played to-be-played voice frame signal to obtain a recording frame signal corresponding to the to-be-played voice frame signal;
[0009] Based on the attribute information, the current voice frame number is verified, and a cross-correlation value of the recording frame signal and the to-be-played voice frame signal is calculated, the cross-correlation value being used to indicate the similarity of the recording frame signal and the to-be-played voice frame signal;
[0010] According to the verification result of the current voice frame number and the cross-correlation value, the current call state with the first terminal is determined.
[0011] Correspondingly, the embodiment of the present application provides a call state detection device, comprising:
[0012] a receiving unit configured to receive current call information sent by a first terminal, wherein the current call information comprises a voice frame signal and attribute information of the voice frame signal;
[0013] a parsing unit configured to parse the voice frame signal according to the attribute information to obtain a to-be-played voice frame signal and a current voice frame number of the voice frame signal;
[0014] a collecting unit configured to play the to-be-played voice frame signal to collect a recorded voice frame signal corresponding to the to-be-played voice frame signal;
[0015] a checking unit configured to check the current voice frame number based on the attribute information and calculate a cross-correlation value of the recorded voice frame signal and the to-be-played voice frame signal, wherein the cross-correlation value is used to indicate a similarity degree of the recorded voice frame signal and the to-be-played voice frame signal;
[0016] a determining unit configured to determine a current call state of the first terminal according to a checking result of the current voice frame number and the cross-correlation value.
[0017] Optionally, in some embodiments, the checking unit can be specifically configured to extract a basic voice frame number of the voice frame signal from the attribute information; compare the basic voice frame number with the current voice frame number when the basic voice frame number and the current voice frame number exceed a preset frame number threshold; and determine the checking result of the current voice frame number based on a comparison result.
[0018] Optionally, in some embodiments, the checking unit can be specifically configured to calculate a ratio of the received voice frame number and the basic voice frame number to obtain a first frame number ratio; calculate a ratio of the received voice frame number and the played voice frame number to obtain a second frame number ratio; and determine the checking result of the current voice frame number based on a comparison result, including: determining that the current voice frame number passes the check when the first frame number ratio and the second frame number ratio do not exceed a preset frame number ratio threshold.
[0019] Optionally, in some embodiments, the verification unit can be specifically configured to obtain a preset time delay parameter set between the to-be-played voice frame signal and the recorded frame signal; filter out the to-be-played voice frame signal corresponding to each preset time delay parameter in the preset time delay parameter set from the to-be-played voice frame signal, to obtain a candidate to-be-played voice frame signal set corresponding to each recorded frame signal; calculate a candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set, respectively; and determine the current call state of the first terminal according to the verification result of the current voice frame number and the cross-correlation value.
[0020] Optionally, in some embodiments, the verification unit can be specifically configured to calculate a power spectrum value of a frequency point in each recorded frame signal to obtain a first power spectrum value; calculate a power spectrum value of a frequency point in each to-be-played voice frame signal in the candidate to-be-played voice frame signal set to obtain a second power spectrum value; and fuse the first power spectrum value and the second power spectrum value to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set.
[0021] Optionally, in some embodiments, the verification unit can be specifically configured to calculate a mean value of the first power spectrum value to obtain a first power spectrum mean value, and calculate a mean value of the second power spectrum value to obtain a second power spectrum mean value; calculate a difference value between the first power spectrum value and the first power spectrum mean value to obtain a first power spectrum difference value, and calculate a difference value between the second power spectrum value and the second power spectrum mean value to obtain a second power spectrum difference value; and fuse the first power spectrum difference value and the second power spectrum difference value to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set.
[0022] Optionally, in some embodiments, the determination unit can be specifically configured to determine a playing state of the to-be-played voice frame signal based on the candidate cross-correlation value; and when the playing state is normal playing and the current voice frame number passes the verification, determine that the current call state of the first terminal is normal call.
[0023] Optionally, in some embodiments, the determination unit can be specifically configured to compare the candidate cross-correlation value with a preset cross-correlation threshold value; count a number of candidate cross-correlation values exceeding the preset cross-correlation threshold value under each preset time delay parameter in a comparison result to obtain a basic statistical number set; and determine the playing state of the to-be-played voice frame signal based on the basic statistical number set.
[0024] Optionally, in some embodiments, the determining unit can be specifically configured to calculate a mean value of the basis statistical quantities in the basis statistical quantity set, to obtain a basis statistical quantity mean value; filter out a basis statistical quantity with the largest value in the basis statistical quantity set, to obtain a target statistical quantity; when the target statistical quantity is the same as a preset quantity threshold, calculate a quantity ratio between the target statistical quantity and the basis statistical quantity mean value; when the quantity ratio exceeds a preset quantity ratio threshold, determine that the playing state of the to-be-played voice frame signal is normal playing.
[0025] Optionally, in some embodiments, the determining unit can be specifically configured to generate a normal call prompt according to the current call state; and send the prompt to the first terminal, so that the first terminal prompts the current call state according to the prompt.
[0026] Optionally, in some embodiments, the analyzing unit can be specifically configured to decode the voice frame signal to obtain a decoded voice frame signal; post-process the decoded voice frame signal to obtain a to-be-played voice frame signal; and perform voice frame number statistics on the decoded voice frame signal and the to-be-played voice frame signal based on the attribute information, to obtain a current voice frame number of the voice frame signal.
[0027] Optionally, in some embodiments, the analyzing unit can be specifically configured to extract a voice frame identifier of the voice frame signal from the attribute information; perform statistics on the voice frame identifier in the decoded voice frame signal, to obtain a received voice frame number of the decoded voice frame signal; perform voice detection on the to-be-played voice frame signal, to obtain a played voice frame number of the to-be-played voice frame signal; and take the received voice frame number and the played voice frame number as the current voice frame number of the voice frame signal.
[0028] Optionally, in some embodiments, the call state detection apparatus can further include a sending unit, which can be specifically configured to, when a sound signal is detected, collect a current generated sound signal, and perform frame processing on the sound signal to obtain a plurality of sound frame signals; perform encoding processing on the sound frame signals to generate target call information; and send the target call information to a second terminal, so that the second terminal detects a call state based on the target call information.
[0029] Optionally, in some embodiments, the sending unit can be specifically used for voice detection on the sound frame signal, and generating a target voice frame identifier based on a voice detection result; pre-processing the sound frame signal, and performing voice coding processing on the pre-processed sound frame signal to obtain a target voice frame signal; counting the target voice frame identifier in the target voice frame signal to obtain a target voice frame number, and taking the target voice frame identifier and the target voice frame number as target attribute information of the target voice frame signal; fusing the target voice frame signal and the target attribute information of the target voice frame signal to obtain target call information.
[0030] In addition, an electronic device is further provided in the embodiment of the present application, which comprises a processor and a memory, the memory stores an application program, and the processor is used to run the application program in the memory to realize the call state detection method provided in the embodiment of the present application.
[0031] In addition, a computer readable storage medium is further provided in the embodiment of the present application, which stores a plurality of instructions, the instructions are suitable for being loaded by a processor to execute the steps in any one of the call state detection methods provided in the embodiment of the present application.
[0032] After receiving the current call information sent by the first terminal, the attribute information in the current call information is used to analyze the voice frame signal in the current call information to obtain a to-be-played voice frame signal and a current voice frame number of the voice frame signal, then the to-be-played voice frame signal is played, and the to-be-played voice frame signal after playing is collected to obtain a recorded frame signal corresponding to the to-be-played voice frame signal, then the current voice frame number is verified based on the attribute information, and a cross-correlation value of the recorded frame signal and the to-be-played voice frame signal is calculated, then the current voice frame number and the cross-correlation value are used to determine the current call state with the first terminal; since the voice signal collected can be identified, and the voice frame number in the whole process of instant messaging is verified by the listener, and in addition, whether the sound is successfully played is determined by the listener through the cross-correlation calculation of the to-be-played voice frame signal and the recorded frame signal, so that the voice detection in the whole process of call can be realized, and therefore the accuracy of call state detection can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0034] Figure 1 is a scene schematic diagram of a call state detection method provided by an embodiment of the present application;
[0035] Figure 2 is a flow schematic diagram of a call state detection method provided by an embodiment of the present application;
[0036] Figure 3 is a schematic diagram of judging whether a recording function is invalid or not provided by an embodiment of the present application;
[0037] Figure 4 is a page schematic diagram of a call page displayed by a first terminal provided by an embodiment of the present application;
[0038] Figure 5 is a schematic diagram of a detection process of call state detection provided by an embodiment of the present application;
[0039] Figure 6 is another flow schematic diagram of a call state detection method provided by an embodiment of the present application;
[0040] Figure 7 is a structure schematic diagram of a call state detection apparatus provided by an embodiment of the present application;
[0041] Figure 8 is another structure schematic diagram of a call state detection apparatus provided by an embodiment of the present application;
[0042] Figure 9 is a structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0044] The embodiments of the present application provide a call state detection method, apparatus and computer readable storage medium. The call state detection apparatus can be integrated in an electronic device, which can be a server, a terminal or other devices.
[0045] The server can be a stand-alone physical server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, and the like, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0046] For example, referring to Figure 1 For example, referring to
[0047] The call state is used to indicate whether each party in the instant messaging application can receive and play voice information sent by other communication terminals when they conduct audio and video calls through their respective communication terminals. It can include normal call and abnormal call. The normal call means that the sender can normally send call information, the listener can normally receive and play the call information, so that the current call can continue. The abnormal call means that the sender cannot normally send call information, or the listener cannot normally receive or play the call information. In the instant messaging application, the call state has a great influence on the call effect.
[0048] The following will be described in detail. It should be noted that the order of the following embodiments is not limited to the preferred order of the embodiments.
[0049] The embodiment will be described from the perspective of a call state detection device, which can be integrated in an electronic device, which can be a server, a terminal, or the like. The terminal can include a tablet computer, a notebook computer, a personal computer (PC), a wearable device, a virtual reality device, or other smart devices that can detect a call state.
[0050] A call state detection method includes:
[0051] Receiving current call information sent by a first terminal, the current call information including a voice frame signal and attribute information of the voice frame signal, analyzing the voice frame signal according to the attribute information to obtain a to-be-played voice frame signal and a current voice frame number of the voice frame signal, playing the to-be-played voice frame signal, collecting the played to-be-played voice frame signal to obtain a recorded voice frame signal corresponding to the to-be-played voice frame signal, verifying the current voice frame number based on the attribute information, and calculating a cross-correlation value of the recorded voice frame signal and the to-be-played voice frame signal, the cross-correlation value being used to indicate a similarity between the recorded voice frame signal and the to-be-played voice frame signal, and determining a current call state with the first terminal according to a verification result of the current voice frame number and the cross-correlation value.
[0052] As shown in Figure 2 , a specific process of the call state detection method is as follows:
[0053] 101. Receiving current call information sent by a first terminal.
[0054] The current call information can include a voice frame signal and attribute information of the voice frame signal. The voice frame signal can be a voice signal detected by performing voice activity detection on a recorded voice frame signal. The recorded voice frame signal can be a sound emitted by a user and collected by a microphone of the terminal, which is divided into equal-length frame signals according to a certain window size.
[0055] The attribute information of the voice frame signal can be information describing attributes of the voice frame signal, which can include a voice frame identifier of the voice frame signal and a basic voice frame number of the voice frame signal. The voice frame number can be a number of recorded voice frame signals that are voice frame signals.
[0056] The way of receiving the current call information sent by the first terminal can be various, which can be as follows:
[0057] For example, an information transmission channel is established with the first terminal through the network, and current call information sent by the first terminal is received through the information transmission channel, or, when the current call information is sensitive information, encrypted information of the current call information sent by the first terminal is also received, the encrypted information is decrypted based on a preset encryption protocol to obtain the current call information, or, when the current call information has a large memory or a large quantity, the current call information is indirectly obtained from the first terminal.
[0058] The current call information indirectly obtained from the first terminal can be obtained in various ways, for example, when the first terminal sends the current call information to a server or a third terminal, the current call information corresponding to the first terminal sent by the server or the third terminal is received, or, a call request sent by the first terminal is received, the call request carries a storage address of the current call information, and the current call information is obtained from the first terminal, the third terminal or a call information database based on the storage address.
[0059] 102. Analyze the voice frame signal based on the attribute information to obtain the to-be-played voice frame information and a current voice frame number of the voice frame signal.
[0060] The to-be-played voice frame signal can be understood as a voice frame signal that needs to be played after decoding and post-processing of the voice frame signal.
[0061] The current voice frame number can include a received voice frame number and a played voice frame number, the received voice frame number can be the frame number of the received voice frame signal, and the played voice frame number can be the frame number of the to-be-played voice frame signal.
[0062] The voice frame signal can be analyzed in various ways, for example, the voice frame signal can be decoded to obtain a decoded voice frame signal, the decoded voice frame signal can be post-processed to obtain the to-be-played voice frame signal, the decoded voice frame signal and the to-be-played voice frame signal can be counted based on the attribute information to obtain the current voice frame number of the voice frame signal.
[0063] For example, the voice frame signal is decoded to obtain a decoded voice frame signal, the decoded voice frame signal is post-processed to obtain the to-be-played voice frame signal, the decoded voice frame signal and the to-be-played voice frame signal are counted based on the attribute information to obtain the current voice frame number of the voice frame signal.
[0064] The voice frame signal can be decoded in various ways, for example, the voice frame signal can be decoded according to a corresponding decoding strategy based on the encoding format of the voice frame signal, or, a terminal identifier of the first terminal can be obtained, a target decoding protocol corresponding to the terminal identifier is filtered out in a preset decoding protocol, and the voice frame signal is decoded based on the target decoding protocol to obtain a decoded voice frame signal.
[0065] After the speech frame signal is decoded, the decoded speech frame signal can be post-processed. The post-processing can be understood as a processing mode corresponding to the pre-processing of the recording signal by the first terminal. The main role of the post-processing is to convert the decoded speech frame signal into a speech frame signal suitable for playing by the call state detection device. The post-processing mode can be various, such as adapting to the current playing device to be played, or compressing or amplifying the decoded speech frame signal, or setting the playing parameters of the decoded speech frame signal, and the like, so as to convert the decoded speech frame signal into a to-be-played speech frame signal that can be directly played.
[0066] After the post-processing of the decoded speech frame signal, the decoded speech frame signal and the to-be-played speech frame signal can be counted based on the attribute information, so as to obtain the current speech frame number of the speech frame signal. The counting mode can be various, such as extracting the speech frame identifier of the speech frame signal in the attribute information, counting the speech frame identifier in the decoded speech frame signal to obtain the received speech frame number of the decoded speech frame signal, performing speech detection on the to-be-played speech frame signal to obtain the played speech frame number of the to-be-played speech frame signal, and taking the received speech frame number and the played speech frame number as the current speech frame number of the speech frame signal.
[0067] Among them, the counting mode of the speech frame identifier in the decoded speech frame signal can be various, such as detecting the decoded speech frame signal containing the speech frame identifier in the decoded speech frame signal to obtain the target decoded speech frame signal, and counting the number of the target decoded speech frame signal, so as to obtain the received speech frame number, or the number of the decoded speech frame signal containing the speech frame identifier can also be recognized in the decoded speech frame signal, so as to count the received speech frame number.
[0068] Among them, the speech detection mode of the to-be-played speech frame signal can be various, such as using a voice activity detection (vad) algorithm to detect whether the to-be-played speech frame signal is a speech signal, so as to obtain the detection result of each to-be-played speech frame signal, and based on the detection result, the number of to-be-played speech frame signals that are speech signals is counted, so as to obtain the played speech frame number of the to-be-played speech frame signal. Thus, it can be found that the played speech frame number is used to indicate the number of to-be-played speech frame signals of the type of speech frame signal.
[0069] 103, playing the to-be-played speech frame signal and collecting the played to-be-played speech frame signal to obtain the recording frame signal corresponding to the to-be-played speech frame signal.
[0070] The recording frame signal can be collected by a microphone of the call state detection device when playing the to-be-played voice frame signal, and then windowed according to a certain window size to obtain the frame signal.
[0071] The method for collecting the recording frame signal corresponding to the to-be-played voice frame signal can be various, and can be as follows:
[0072] For example, the to-be-played voice frame signal is played through a loudspeaker in the call state detection device, and the to-be-played voice frame signal played is recorded through a microphone at the same time to obtain a recording signal. The recording signal is then framed according to the window length of each frame in the to-be-played voice frame signal, so as to obtain a recording frame signal with the same window length as the to-be-played voice frame signal.
[0073] It should be noted that after the to-be-played voice frame signal is converted into sound through electro-acoustic conversion, the sound is conducted through the air and then collected by the microphone, that is, an echo signal of the to-be-played voice frame signal, which is finally mixed into the recording signal to obtain the recording frame signal. The echo delay is relatively stable.
[0074] 104. Based on the attribute information, the current voice frame number is verified, and the cross-correlation value of the recording frame signal and the to-be-played voice frame signal is calculated.
[0075] The cross-correlation value is used to indicate the similarity between the recording frame signal and the to-be-played voice frame signal. The greater the cross-correlation value, the more similar the to-be-played voice frame signal and the recording frame signal. On the contrary, the smaller the cross-correlation value, the less relevant the to-be-played voice frame signal and the recording frame signal.
[0076] The method for verifying the current voice frame number and calculating the cross-correlation value of the recording frame signal and the to-be-played voice frame signal can be various, and can be as follows:
[0077] S1. Based on the attribute information, the current voice frame number is verified.
[0078] The method for verifying the current voice frame number can be various, and can be as follows:
[0079] For example, the basic voice frame number of the voice frame signal is extracted from the attribute information. When the basic voice frame number and the current voice frame number exceed a preset frame number threshold, the basic voice frame number and the current voice frame number are compared, and based on the comparison result, the verification result of the current voice frame number is determined.
[0080] The base voice frame number can be the number of voice frame signals contained in the current call information sent by the sender. In addition, since the current call information needs to be normally played, the base voice frame number and the current voice frame number both need to be greater than 0, so the preset frame number threshold can be 0 or other positive integers.
[0081] The current voice frame number can include the received voice frame number and the played voice frame number, so there can be multiple ways to compare the base voice frame number with the current voice frame number, such as calculating the ratio of the received voice frame number to the base voice frame number to obtain a first frame number ratio, calculating the ratio of the received voice frame number to the played voice frame number to obtain a second frame number ratio, and determining that the current voice frame number passes the verification when the first frame number ratio and the second frame number ratio do not exceed a preset frame number ratio threshold.
[0082] It can be found that the purpose of verifying the current voice frame number is to determine whether there is packet loss or other abnormal conditions when receiving and playing voice frame signals, which causes the voice frame signals to be unable to be normally played. In addition, taking the base voice frame number as Vcnt_0, the received voice frame number as Vcnt_1, and the played voice frame number as Vcnt_2, the preset frame number threshold as 0, and the preset frame number ratio threshold as 1.25 as an example, the verification of the current voice frame number can be summarized as needing to meet the following three verification conditions:
[0083] (1) Vcnt_0, Vcnt_1, and Vcnt_2 are all greater than 0;
[0084] (2) Verification of the base voice frame number and the received voice frame number, i.e., whether Vcnt_0 / Vcnt_1 is less than 1.25. If it is greater, the condition is not met;
[0085] (3) Verification of the played voice frame number and the received frame number of the listener, i.e., whether Vcnt_1 / Vcnt_2 is less than 1.25. If it is greater, the condition is not met.
[0086] When the above three verification conditions are met, it can be determined that the verification result of the current voice frame number is that the verification passes, and the lost voice frame signals in the transmission and playing process of the voice frame signals are within a normal range, meeting the standard of normal call. In addition, the preset frame number ratio thresholds corresponding to the first frame number ratio and the second frame number ratio can be the same or different.
[0087] S2, calculate the cross-correlation value of the recorded frame signal and the to-be-played voice frame signal.
[0088] For example, a set of preset time delay parameters between the to-be-played voice frame signal and the recorded frame signal can be acquired, each to-be-played voice frame signal corresponding to each preset time delay parameter in the to-be-played voice frame signal is screened, a set of candidate to-be-played voice frame signals corresponding to each recorded frame signal is obtained, and a candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the set of candidate to-be-played voice frame signals corresponding to the recorded frame signal is calculated.
[0089] The set of preset time delay parameters can be a set of time delay parameters between the to-be-played voice frame signal and the recorded frame signal that is preset. Since the to-be-played voice frame signal is converted into sound for playing, the sound is conducted through the air and is collected by a sound collecting device such as a microphone into an echo signal, and finally mixed into the recorded frame signal, so that the echo delay is relatively fixed. Therefore, the recorded frame signal of the current frame and the to-be-played voice frame signal at a position k0 (a corresponding echo delay position) in the playing buffer have a high cross-correlation P(i, k0), and the cross-correlation values P(i, k) between the recorded frame signal and the to-be-played voice frame signal at other positions in the playing buffer are small. However, in addition to the echo signal of the to-be-played voice frame signal, the recorded frame signal also includes other sounds (for example, a near-end human voice and other environmental sounds), and these remaining sounds will interfere with the cross-correlation value between the to-be-played voice frame signal and the recorded frame signal, so that some statistical methods are needed to filter. Therefore, the time delay parameter is introduced. The time delay parameter can be understood as a statistical variable with a size of M. The statistical variable includes a plurality of possible time delay positions. The time delay position means that the echo signal corresponding to the voice frame signal of the ith frame is the K value in the to-be-played voice frame signal of the (i-k)th frame, that is, the preset time delay parameter is K, and k∈M.
[0090] The to-be-played voice frame signal corresponding to each preset time delay parameter in the set of preset time delay parameters can be screened in multiple ways. For example, taking the preset time delay parameters K in the set of preset time delay parameters as 2, 3, and 4, the candidate to-be-played voice frame signal corresponding to the 5th recorded frame signal can be the first 2, 3, or 4 frames of the 5th frame in the playing buffer. That is, the candidate to-be-played voice frame signal corresponding to the 5th recorded frame signal can be the 3rd frame, the 2nd frame, or the 1st frame. The three candidate to-be-played voice frame signals are taken as the set of candidate to-be-played voice frame signals corresponding to the 5th recorded frame signal. In this way, the set of candidate to-be-played voice frame signals corresponding to each recorded frame signal can be obtained.
[0091] After the candidate to-be-played speech frame signal set corresponding to each recording frame signal is screened out, the candidate cross-correlation value between each to-be-played speech frame in the candidate to-be-played speech frame signal set corresponding to each recording frame signal can be calculated. The candidate cross-correlation value can be calculated in multiple ways. For example, the power spectrum value of a frequency point in each recording frame signal can be calculated, the power spectrum value of a frequency point in each to-be-played speech frame signal in the candidate to-be-played speech frame signal set can be calculated, a second power spectrum value can be obtained, and the first power spectrum value and the second power spectrum value can be fused to obtain the candidate cross-correlation value between each recording frame signal and the corresponding candidate to-be-played speech frame signal.
[0092] The power spectrum can be calculated in multiple ways. For example, the recording frame signal of the to-be-played speech frame of the current frame can be subjected to Fourier transform, then the power spectrum X(i, j) and D(i, j) of the i th frame and the j th frequency point can be calculated, and thus the first power spectrum value and the second power spectrum value can be obtained. The first power spectrum value X(i, j) of the to-be-played speech frame is buffered through a ring buffer (maximum buffer M frame playback signal power spectrum) for cross-correlation calculation.
[0093] After the first power spectrum value and the second power spectrum value are calculated, the first power spectrum value and the second power spectrum value can be fused. The fusion can be performed in multiple ways. For example, the mean value of the first power spectrum value can be calculated to obtain a first power spectrum mean value, the mean value of the second power spectrum value can be calculated to obtain a second power spectrum mean value. The difference between the first power spectrum value and the first power spectrum mean value is calculated to obtain a first power spectrum difference value, and the difference between the second power spectrum value and the second power spectrum mean value is calculated to obtain a second power spectrum difference value. The first power spectrum difference value and the second power spectrum difference value are fused to obtain the candidate cross-correlation value between each recording frame signal and each to-be-played speech frame signal in the corresponding candidate to-be-played speech frame signal set. The specific process can be shown in formula (1):
[0094]
[0095] wherein P(i, k) is the cross-correlation value between the i th recording frame signal and the to-be-played speech frame signal in the playback buffer that is k frames away from the to-be-played speech frame signal of the i th frame, k is a preset time delay parameter, k corresponds to the to-be-played speech frame signal in the playback buffer that is k frames away from the current i th frame, X(i-k, j) is the first power spectrum value of the j th frequency point in the i-k th to-be-played speech frame signal, is the first power spectrum mean value of the i-k th to-be-played speech frame playback signal on the frequency domain from the frequency point N1 to the frequency point N2, is the second power spectrum mean value of the i th recording frame signal on the frequency domain from the frequency point N1 to the frequency point N2, and D(i, j) is the second power spectrum value of the j th frequency point in the i th recording frame signal.
[0096] 105. Determine the current call state with the first terminal according to the check result of the current number of voice frames and the cross-correlation value.
[0097] For example, the playing state of the to-be-played voice frame signal can be determined based on the candidate cross-correlation value. When the playing state is normal playing and the current number of voice frames passes the check, it is determined that the current call state with the first terminal is normal call.
[0098] The playing state is used to indicate whether the to-be-played voice frame signal is normally played by producing sound through a loudspeaker or the like. Therefore, it can include normal playing and abnormal playing. The so-called normal playing can be understood as that the to-be-played voice frame signal is converted into sound through electro-acoustic conversion and played normally. When the playing state is normal, the microphone or the like can collect the corresponding recording signal. When the playing state is abnormal, the to-be-played voice frame signal is abnormal during electro-acoustic conversion or sound playing. At this time, the microphone or the like cannot accurately collect the corresponding recording signal. The playing state of the to-be-played voice frame signal can be determined in various ways based on the candidate cross-correlation value. For example, the candidate cross-correlation value can be compared with a preset cross-correlation threshold. The number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter is counted in the comparison result to obtain a basic statistical number set. The playing state of the to-be-played voice frame signal is determined based on the basic statistical number set.
[0099] The number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter can be counted in various ways in the comparison result. For example, taking the preset time delay parameter k=2 as an example, the number of candidate cross-correlation values exceeding the preset cross-correlation threshold is counted between each recording frame signal and the corresponding candidate to-be-played voice frame signal under this time delay parameter, so that the basic statistical number under this time delay parameter can be obtained. By analogy, the basic statistical number corresponding to each time delay parameter can be obtained. The basic statistical numbers are combined to obtain the basic statistical number set.
[0100] After the basic number set is counted, the playing state of the to-be-played voice frame signal can be determined based on the basic statistical number set. The playing state can be determined in various ways. For example, the mean value of the basic statistical numbers in the basic statistical number set is calculated to obtain a basic statistical number mean value. The largest basic statistical number in the basic statistical number set is selected to obtain a target statistical number. When the target statistical number is equal to a preset number threshold, the number ratio between the target statistical number and the basic statistical number mean value is calculated. When the number ratio exceeds a preset number ratio, it is determined that the playing state of the to-be-played voice frame signal is normal playing.
[0101] Wherein, since the recording signal includes other sounds (such as the voice of the near-end person, other environmental sounds) in addition to the echo signal of the to-be-played voice signal, these remaining sounds will cause certain interference to the cross-correlation value between the to-be-played voice frame signal and the recording frame signal, therefore, the purpose of determining the playing state of the to-be-played voice frame signal based on the basic statistical quantity set is to filter these interferences, so as to more accurately determine the preset time delay parameter between the to-be-played voice frame signal and the recording frame signal, and then judge the playing state of the to-be-played voice frame signal. Here, taking the statistical variable C(k) of the preset time delay parameter set M as an example, the filtering method mainly calculates the candidate cross-correlation degree of the continuous recording frame signal, and each recording frame signal will obtain M P(i, k) values (candidate cross-correlation values). Whenever the P(i, k1) value is greater than the preset cross-correlation threshold THRD_P (for example, 0.5), the statistical variable C(k1) is added by 1, and the rest C(k) remains unchanged, k≠k1. When the maximum value of the M C(k) statistical variables is equal to the preset quantity threshold THRD_L, such as C(k2)=THRD_L, THRD_L=200; and satisfies THRD_C, here is the average value of C(k), it is determined that the to-be-played voice frame signal is normally played, which means that the recording function is normal and effective, otherwise, it is determined that the to-be-played voice frame signal is not normally played, which means that the recording function is invalid. Thus, it can be found that the process of judging the recording function failure by calculating the cross-correlation value can be as shown in Figure 3 , first, the power spectrum of the to-be-played voice frame signal and the recording frame signal is calculated respectively to obtain the first power spectrum value and the second power spectrum value, then the first power spectrum value of the to-be-played voice frame signal is cached, then the cross-correlation value is calculated based on the first power spectrum value and the second power spectrum value, and then the cross-correlation value is counted to judge the playing state of the to-be-played voice frame signal and the failure of the recording function based on the counting result.
[0102] Optionally, when it is determined that the current call state with the first terminal is a normal call, prompt information of the current call state can also be sent to the first terminal, for example, prompt information of the normal call can be generated according to the current call state, and the prompt information is sent to the first terminal, so that the first terminal prompts the current call state according to the prompt information.
[0103] Wherein, there are many ways to prompt the call state, for example, when the first terminal is a hardware device product, light prompting can be performed on the prompt light on the first terminal, or vibration prompting can be performed on the vibration motor on the first terminal, or the call state can also be displayed on the call page.
[0104] The call state display manner on the call page can be various, for example, a prompt information of normal call can be directly displayed on the call page, or a target call icon corresponding to the normal call can be filtered out from the call state icon according to the prompt information and displayed. The call icon is mainly used to prompt that the current call state of the speaking party corresponding to the first terminal is normal call. The form of the call icon can be various. Taking a small horn as an example, the call page displayed on the first terminal can be as shown in FIG. 8. Figure 4
[0105] Optionally, when the call terminal detection apparatus detects that the user makes a sound, the current sound signal can also be collected and processed. The processing manner can be various, for example, the current generated sound signal can be collected and subjected to frame processing to obtain a plurality of sound frame signals. The sound frame signals are subjected to encoding processing to generate target call information. The target call information is sent to the second terminal, so that the second terminal detects the call state based on the target call information.
[0106] The encoding processing of the sound frame signal to generate the target call information can be various, for example, the sound frame signal can be subjected to voice detection, and a target voice frame identifier is generated based on the voice detection result. The sound frame signal is subjected to pre-processing, and the pre-processed sound frame signal is subjected to voice encoding processing to obtain a target voice frame signal. The target voice frame identifier is counted in the target voice frame signal to obtain a target voice frame number. The target voice frame identifier and the target voice frame number are taken as target attribute information of the target voice frame signal. The target voice frame signal and the target attribute information of the target voice frame signal are fused to obtain the target call information.
[0107] The voice detection of the sound frame signal can be various, for example, each sound frame signal can be subjected to vad algorithm detection, and the sound signal with vad of 1 is filtered out from the sound frame signal as a target sound frame signal corresponding to voice, so that the target sound frame signal is taken as the voice detection result.
[0108] After voice detection, the target voice frame identifier can be generated based on the voice detection result in various manners, for example, each target sound frame signal is numbered to obtain a target voice frame identifier corresponding to each target sound frame signal, or each target sound frame signal is sorted, and the sorting result is converted into an identifier of each target sound frame signal to obtain the target voice frame identifier.
[0109] After the target speech frame identifier is generated, the pre-processed sound frame signal can be subjected to speech encoding processing. The speech encoding processing can be performed in various ways, such as using a preset speech encoding protocol or strategy to encode the pre-processed sound frame signal to obtain an initial speech frame signal, and identifying the initial speech frame signal through the target speech frame identifier to obtain a target speech frame signal.
[0110] The initial speech frame signal can be identified through the target speech frame identifier in various ways, such as screening a basic speech frame signal corresponding to the target speech frame identifier from the initial speech frame signal, adding the target speech frame identifier to the basic speech frame identifier, and thus obtaining the target speech frame signal.
[0111] After the pre-processed sound frame signal is subjected to speech encoding processing, the target speech frame identifier can be counted in the target speech frame signal. The counting can be performed in various ways, such as screening a target speech frame signal having the target speech frame identifier from the target speech frame signal, and counting the number of the screened target speech frame signal, so as to obtain a target speech frame number. The target speech frame number can indicate the number of the target speech frame signal included in the current call information sent by the first terminal to the call state detection device, and further determine whether the second terminal completely receives and plays the target speech frame signal.
[0112] After the target call information is sent to the second terminal, target prompt information corresponding to the call state of the target call information can be received from the second terminal, such as sending the target call information to the second terminal, detecting the call state according to the target call information, and when the second terminal detects that the call state is normal call, the target prompt information corresponding to the normal call returned by the second terminal can be received, and the call state can be prompted based on the target prompt information. The prompt call state can be prompted in the same way as the first terminal, which will not be described here. In addition, the second terminal can be the same as or different from the first terminal.
[0113] In this scheme, it can be found that the completion of the call state detection requires the participation of both parties of the instant messaging service, i.e., the speaker and the listener jointly complete the transmission and playing of the current call information, and the current call state is detected through the transmission and playing results. The whole process of the call state detection can be as follows Figure 5As shown, in the speaker, the recording signal is obtained through the microphone of the first terminal corresponding to the speaker, the recording signal is divided into equal length frame signals according to a certain window length (for example, 20ms), each frame recording signal is judged whether it is a speech signal through the vad algorithm, when vad = 1, the frame is a speech signal, and the speech frame signal is identified, the speech frame signal is bundled with the speech coding data and sent to the listener, at the same time, the speaker will count the number of speech frames identified in a certain time period to obtain the basic speech frame number, and the encoded speech frame signal, the basic speech frame number and the speech frame identification are transmitted to the terminal of the listener as the target call information through the channel. In the listener, the speech frame signal is decoded, and the speech frame identification is detected, so as to count the received speech frame number, the decoded speech frame signal is post-processed to obtain the to-be-played speech frame signal, then, the to-be-played speech frame signal is detected through the vad algorithm, when vad = 1, the to-be-played speech frame signal is counted, so as to count the played speech frame number. The played speech frame number, the received speech frame number and the basic speech frame number are checked. For the to-be-played speech frame signal, the to-be-played speech frame signal is played through the loudspeaker, the sound corresponding to the to-be-played speech frame signal is collected to obtain the recording frame signal, the cross-correlation value of the recording frame signal and the to-be-played speech frame signal is calculated, and the cross-correlation value is checked. When the speech frame number check and the cross-correlation value check pass, the prompt information is sent to the first terminal, so that the first terminal displays the small loudspeaker state identification on the call page.
[0114] As can be seen from the above, after receiving the current call information sent by the first terminal, the speech frame signal in the current call information is parsed according to the attribute information in the current call information to obtain the to-be-played speech frame signal and the current speech frame number of the speech frame signal, then, the to-be-played speech frame signal is played, and the to-be-played speech frame signal is collected to obtain the recording frame signal corresponding to the to-be-played speech frame signal, then, based on the attribute information, the current speech frame number is checked, and the cross-correlation value of the recording frame signal and the to-be-played speech frame signal is calculated, then, according to the checking result of the current speech frame number and the cross-correlation value, the current call state with the first terminal is determined; Since the scheme can identify the collected speech signal, and then the speech frame number of the speech signal is checked in the whole process of instant messaging through the listener, in addition, whether the sound is successfully played is determined through the cross-correlation calculation of the to-be-played speech frame signal and the recording frame signal in the listener, so that the sound detection of the whole call process can be realized, and therefore, the accuracy of the call state detection can be improved.
[0115] According to the method described in the above embodiment, the following will be further described in detail by way of example.
[0116] In the embodiment, the call state detection device is integrated in an electronic device, which is a terminal. In order to distinguish the first terminal of the speaker, the terminal can be taken as an example for illustration.
[0117] As shown in the figure, a call state detection method includes the following specific processes: Figure 6
[0118] 201. The target terminal receives the current call information sent by the first terminal.
[0119] For example, the target terminal establishes an information transmission channel with the first terminal through a network, receives the current call information sent by the first terminal through the information transmission channel, or, when the current call information is sensitive information, also receives the encrypted information of the current call information sent by the first terminal, decrypts the encrypted information based on a preset encryption protocol to obtain the current call information, or, when the first terminal sends the current call information to a server or a third-party terminal, receives the current call information corresponding to the first terminal sent by the server or the third-party terminal, or, receives a call request sent by the first terminal, the call request carrying a storage address of the current call information, and obtains the current call information from the first terminal, the third-party terminal or a call information database based on the storage address.
[0120] 202. The target terminal decodes the voice frame signal to obtain a decoded voice frame signal.
[0121] For example, the target terminal can decode the voice frame signal according to the encoding format of the voice frame signal, or can obtain the terminal identifier of the first terminal, filter out the target decoding protocol corresponding to the terminal identifier in a preset decoding protocol, and decode the voice frame signal based on the target decoding protocol to obtain the decoded voice frame signal.
[0122] 203. The target terminal post-processes the decoded voice frame signal to obtain a to-be-played voice frame signal.
[0123] For example, the target terminal can adapt to the current playback device, or compress or amplify the decoded voice frame signal, or set playback parameters for the decoded voice frame signal, etc., so as to convert the decoded voice frame signal into the to-be-played voice frame signal which can be directly played.
[0124] 204. The target terminal counts the number of voice frames of the decoded voice frame signal and the to-be-played voice frame signal based on the attribute information to obtain the current number of voice frames of the voice frame signal.
[0125] For example, the target terminal extracts the speech frame identifier of the speech frame signal in the attribute information, detects the decoded speech frame signal containing the speech frame identifier in the decoded speech frame signal, obtains the target decoded speech frame signal, and counts the number of the target decoded speech frame signal, so as to obtain the received speech frame number, or the number of the decoded speech frame signal containing the speech frame identifier in the decoded speech frame signal is also identified, so as to count the received speech frame number. The vad algorithm is used to detect whether the to-be-played speech frame signal is a speech signal, so as to obtain the detection result of each to-be-played speech frame signal, and based on the detection result, the number of to-be-played speech frame signals that are speech signals is counted, so as to obtain the played speech frame number of the to-be-played speech frame signal. The received speech frame number and the played speech frame number are taken as the current speech frame number of the speech frame signal.
[0126] 205、The target terminal plays the to-be-played speech frame signal, and collects the played to-be-played speech frame signal to obtain the recorded frame signal corresponding to the to-be-played speech frame signal.
[0127] For example, the target terminal converts the to-be-played speech frame signal into sound through electro-acoustic conversion through the loudspeaker, and then the sound is conducted through the air and collected by the microphone again, that is, the echo signal of the to-be-played speech frame signal, and finally mixed into the recorded frame signal, so as to obtain the recorded frame signal.
[0128] 206、The target terminal verifies the current speech frame number based on the attribute information.
[0129] For example, the target terminal extracts the basic speech frame number of the speech frame signal in the attribute information, and when the basic speech frame number and the current speech frame number exceed the preset frame number threshold, the ratio of the received speech frame number to the basic speech frame number is calculated to obtain a first frame number ratio, and the ratio of the received speech frame number to the played speech frame number is calculated to obtain a second frame number ratio, and when the first frame number ratio and the second frame number ratio do not exceed the preset frame number ratio threshold, it is determined that the current speech frame number verification is passed.
[0130] 207、The target terminal calculates the cross-correlation value of the recorded frame signal and the to-be-played speech frame signal.
[0131] For example, the target terminal can obtain a preset time delay parameter set between the to-be-played voice frame signal and the recorded frame signal, filter out the to-be-played voice frame signal corresponding to each preset time delay parameter in the to-be-played voice frame signal, obtain a candidate to-be-played voice frame signal set corresponding to each recorded frame signal, perform Fourier transform on the to-be-played voice frame signal and the recorded frame signal of the current frame, then calculate the power spectrum X(i, j) and D(i, j) of the i-th frame and the j-th frequency point, so as to obtain the first power spectrum value and the second power spectrum value. The first power spectrum value X(i, j) of the to-be-played voice frame signal is cached through a ring buffer (maximum cache M frame play signal power spectrum) for cross-correlation degree calculation. The mean value of the first power spectrum value is calculated to obtain the first power spectrum mean value, and the mean value of the second power spectrum value is calculated to obtain the second power spectrum mean value. The difference between the first power spectrum value and the first power spectrum mean value is calculated to obtain the first power spectrum difference value, and the difference between the second power spectrum value and the second power spectrum mean value is calculated to obtain the second power spectrum difference value. The first power spectrum difference value and the second power spectrum difference value are fused to obtain a candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set, which can be specifically shown as formula (1).
[0132] 208、The target terminal determines the current call state with the first terminal according to the check result of the current voice frame number and the cross-correlation value.
[0133] For example, the target terminal can compare the candidate cross-correlation value with a preset cross-correlation threshold, count the number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter in the comparison result to obtain a basic statistical number set, calculate the mean value of the basic statistical number in the basic statistical number set to obtain a basic statistical number mean value, filter out the basic statistical number with the largest value in the basic statistical number set to obtain a target statistical number, and when the target statistical number is the same as a preset number threshold, calculate the number ratio between the target statistical number and the basic statistical number mean value, and when the number ratio exceeds a preset number ratio, determine that the play state of the to-be-played voice frame signal is normal play. When the play state is normal play and the current voice frame number passes the check, the target terminal determines that the current call state with the first terminal is normal call.
[0134] 209、The target terminal sends prompt information of the current call state to the first terminal.
[0135] For example, the target terminal can generate prompt information of normal call according to the current call state, and send the prompt information to the first terminal.
[0136] 210、The first terminal prompts the current call state according to the prompt information.
[0137] For example, when the first terminal is a hardware device product, the first terminal can give a light prompt through a prompt light, or give a vibration prompt through a vibration motor, or directly display a normal call prompt information on a call page, or can also display a target call icon corresponding to the normal call in a call status icon according to the prompt information. The call icon is mainly used to prompt that the current call status of the first terminal corresponding to the speaker is a normal call. The form of the call icon can be various. Taking a small horn as an example, the call page displayed on the first terminal can be as shown in FIG. 9. Figure 4
[0138] 211. When the target terminal detects the sound signal, the target terminal collects the current sound signal, and generates the target call information according to the current sound signal.
[0139] For example, when the target terminal detects the sound signal, the target terminal collects the current generated sound signal, and performs frame processing on the sound signal to obtain a plurality of sound frame signals. Each sound frame signal is detected through a vad algorithm, and the sound signal with vad being 1 is selected from the sound frame signal as a target sound frame signal corresponding to the voice, so as to take the target sound frame signal as a voice detection result. Each target sound frame signal is numbered, so as to obtain a target voice frame identifier corresponding to each target sound frame signal, or each target sound frame signal is sorted, and the sorting result is converted into an identifier of each target sound frame signal, so as to obtain the target voice frame identifier. The sound frame signal is preprocessed, a preset voice coding protocol or coding strategy is adopted to code the preprocessed sound frame signal to obtain an initial voice frame signal, the target voice frame identifier corresponding to the basic voice frame signal is selected from the initial voice frame signal, the target voice frame identifier is added to the basic voice frame identifier, so as to obtain the target voice frame signal.
[0140] The target terminal selects the target voice frame signal with the target voice frame identifier from the target voice frame signal, and counts the number of the selected target voice frame signal, so as to obtain the target voice frame number. The target voice frame identifier and the target voice frame number are taken as target attribute information of the target voice frame signal, and the target voice frame signal and the target attribute information of the target voice frame signal are fused to obtain the target call information.
[0141] 212. The target terminal sends the target call information to the second terminal, and receives target prompt information of a call status corresponding to the target call information returned by the second terminal.
[0142] For example, the target terminal sends the target call information to the second terminal, so that the second terminal detects the call state according to the target call information, and when the second terminal detects that the call state is a normal call, the second terminal returns the target prompt information corresponding to the normal call, and prompts the call state based on the target prompt information. The way of prompting the call state can be the same as that of the first terminal, which will not be repeated here.
[0143] The first terminal and the second terminal can be the same or different.
[0144] As can be seen from the above, after receiving the current call information sent by the first terminal, the target terminal in the embodiment analyzes the voice frame signal in the current call information according to the attribute information in the current call information, obtains the to-be-played voice frame signal and the current voice frame number of the voice frame signal, then plays the to-be-played voice frame signal, collects the played to-be-played voice frame signal to obtain the recording frame signal corresponding to the to-be-played voice frame signal, then verifies the current voice frame number based on the attribute information, calculates the cross-correlation value of the recording frame signal and the to-be-played voice frame signal, and then determines the current call state with the first terminal according to the verification result of the current voice frame number and the cross-correlation value. Since this scheme can identify the collected voice signal, and then the listening party verifies the voice frame number in the whole process of instant messaging, in addition, the listening party determines whether the sound is successfully played by calculating the cross-correlation of the to-be-played voice frame signal and the recording frame signal, so that the whole process of the call can be realized. Therefore, the accuracy of call state detection can be improved.
[0145] In order to better implement the above method, the embodiment of the application also provides a call state detection device, which can be integrated in an electronic device, such as a server or a terminal device. The terminal device can include a tablet computer, a notebook computer, and / or a personal computer, etc.
[0146] For example, as shown in Figure 7 The call state device can include a receiving unit 301, an analyzing unit 302, a collecting unit 303, a verifying unit 304, and a determining unit 305, as follows:
[0147] (1) the receiving unit 301;
[0148] The receiving unit 301 is configured to receive the current call information sent by the first terminal, and the current call information includes a voice frame signal and attribute information of the voice frame signal.
[0149] For example, the receiving unit 301 can be specifically configured to establish an information transmission channel with the first terminal through a network, receive the current call information sent by the first terminal through the information transmission channel, or when the current call information is sensitive information, receive the encrypted information of the current call information sent by the first terminal, decrypt the encrypted information based on a preset encryption protocol to obtain the current call information, or when the current call information is large in memory or large in quantity, indirectly obtain the current call information from the first terminal.
[0150] (2) the parsing unit 302;
[0151] The parsing unit 302 is configured to parse the voice frame signal based on the attribute information to obtain the to-be-played voice frame signal and the current number of voice frames of the voice frame signal.
[0152] For example, the parsing unit 302 can be specifically configured to decode the voice frame signal to obtain a decoded voice frame signal, post-process the decoded voice frame signal to obtain the to-be-played voice frame signal, and count the number of voice frames of the decoded voice frame signal and the to-be-played voice frame signal based on the attribute information to obtain the current number of voice frames of the voice frame signal.
[0153] (3) the collecting unit 303;
[0154] The collecting unit 303 is configured to play the to-be-played voice frame signal and collect the played to-be-played voice frame signal to obtain a recording frame signal corresponding to the to-be-played voice frame signal.
[0155] For example, the collecting unit 303 can be specifically configured to play the to-be-played voice frame signal through a loudspeaker, record the played to-be-played voice frame signal through a microphone at the same time to obtain a recording signal, and frame the recording signal according to the window length of each frame in the to-be-played voice frame signal, so as to obtain a recording frame signal with the same window length as the to-be-played voice frame signal.
[0156] (4) the verifying unit 304;
[0157] The verifying unit 304 is configured to verify the current number of voice frames based on the attribute information, and calculate a cross-correlation value of the recording frame signal and the to-be-played voice frame signal, the cross-correlation value being used to indicate the similarity of the recording frame signal and the to-be-played voice frame signal.
[0158] For example, the checking unit 304 can be specifically configured to extract the basic voice frame number of the voice frame signal from the attribute information, compare the basic voice frame number with the current voice frame number when the basic voice frame number and the current voice frame number exceed the preset frame number threshold, determine the checking result of the current voice frame number based on the comparison result, obtain a preset time delay parameter set between the to-be-played voice frame signal and the recorded frame signal, select the to-be-played voice frame signal corresponding to each preset time delay parameter from the to-be-played voice frame signal, obtain a candidate to-be-played voice frame signal set corresponding to each recorded frame signal, and calculate the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding to the recorded frame signal.
[0159] The determining unit 305 can be specifically configured to determine the current call state of the first terminal according to the checking result of the current voice frame number and the cross-correlation value.
[0160] The determining unit 305 can be specifically configured to determine the current call state of the first terminal according to the checking result of the current voice frame number and the cross-correlation value.
[0161] For example, the determining unit 305 can be specifically configured to compare the candidate cross-correlation value with a preset cross-correlation threshold, count the number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter in the comparison result to obtain a basic statistical number set, and determine the playing state of the to-be-played voice frame signal based on the basic statistical number set. When the playing state is normal playing and the current voice frame number passes the checking, it is determined that the current call state of the first terminal is normal call.
[0162] Optionally, the call state detection apparatus can further include a sending unit 306, as shown in Figure 8 The sending unit 306 can be specifically configured as follows:
[0163] The sending unit 306 can be specifically configured to, when the sound signal is detected, collect the current sound signal to generate target call information, and send the target call information to the second terminal so that the second terminal detects the call state based on the target call information.
[0164] For example, the sending unit 306 can be specifically configured to, when the sound signal is detected, collect the current sound signal, perform frame processing on the sound signal to obtain a plurality of sound frame signals, perform encoding processing on the sound frame signals to generate target call information, and send the target call information to the second terminal so that the second terminal detects the call state based on the target call information.
[0165] In specific implementation, the above units can be implemented as independent entities, or can be combined as the same or several entities, and the specific implementation of the above units can be referred to the method embodiments above, which will not be described herein.
[0166] From the above, after the receiving unit 3012 receives the current call information sent by the first terminal, the parsing unit 302 parses the voice frame signal in the current call information according to the attribute information in the current call information, obtains the to-be-played voice frame signal and the current voice frame number of the voice frame signal, then the collecting unit 303 plays the to-be-played voice frame signal and collects the to-be-played voice frame signal to obtain the recording frame signal corresponding to the to-be-played voice frame signal, then the checking unit 304 checks the current voice frame number based on the attribute information and calculates the cross-correlation value of the recording frame signal and the to-be-played voice frame signal, and then the determining unit 305 determines the current call state with the first terminal according to the checking result of the current voice frame number and the cross-correlation value. Since the scheme can identify the collected voice signal, and then the voice frame number of the voice signal is checked by the listening party in the whole process of instant messaging, in addition, whether the sound is successfully played is determined by the listening party through the cross-correlation calculation of the to-be-played voice frame signal and the recording frame signal, so that the whole process of voice detection of the call can be realized, and therefore the accuracy of the call state detection can be improved.
[0167] The embodiment of the present application also provides an electronic device, as shown in the figure, which shows a structural schematic diagram of the electronic device related to the embodiment of the present application, in particular: Figure 9
[0168] The electronic device can include a processor 401 with one or more processing cores, a memory 402 with one or more computer readable storage media, a power supply 403, an input unit 404 and the like. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements. Among them: Figure 9
[0169] The processor 401 is the control center of the electronic device, which connects all parts of the electronic device through various interfaces and lines, executes the software programs and / or modules stored in the memory 402 and the data stored in the memory 402, and processes data to perform various functions of the electronic device, thereby overall detecting the electronic device. Optionally, the processor 401 can include one or more processing cores; preferably, the processor 401 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 401.
[0170] The memory 402 can be used to store software programs and modules, and the processor 401 executes various functions and data processing by running the software programs and modules stored in the memory 402. The memory 402 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 402 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0171] The electronic device also includes a power supply 403 for powering the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so that the power management system can be used to manage charging, discharging, and power consumption management, etc. The power supply 403 can also include one or more direct current or alternating current power sources, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and any other components.
[0172] The electronic device can also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0173] Although not shown, the electronic device can also include a display unit, etc., which will not be described here. In particular, in the present embodiment, the processor 401 in the electronic device loads the executable file corresponding to the process of one or more application programs into the memory 402 according to the following instructions, and runs the application programs stored in the memory 402 by the processor 401, thereby realizing various functions, as follows:
[0174] Receiving current call information sent by the first terminal, the current call information including voice frame signals and attribute information of the voice frame signals, parsing the voice frame signals according to the attribute information to obtain to-be-played voice frame signals and a current number of voice frames of the voice frame signals, playing the to-be-played voice frame signals, collecting the played to-be-played voice frame signals to obtain recording frame signals corresponding to the to-be-played voice frame signals, verifying the current number of voice frames based on the attribute information, and calculating a cross-correlation value of the recording frame signals and the to-be-played voice frame signals, the cross-correlation value being used to indicate a similarity degree of the recording frame signals and the to-be-played voice frame signals, and determining a current call state with the first terminal according to a verification result of the current number of voice frames and the cross-correlation value.
[0175] For example, an information transmission channel is established with the first terminal through the network, current call information sent by the first terminal is received through the information transmission channel, or when the current call information is sensitive information, encrypted information of the current call information sent by the first terminal is also received, the encrypted information is decrypted based on a preset encryption protocol to obtain the current call information, or when the current call information has a large memory or a large quantity, the current call information is indirectly obtained from the first terminal. The voice frame signal is decoded to obtain a decoded voice frame signal, the decoded voice frame signal is post-processed to obtain a to-be-played voice frame signal, the decoded voice frame signal and the to-be-played voice frame signal are counted based on the attribute information to obtain a current voice frame number of the voice frame signal. The to-be-played voice frame signal is played through the loudspeaker, and at the same time of playing the to-be-played voice frame signal, the played to-be-played voice frame signal is recorded through the microphone to obtain a recording signal, and the recording signal is framed according to the window length of each frame in the to-be-played voice frame signal, so that a recording frame signal with the same window length as the to-be-played voice frame signal is obtained. The basic voice frame number of the voice frame signal is extracted from the attribute information, when the basic voice frame number and the current voice frame number exceed a preset frame number threshold, the basic voice frame number and the current voice frame number are compared, and based on the comparison result, a verification result of the current voice frame number is determined. A set of preset time delay parameters between the to-be-played voice frame signal and the recording frame signal is obtained, the to-be-played voice frame signal corresponding to each preset time delay parameter is selected from the to-be-played voice frame signal to obtain a candidate to-be-played voice frame signal set corresponding to each recording frame signal, and a candidate cross-correlation value between each recording frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding to the recording frame signal is calculated. The candidate cross-correlation value is compared with a preset cross-correlation threshold, the number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter is counted in the comparison result to obtain a basic statistical number set, and based on the basic statistical number set, a playing state of the to-be-played voice frame signal is determined. When the playing state is normal playing and the current voice frame number passes the verification, it is determined that the current call state with the first terminal is normal call. When a sound signal is detected, a current generated sound signal is collected, the sound signal is framed to obtain a plurality of sound frame signals, the sound frame signal is encoded to generate target call information, and the target call information is sent to the second terminal so that the second terminal detects a call state based on the target call information.
[0176] The specific implementation of each operation can refer to the foregoing embodiments, which will not be repeated here.
[0177] From the above, the embodiment of the present application receives the current call information sent by the first terminal, analyzes the voice frame signal in the current call information according to the attribute information in the current call information, obtains the to-be-played voice frame signal and the current voice frame number of the voice frame signal, then plays the to-be-played voice frame signal, collects the played to-be-played voice frame signal, obtains the recording frame signal corresponding to the to-be-played voice frame signal, then checks the current voice frame number based on the attribute information, calculates the cross-correlation value of the recording frame signal and the to-be-played voice frame signal, then determines the current call state with the first terminal according to the checking result of the current voice frame number and the cross-correlation value. Since the scheme can identify the collected voice signal, and then check the voice frame number of the voice signal in the whole process of instant messaging by the listening party, and in addition, the listening party determines whether the sound is successfully played by calculating the cross-correlation of the to-be-played voice frame signal and the recording frame signal, so that the whole process of voice detection of the call can be realized, and therefore the accuracy of call state detection can be improved.
[0178] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by related hardware controlled by the instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.
[0179] To this end, the embodiment of the present application provides a computer readable storage medium, which stores a plurality of instructions capable of being loaded by a processor to execute the steps in any one of the call state detection methods provided by the embodiment of the present application. For example, the instructions can execute the following steps:
[0180] Receiving current call information sent by the first terminal, the current call information including a voice frame signal and attribute information of the voice frame signal, analyzing the voice frame signal according to the attribute information, obtaining a to-be-played voice frame signal and a current voice frame number of the voice frame signal, playing the to-be-played voice frame signal, collecting the played to-be-played voice frame signal, obtaining a recording frame signal corresponding to the to-be-played voice frame signal, checking the current voice frame number based on the attribute information, and calculating a cross-correlation value of the recording frame signal and the to-be-played voice frame signal, the cross-correlation value being used to indicate the similarity of the recording frame signal and the to-be-played voice frame signal, and determining the current call state with the first terminal according to the checking result of the current voice frame number and the cross-correlation value.
[0181] For example, an information transmission channel is established with the first terminal through the network, current call information sent by the first terminal is received through the information transmission channel, or when the current call information is sensitive information, encrypted information of the current call information sent by the first terminal is also received, the encrypted information is decrypted based on a preset encryption protocol to obtain the current call information, or when the current call information has a large memory or a large quantity, the current call information is indirectly obtained from the first terminal. The voice frame signal is decoded to obtain a decoded voice frame signal, the decoded voice frame signal is post-processed to obtain a to-be-played voice frame signal, the decoded voice frame signal and the to-be-played voice frame signal are counted based on the attribute information to obtain a current voice frame number of the voice frame signal. The to-be-played voice frame signal is played through the loudspeaker, and at the same time of playing the to-be-played voice frame signal, the played to-be-played voice frame signal is recorded through the microphone to obtain a recording signal, and the recording signal is framed according to the window length of each frame in the to-be-played voice frame signal, so that a recording frame signal with the same window length as the to-be-played voice frame signal is obtained. The basic voice frame number of the voice frame signal is extracted from the attribute information, when the basic voice frame number and the current voice frame number exceed a preset frame number threshold, the basic voice frame number and the current voice frame number are compared, and based on the comparison result, a verification result of the current voice frame number is determined. A set of preset time delay parameters between the to-be-played voice frame signal and the recording frame signal is obtained, the to-be-played voice frame signal corresponding to each preset time delay parameter is selected from the to-be-played voice frame signal to obtain a candidate to-be-played voice frame signal set corresponding to each recording frame signal, and a candidate cross-correlation value between each recording frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding to the recording frame signal is calculated. The candidate cross-correlation value is compared with a preset cross-correlation threshold, the number of candidate cross-correlation values exceeding the preset cross-correlation threshold under each preset time delay parameter is counted in the comparison result to obtain a basic statistical number set, and based on the basic statistical number set, a playing state of the to-be-played voice frame signal is determined. When the playing state is normal playing and the current voice frame number passes the verification, it is determined that the current call state with the first terminal is normal call. When a sound signal is detected, a current generated sound signal is collected, the sound signal is framed to obtain a plurality of sound frame signals, the sound frame signal is encoded to generate target call information, and the target call information is sent to the second terminal so that the second terminal detects a call state based on the target call information.
[0182] The specific implementation of each operation can refer to the foregoing embodiments, which will not be described here.
[0183] The computer readable storage medium can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0184] The steps of any of the call state detection methods provided by the embodiments of the present application can be executed due to the instructions stored in the computer readable storage medium, thus the beneficial effects of any of the call state detection methods provided by the embodiments of the present application can be achieved, which will be described in detail in the foregoing embodiments and will not be repeated here.
[0185] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in any of the various optional implementation manners of the call state detection aspect.
[0186] The above describes in detail a call state detection method, device and computer readable storage medium provided by the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, the specific implementation manners and application range will be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A call state detection method characterized by, The method comprises the following steps: receiving current call information sent by a first terminal, the current call information comprising a voice frame signal and attribute information of the voice frame signal; analyzing the voice frame signal according to the attribute information to obtain a to-be-played voice frame signal and a current voice frame number of the voice frame signal, comprising: decoding the voice frame signal to obtain a decoded voice frame signal; post-processing the decoded voice frame signal to obtain the to-be-played voice frame signal; and counting the voice frame number of the decoded voice frame signal and the to-be-played voice frame signal based on the attribute information to obtain the current voice frame number of the voice frame signal; playing the to-be-played voice frame signal and collecting the played to-be-played voice frame signal to obtain a recorded frame signal corresponding to the to-be-played voice frame signal; verifying the current voice frame number based on the attribute information and calculating a cross-correlation value of the recorded frame signal and the to-be-played voice frame signal, the cross-correlation value being used to indicate the similarity between the recorded frame signal and the to-be-played voice frame signal; determining a current call state of the first terminal according to the verification result of the current voice frame number and the cross-correlation value.
2. The call state detection method according to claim 1, wherein The verification of the current voice frame number based on the attribute information comprises: extracting a basic voice frame number of the voice frame signal from the attribute information; comparing the basic voice frame number with the current voice frame number when the basic voice frame number and the current voice frame number exceed a preset frame number threshold; determining the verification result of the current voice frame number based on the comparison result.
3. The call state detection method according to claim 2, wherein The current voice frame number comprises a received voice frame number and a played voice frame number, and the comparison of the basic voice frame number with the current voice frame number comprises: calculating a first frame number ratio by taking the received voice frame number and the basic voice frame number as a divisor and a dividend respectively; calculating a second frame number ratio by taking the received voice frame number and the played voice frame number as a divisor and a dividend respectively; The determination of the verification result of the current voice frame number based on the comparison result comprises: determining that the verification of the current voice frame number is passed when the first frame number ratio and the second frame number ratio do not exceed a preset frame number ratio threshold.
4. The talker state detection method of claim 1, wherein, The calculation of the cross-correlation value of the recorded frame signal and the to-be-played voice frame signal comprises: obtaining a preset time delay parameter set between the to-be-played voice frame signal and the recorded frame signal; selecting, from the to-be-played voice frame signal, a to-be-played voice frame signal corresponding to each preset time delay parameter in the preset time delay parameter set to obtain a candidate to-be-played voice frame signal set corresponding to each recorded frame signal; calculating a candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set; The determination of the current call state of the first terminal according to the verification result of the current voice frame number and the cross-correlation value comprises: determining the current call state of the first terminal according to the verification result of the current voice frame number and the candidate cross-correlation value.
5. The talk state detection method according to claim 4, wherein The calculation of the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set comprises: Calculate the power spectrum value of each frequency point in each recorded frame signal to obtain a first power spectrum value; Calculate the power spectrum value of each frequency point in each to-be-played voice frame signal in the candidate to-be-played voice frame signal set to obtain a second power spectrum value; Fuse the first power spectrum value and the second power spectrum value to obtain a candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set.
6. The talker state detection method of claim 5, wherein, The fusing of the first power spectrum value and the second power spectrum value to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set comprises: Calculate the mean value of the first power spectrum value to obtain a first power spectrum mean value, and calculate the mean value of the second power spectrum value to obtain a second power spectrum mean value; Calculate the difference value of the first power spectrum value and the first power spectrum mean value to obtain a first power spectrum difference value, and calculate the difference value of the second power spectrum value and the second power spectrum mean value to obtain a second power spectrum difference value; Fuse the first power spectrum difference value and the second power spectrum difference value to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the corresponding candidate to-be-played voice frame signal set.
7. The talker state detection method of claim 4, wherein, The determination of the current call state of the first terminal according to the verification result of the current voice frame number and the candidate cross-correlation value comprises: Determine the playing state of the to-be-played voice frame signal based on the candidate cross-correlation value; When the playing state is normal playing and the verification of the current voice frame number passes, determine that the current call state of the first terminal is normal call.
8. The talk state detection method according to claim 7, wherein The determination of the playing state of the to-be-played voice frame signal based on the candidate cross-correlation value comprises: Compare the candidate cross-correlation value with a preset cross-correlation threshold value; In the comparison result, count the number of candidate cross-correlation values exceeding the preset cross-correlation threshold value under each preset time delay parameter to obtain a basic statistical number set; Determine the playing state of the to-be-played voice frame signal based on the basic statistical number set.
9. The talk state detection method according to claim 8, wherein The determination of the playing state of the to-be-played voice frame signal based on the basic statistical number set comprises: Calculate the mean value of the basic statistical number in the basic statistical number set to obtain a basic statistical number mean value; Filter out the basic statistical number with the largest value in the basic statistical number set to obtain a target statistical number; When the target statistical number is the same as a preset number threshold value, calculate the number ratio value between the target statistical number and the basic statistical number mean value; When the number ratio value exceeds a preset number ratio threshold value, determine that the playing state of the to-be-played voice frame signal is normal playing.
10. The talk state detection method of claim 7, wherein, After determining that the current call state of the first terminal is normal call, further comprising: Generate a prompt information of normal call according to the current call state; Send the prompt information to the first terminal, so that the first terminal prompts the current call state according to the prompt information.
11. The talkers state detection method of claim 1, wherein, The speech frame number statistics are performed on the decoded speech frame signal and the to-be-played speech frame signal based on the attribute information, to obtain a current speech frame number of the speech frame signal, including: extracting a speech frame identifier of the speech frame signal from the attribute information; counting the speech frame identifier in the decoded speech frame signal to obtain a received speech frame number of the decoded speech frame signal; performing speech detection on the to-be-played speech frame signal to obtain a played speech frame number of the to-be-played speech frame signal; taking the received speech frame number and the played speech frame number as the current speech frame number of the speech frame signal.
12. The talk state detection method according to any one of claims 1 to 10, wherein, Further comprising: when a sound signal is detected, collecting a current generated sound signal, and performing frame processing on the sound signal to obtain a plurality of sound frame signals; performing encoding processing on the sound frame signal to generate target call information; sending the target call information to a second terminal, so that the second terminal detects a call state based on the target call information.
13. The talk state detection method of claim 12, wherein, The encoding processing on the sound frame signal to generate target call information includes: performing speech detection on the sound frame signal, and generating a target speech frame identifier based on a speech detection result; performing pre-processing on the sound frame signal, and performing speech encoding processing on the pre-processed sound frame signal to obtain a target speech frame signal; counting the target speech frame identifier in the target speech frame signal to obtain a target speech frame number, and taking the target speech frame identifier and the target speech frame number as target attribute information of the target speech frame signal; fusing the target speech frame signal and the target attribute information of the target speech frame signal to obtain target call information.
14. A call state detecting apparatus characterized by comprising: Comprise: a receiving unit, configured to receive current call information sent by a first terminal, the current call information comprising a speech frame signal and attribute information of the speech frame signal; an analysis unit, configured to analyze the speech frame signal based on the attribute information to obtain a to-be-played speech frame signal and a current speech frame number of the speech frame signal, including: decoding the speech frame signal to obtain a decoded speech frame signal; performing post-processing on the decoded speech frame signal to obtain the to-be-played speech frame signal; and performing speech frame number statistics on the decoded speech frame signal and the to-be-played speech frame signal based on the attribute information to obtain the current speech frame number of the speech frame signal; a collecting unit, configured to play the to-be-played speech frame signal, and collect a played to-be-played speech frame signal to obtain a recording frame signal corresponding to the to-be-played speech frame signal; a checking unit, configured to check the current speech frame number based on the attribute information, and calculate a cross-correlation value of the recording frame signal and the to-be-played speech frame signal, the cross-correlation value being used to indicate a similarity degree of the recording frame signal and the to-be-played speech frame signal; a determination unit, configured to determine a current call state with the first terminal according to a checking result of the current speech frame number and the cross-correlation value.
15. The call state detection apparatus according to claim 14, characterized in that, The check unit is specifically configured to extract a basic voice frame number of the voice frame signal from the attribute information; and compare the basic voice frame number with a current voice frame number when the basic voice frame number and the current voice frame number exceed a preset frame number threshold. The check result of the current voice frame number is determined based on a comparison result.
16. The apparatus according to claim 15, wherein The current voice frame number includes a received voice frame number and a played voice frame number. The check unit is specifically configured to calculate a ratio of the received voice frame number to the basic voice frame number to obtain a first frame number ratio. A ratio of the received voice frame number to the played voice frame number is calculated to obtain a second frame number ratio. The check result of the current voice frame number is determined based on the comparison result, including determining that the current voice frame number passes the check when the first frame number ratio and the second frame number ratio do not exceed a preset frame number ratio threshold.
17. The call state detection apparatus of claim 14, wherein The check unit is specifically configured to obtain a preset time delay parameter set between the to-be-played voice frame signal and the recorded frame signal; and filter out, from the to-be-played voice frame signal, a to-be-played voice frame signal corresponding to each preset time delay parameter in the preset time delay parameter set to obtain a candidate to-be-played voice frame signal set corresponding to each recorded frame signal. A candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding thereto is calculated. The current call state of the first terminal is determined according to the check result of the current voice frame number and the cross-correlation value.
18. The call state detection apparatus of claim 17, wherein The check unit is specifically configured to calculate a power spectrum value of a frequency point in each recorded frame signal to obtain a first power spectrum value; and calculate a power spectrum value of a frequency point in each to-be-played voice frame signal in the candidate to-be-played voice frame signal set to obtain a second power spectrum value. The first power spectrum value and the second power spectrum value are fused to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding thereto.
19. The call state detection apparatus of claim 18, wherein The check unit is specifically configured to calculate a mean value of the first power spectrum value to obtain a first power spectrum mean value, and calculate a mean value of the second power spectrum value to obtain a second power spectrum mean value; calculate a difference value between the first power spectrum value and the first power spectrum mean value to obtain a first power spectrum difference value, and calculate a difference value between the second power spectrum value and the second power spectrum mean value to obtain a second power spectrum difference value. The first power spectrum difference value and the second power spectrum difference value are fused to obtain the candidate cross-correlation value between each recorded frame signal and each to-be-played voice frame signal in the candidate to-be-played voice frame signal set corresponding thereto.
20. The call state detection apparatus of claim 17, wherein The determining unit is specifically configured to determine a playing state of the to-be-played voice frame signal based on the candidate cross-correlation values; and when the playing state is normal playing and the current voice frame number check passes, determine that a current call state with the first terminal is normal call. 21.The call state detection apparatus of claim 20, characterized in that, The determining unit is specifically configured to compare the candidate cross-correlation values with a preset cross-correlation threshold; count, in a comparison result, a number of candidate cross-correlation values that exceed the preset cross-correlation threshold under each preset time delay parameter to obtain a basic statistical number set; and determine the playing state of the to-be-played voice frame signal based on the basic statistical number set. 22.The call state detection apparatus of claim 21, characterized in that, The determining unit is specifically configured to calculate a mean value of the basic statistical numbers in the basic statistical number set to obtain a basic statistical number mean value; select, from the basic statistical number set, a basic statistical number with the largest value to obtain a target statistical number; when the target statistical number is the same as a preset number threshold, calculate a number ratio value between the target statistical number and the basic statistical number mean value; and when the number ratio value exceeds a preset number ratio threshold, determine that the playing state of the to-be-played voice frame signal is normal playing. 23.The call state detection apparatus of claim 20, characterized in that, The determining unit is specifically configured to generate a normal call prompt information according to the current call state; and send the prompt information to the first terminal so that the first terminal prompts the current call state according to the prompt information. 24.The call state detection apparatus of claim 14, characterized in that, The analyzing unit is specifically configured to extract a voice frame identifier of the voice frame signal from the attribute information; count the voice frame identifier in the decoded voice frame signal to obtain a received voice frame number of the decoded voice frame signal; perform voice detection on the to-be-played voice frame signal to obtain a played voice frame number of the to-be-played voice frame signal; and take the received voice frame number and the played voice frame number as a current voice frame number of the voice frame signal. The call state detection apparatus further comprises a sending unit.
25. The talk state detection apparatus according to any one of claims 14 to 23, wherein The sending unit is specifically configured to, when a sound signal is detected, collect a current generated sound signal, perform frame processing on the sound signal to obtain a plurality of sound frame signals, perform encoding processing on the sound frame signals to generate target call information, and send the target call information to a second terminal so that the second terminal detects a call state based on the target call information. 26.The call state detection apparatus of claim 25, characterized in that, The sending unit is specifically configured to perform voice detection on the sound frame signal, generate a target voice frame identifier based on a voice detection result, perform pre-processing on the sound frame signal, perform voice coding processing on the sound frame signal after pre-processing to obtain a target voice frame signal, count the target voice frame identifier in the target voice frame signal to obtain a target voice frame number, and take the target voice frame identifier and the target voice frame number as target attribute information of the target voice frame signal; and fuse the target voice frame signal and the target attribute information of the target voice frame signal to obtain target call information.
27. An electronic device, comprising: The computer program / instruction is executed by the processor to implement the steps in the call state detection method of any one of claims 1 to 13.
28. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the steps in the call state detection method of any one of claims 1 to 13.
29. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of instructions adapted to be loaded by the processor to execute the steps in the call state detection method of any one of claims 1 to 13.
Citation Information
Patent Citations
Method and device for detecting voice service noise
CN102695206A
Voice signal processing method and device, and storage medium
CN111768800A