Risk voice call identification method and device

By converting the speech signal into a quantum state and using the self-attention matrix and long-term short-term memory network to process it, the recognition accuracy and efficiency of traditional speech recognition in complex environments is solved, and the rapid and accurate recognition of risk voice calls is achieved.

CN120354096APending Publication Date: 2025-07-22CHINA TELECOM CORP LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510429665.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional voice recognition technology has insufficient noise resistance in complex call environments, making it difficult to capture non-verbal information in risky voice, and has low processing efficiency, resulting in reduced recognition accuracy and insufficient real-time performance.

Method used

By converting time domain signals into quantum state representations, speech features are extracted using self-attention matrix and long and short-term memory networks, combined with the parallel processing capabilities of quantum computing, and dynamically adjust the weights to achieve efficient recognition of risk speech.

Benefits of technology

It improves the speed and efficiency of voice recognition, and can accurately identify risky voice calls in a noisy environment, meet real-time needs and avoid recognition lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354096A_ABST
    Figure CN120354096A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for identifying a risk voice call. The method comprises the following steps: acquiring a time domain signal of a voice call, and determining a first quantum state representation corresponding to the time domain signal; target features are extracted from the first quantum state representation, a self-attention matrix is constructed according to the target features, and the target features at least comprise a frequency spectrum feature, an energy feature and a fundamental frequency feature; performing weighting processing on the target features by using the self-attention matrix to obtain target quantum features; analyzing the target quantum features by using a long short-term memory network to obtain intermediate feature representation; and performing voice classification according to the intermediate feature representation to obtain a classification result, the classification result being used for reflecting whether the voice call is a risk voice call. The technical problem that a traditional voice recognition algorithm is difficult to efficiently and accurately recognize a risk voice call is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology. Specifically, it relates to a method and device for identifying risk voice calls. Background Art

[0002] In a communication environment, speech recognition technology is widely used in the process of voice calls to monitor voice calls. However, existing speech recognition technologies have obvious limitations and cannot effectively meet the identification requirements of risk voice calls. First, the anti-noise performance is insufficient. The real call environment is complex and background noise occurs frequently. The recognition accuracy of traditional recognition systems drops significantly under noise interference, especially for those subtle voice changes in risk calls. Second, the feature extraction is rough. Traditional algorithms that rely on basic parameters such as spectrum and energy are difficult to capture the rich non-verbal information contained in risk voices, such as emotional fluctuations and psychological states. The lack of this information reduces the comprehensiveness and accuracy of recognition. Finally, the processing efficiency is low. With the rapid increase in data volume, the computational burden of traditional technologies increases, and the real-time feedback ability is limited. As a result, in the risk voice monitoring scenario, the recognition process lags behind, and the best intervention time may be missed.

[0003] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a method and device for identifying risk voice calls to at least solve the technical problem that traditional speech recognition algorithms are difficult to efficiently and accurately identify risk voice calls.

[0005] According to one aspect of the embodiments of this application, a method for identifying risk voice calls is provided, including: obtaining a time-domain signal of a voice call and determining a first quantum state representation corresponding to the time-domain signal; extracting target features from the first quantum state representation and constructing a self-attention matrix based on the target features, where the target features at least include: spectral features, energy features, and fundamental frequency features; using the self-attention matrix to perform weighted processing on the target features to obtain target quantum features; using a long short-term memory network to analyze the target quantum features to obtain an intermediate feature representation; and performing speech classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risk voice call.

[0006] Optionally, obtaining a time-domain signal of a voice call and determining a first quantum state representation corresponding to the time-domain signal includes: obtaining a time-domain signal of a voice call and performing preprocessing on the time-domain signal to obtain a target time-domain signal, where the preprocessing includes at least one of the following: noise reduction processing, normalization processing; using Fourier transform to convert the target time-domain signal into a target frequency-domain signal; and using quantum Fourier transform to convert the target frequency-domain signal into a first quantum state representation.

[0007] Optionally, extracting target features from the first quantum state representation includes: performing quantum measurement on the first quantum state representation to obtain spectral features for reflecting the energy distribution state of the voice call in the frequency domain; calculating the first quantum state representation based on a preset energy operator to obtain a second quantum state representation, performing quantum measurement on the second quantum state representation to obtain energy features for reflecting the energy distribution state of the voice call in the time domain; analyzing the first quantum state representation using a quantum phase estimation algorithm to obtain a third quantum state representation, performing quantum measurement on the third quantum state representation to obtain fundamental frequency features for reflecting the fundamental frequency change state of the voice call.

[0008] Optionally, constructing a self-attention matrix based on the target features includes: determining the product of the spectral features and a preset first weight matrix as the query vector, and determining the product of the energy features and a preset second weight matrix as the key vector; constructing a self-attention matrix based on the query vector and the key vector, where the calculation formula for each element in the matrix is as follows:

[0009]

[0010] In the formula, Q i is the i-th feature in the query vector Q, K j and K k are the j-th feature and the k-th feature in the key vector K respectively, d is the feature dimension of the vector to be weighted, and A ij is the attention weight of the j-th feature in the vector to be weighted to the i-th feature.

[0011] Optionally, performing weighted processing on the target features using the self-attention matrix to obtain target quantum features includes: determining the product of the fundamental frequency features and a preset third weight matrix as the value vector; performing weighted summation on the features in the value vector using the self-attention matrix to obtain target quantum features.

[0012] Optionally, performing voice classification based on the intermediate feature representation to obtain a classification result includes: respectively determining a speech rate index, an intonation index, and a tone index based on the intermediate feature representation; performing weighted summation on the speech rate index, the intonation index, and the tone index using preset weight coefficients to obtain a comprehensive score of the voice call; in the case where the comprehensive score is greater than a preset score threshold, determining that the voice call is a risky voice call; in the case where the comprehensive score is not greater than the preset score threshold, determining that the voice call is a normal voice call.

[0013] Optionally, respectively determining a speech rate index, an intonation index, and a tone index based on the intermediate feature representation includes: respectively calculating the speech rate index, the intonation index, and the tone index using the following formulas:

[0014]

[0015] S tone = max(H) - min(H)

[0016]

[0017] wherein, S speed is the speech rate index, S tone is the intonation index, S stability is the tone index, H is the intermediate feature representation, n is the length of H, H i and H i+1 are the i-th element and the (i + 1)-th element in H respectively, max(H) and min(H) are the maximum element and the minimum element in H respectively, represents the average value of all elements in H.

[0018] Optionally, perform speech classification based on the intermediate feature representation to obtain a classification result, including: performing normalization processing on the intermediate feature representation, and using the quantum Fourier transform to convert the normalized intermediate feature representation into a fourth quantum state representation; using a quantum algorithm to classify the fourth quantum state representation to obtain a classification result, wherein the quantum algorithm includes at least one of the following: quantum support vector machine, quantum neural network, quantum nearest neighbor algorithm, quantum Bayesian classifier, Grover search algorithm.

[0019] Optionally, the quantum algorithm is the Grover search algorithm. Using the quantum algorithm to classify the fourth quantum state representation to obtain a classification result, including: initializing the fourth quantum state representation into a superposition state; using a preset marking function to perform marking processing on each quantum state in the fourth quantum state representation, wherein the marking function is used to analyze the input quantum state from the dimensions of speech rate, intonation and tone to determine whether the input quantum state corresponds to a risky voice call. If so, mark the input quantum state, otherwise do not mark the input quantum state. The marking includes: quantum phase flip; performing diffusion processing on the marked quantum state space, wherein the diffusion processing is used to enhance the probability amplitude of the marked quantum state and suppress the probability amplitude of the unmarked quantum state; repeatedly performing the marking processing and the diffusion processing for a preset number of times, and performing quantum measurement on the obtained quantum state representation, and determining the classification result according to the measurement result.

[0020] Optionally, after obtaining the classification result, the above method further includes: when the classification result indicates that the voice call is a risky voice call, performing a preset risk control action, wherein the risk control action includes at least one of the following: hanging up the voice call, sending a risk call prompt message to the user.

[0021] According to another aspect of the embodiments of the present application, there is also provided an identification device for risk voice calls, including: an acquisition module, configured to acquire the time-domain signal of a voice call and determine the first quantum state representation corresponding to the time-domain signal;

[0022] an extraction module, configured to extract target features from the first quantum state representation and construct a self-attention matrix based on the target features, where the target features at least include: spectral features, energy features, and fundamental frequency features; a weighting module, configured to perform a weighting process on the target features by using the self-attention matrix to obtain target quantum features; an analysis module, configured to analyze the target quantum features by using a long short-term memory network to obtain an intermediate feature representation; a classification module, configured to perform voice classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risk voice call.

[0023] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes: a computer program, where when the computer program is executed by a processor, it implements the above-mentioned identification method for risk voice calls.

[0024] According to another aspect of the embodiments of the present application, there is also provided an electronic device, which includes: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned identification method for risk voice calls through the computer program.

[0025] In the embodiments of the present application, by converting the time-domain signal into a quantum state representation, extracting features from the quantum state, then introducing a self-attention matrix to perform a weighting process on the target feature matrix to obtain target quantum features, and analyzing the target quantum features by a long short-term memory network, the self-attention matrix can better extract and process voice features, and the self-attention mechanism dynamically adjusts the weights, which can highlight and identify voices with risk features from the voice features. Throughout the process, the introduced quantum computing has powerful parallel processing capabilities, which can significantly improve the speed and efficiency of voice recognition. Especially when dealing with large-scale data sets, it can achieve fast response and meet real-time requirements, thereby solving the technical problem that traditional voice recognition algorithms are difficult to efficiently and accurately identify risk voice calls. Description of the Drawings

[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0027] Figure 1 is a schematic flowchart of an alternative identification method for risk voice calls according to the embodiments of the present application;

[0028] Figure 2 It is a schematic structural diagram of an optional risk voice call recognition device according to an embodiment of the present application;

[0029] Figure 3 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] To better understand the embodiments of the present application, some nouns or terms that appear in the description process of the embodiments of the present application are first translated and explained as follows:

[0033] Self-attention mechanism: It is an algorithmic concept used to capture the internal correlations in sequence data, especially in the fields of natural language processing and speech recognition. It allows the model to refer to other positions in the entire sequence when processing data at a certain position, rather than relying solely on fixed neighboring information before and after. The self-attention mechanism calculates the similarity between parts of the sequence and dynamically assigns attention weights, enabling the model to focus on the most informative parts of the input data. In telecommunication information analysis, the self-attention mechanism can help algorithms more accurately identify key feature changes, such as the intonation of speech or the sentiment of text, thereby improving the overall performance of the model. Specifically, the self-attention mechanism is achieved through three main steps: calculating query, key, and value vectors, calculating attention weights, and weighted summing the value vectors according to the weights, finally outputting a vector that aggregates global information and can better represent the important content in the sequence.

[0034] Long Short-Term Memory network (LSTM): It is a special type of recurrent neural network designed to address the problems of vanishing gradients and exploding gradients in processing long sequence data. It can effectively learn and remember long-term dependencies through a unique gating mechanism and maintain good performance even when the sequence length is very large. In the LSTM network, the information flow passes through three gates: the input gate, the forget gate, and the output gate, which control the storage, update, and retrieval of information. The input gate determines which new information will be stored in the cell state, the forget gate determines which information will be retained or discarded, and the output gate controls which information will be passed to the next layer of the network.

[0035] Grover search algorithm: A quantum algorithm specifically designed for efficient search in an unsorted database. It utilizes the principle of quantum interference and can find the target item in far less time than a classical computer requires. In traditional search, finding an item in an unsorted database may take linear time, i.e., O(N), while the Grover algorithm can reduce the search time to O(√N), greatly improving the search efficiency. The core of the Grover algorithm lies in creating a superposition of quantum states, where the target item has a relatively large probability amplitude compared to other items. Through a series of carefully designed quantum operations, including marking the target item and diffusion operations, the algorithm can gradually enhance the probability amplitude of the target item until it becomes almost the only possible result in the measurement. This method is particularly effective in dealing with search problems of large amounts of data, such as finding call samples that meet specific characteristics in a vast number of call records.

[0036] Fourier Transform: A mathematical tool widely used in signal processing and data analysis. It can transform a signal from the time domain or spatial domain to the frequency domain, revealing the frequency composition of the signal. Through the Fourier Transform, a signal varying with time can be decomposed into the superposition of its various frequency components, which provides a new perspective for analyzing and understanding the signal.

[0037] In the embodiments of the present application, the information collected is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions. Necessary confidentiality measures are taken, which do not violate public order and good customs, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0038] Embodiment 1

[0039] According to the embodiments of the present application, a method for identifying risky voice calls is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0040] Figure 1 is a schematic flowchart of a method for identifying risky voice calls provided according to the embodiments of the present application, as Figure 1 shown, the method includes the following steps:

[0041] Step S102, obtain the time-domain signal of the voice call and determine the first quantum state representation corresponding to the time-domain signal;

[0042] Step S104, extract target features from the first quantum state representation and construct a self-attention matrix based on the target features, where the target features at least include: spectral features, energy features, fundamental frequency features;

[0043] Step S106, perform weighted processing on the target features using the self-attention matrix to obtain target quantum features;

[0044] Step S108, analyze the target quantum features using a long short-term memory network to obtain an intermediate feature representation;

[0045] Step S110, perform voice classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risky voice call.

[0046] The following describes each step of the method for identifying risky voice calls in combination with the specific implementation process.

[0047] As an alternative implementation, after obtaining the time-domain signal in a voice call, the first quantum state representation corresponding to the time-domain signal can be determined in the following manner: Obtain the time-domain signal of the voice call, and perform preprocessing on the time-domain signal to obtain a target time-domain signal. Use Fourier transform to convert the target time-domain signal into a target frequency-domain signal; use quantum Fourier transform to convert the target frequency-domain signal into the first quantum state representation.

[0048] Suppose the time-domain signal of the collected original voice call is x(t) = [x1, x2, …, x n , where n is the number of sampling points, which refers to the number of times of sampling a continuous analog signal within a certain period. The original voice signal can be an electrical signal obtained from a microphone or other audio acquisition devices, and then the time-domain signal of the original voice call is processed through preprocessing operations.

[0049] Among them, the preprocessing operations include but are not limited to one of the following: noise reduction processing, normalization processing. The noise in noise reduction processing can be achieved through a filter. By noise reduction processing, the background noise in the recording can be removed or reduced, and the clarity of the signal can be improved, making the subsequent feature extraction more accurate. And normalization representation can adjust the amplitude of the signal to [-1, 1], which helps to eliminate the amplitude differences between different input signals and ensures that the recognition algorithm will not be biased due to different signal intensities.

[0050] Since in the frequency-domain signal, features related to the speed, intonation, tone, etc. in a risky voice call will be more obvious, it is necessary to convert the preprocessed signal x'(t) into a frequency-domain signal X(f). The Fourier transform is used to convert the time-domain signal into a frequency-domain signal for better analysis of the frequency components of the signal. X(f) reflects the energy distribution of the processed signal x'(t) at different frequencies f, and can be specifically expressed by the following formula:

[0051]

[0052] In order to provide an efficient basis for subsequent feature extraction and processing, the frequency-domain signal can be efficiently converted into a quantum state representation by using the parallel processing ability of quantum computing through quantum Fourier transform (QFT). The frequency-domain signal X(f) is converted into the first quantum state representation, and the first quantum state representation is denoted as: QFT(X(f)).

[0053] As an alternative implementation, extracting target features from the quantum state representation can be achieved in the following ways: performing a quantum measurement on the first quantum state representation to obtain a spectral feature that reflects the energy distribution state of the voice call in the frequency domain; calculating the first quantum state representation based on a preset energy operator to obtain a second quantum state representation, performing a quantum measurement on the second quantum state representation to obtain an energy feature that reflects the energy distribution state of the voice call in the time domain; using a quantum phase estimation algorithm to analyze the first quantum state representation to obtain a third quantum state representation, performing a quantum measurement on the third quantum state representation to obtain a fundamental frequency feature that reflects the fundamental frequency change state of the voice call.

[0054] Specifically, before performing a quantum measurement, a suitable measurement basis needs to be selected first. The measurement basis usually consists of a set of orthogonal quantum states, and these quantum states correspond to the spectral features of the voice signal. Selecting the correct measurement basis is the key to extracting accurate spectral features. Performing a quantum measurement on the prepared frequency-domain quantum state, the quantum measurement is a non-deterministic process that collapses the quantum state and gives a measurement result with a certain probability. The choice of the measurement basis determines the information that can be extracted from the quantum state, and the measurement result provides the energy distribution information of the signal at each frequency. Specifically, the probability of measuring a certain specific quantum state is actually the square of the amplitude of the projection of this quantum state on the original frequency-domain quantum state. This means that through measurement, we can directly obtain the features of the signal in the frequency domain, that is, the spectral features, which reflect the energy distribution of different frequency components of the signal.

[0055] The energy operator can be regarded as a quantization of the signal intensity, and it can act on the quantum state representation, emphasizing the parts with higher energy in the signal. In the quantum representation of the voice signal, this means that the energy operator will act on those quantum state components corresponding to strong vibrations of the sound, which is usually related to the parts of the voice with large pronunciation force or high volume. Operating on the first quantum state QFT(X(f)) with the energy operator to generate a new quantum state representation, the newly generated quantum state representation is marked as the second quantum state representation. This process is equivalent to weighting the signal intensity in the quantum domain, making the energy features more prominent. Finally, performing a measurement on the second quantum state representation to extract the energy feature. The measurement here is essentially observing the energy distribution at each time point in the signal, and the measurement result will reveal the energy intensity of the signal at different time periods, and this information can be used to construct the distribution of the energy feature.

[0056] During the calculation of the quantum phase estimation algorithm, two quantum registers need to be prepared. The first register is used to store the quantum state to be analyzed, that is, the input state; the second register acts as a counter, and the final change of its quantum state will reflect the result of the phase estimation. The core step of the quantum phase estimation algorithm is to perform the controlled-U operation multiple times, where U is a matrix unit related to the target quantum state, and the controlled-U operation depends on the state of the second register to determine whether to apply U to the first register. This operation is repeated a specific number of times, usually corresponding to the number of binary digits of the required phase accuracy. After the controlled-U operation is completed, the inverse quantum Fourier transform is performed on the second quantum register again. This step converts the state of the second register into a probability distribution containing phase information, so that the phase value can be extracted by measurement. The second quantum register is measured. The measurement result will be a number, which corresponds to an approximation of the estimated phase value in binary representation. The accuracy of the phase estimation algorithm depends on the size of the second register, and increasing the number of its qubits can improve the estimation accuracy.

[0057] In the above manner, the target feature that fuses the spectral feature F, the energy feature E, and the fundamental frequency feature P is extracted.

[0058] As an alternative implementation, constructing a self-attention mechanism based on the target feature to construct a self-attention matrix can be achieved in the following way:

[0059] Step S1, determine that the product of the spectral feature and a preset first weight matrix is the query vector, and determine that the product of the energy feature and a preset second weight matrix is the key vector.

[0060] In the above process, the spectral feature F can reflect the distribution of the signal at different frequencies, which helps to capture the frequency characteristics of the signal. The query vector can be expressed as Q = W Q F; the energy feature E can reflect the energy magnitude of the signal in the time domain, which helps to capture the energy distribution of the signal. The key vector can be expressed as K = W K E, where W Q and W K represent the first weight matrix and the second weight matrix respectively.

[0061] Step S2, construct a self-attention matrix based on the query vector and the key vector, which can be specifically expressed by the following formula:

[0062]

[0063] where Q i is the i-th feature in the query vector Q, K j and K k are the j-th feature and the k-th feature in the key vector K respectively, d is the feature dimension of the vector to be weighted, and Aij is the attention weight of the j-th feature in the vector to be weighted with respect to the i-th feature.

[0064] Considering that risk voice features generally involve the following points: 1) Speech rate, for example, the main caller usually speaks quickly to increase the sense of urgency; 2) Intonation, for example, the main caller may use exaggerated intonation to attract the victim's attention; 3) Tone, for example, the main caller may use a high-pressure or threatening tone to force the victim to take action. Therefore, it is possible to consider introducing the self-attention mechanism to dynamically adjust the weights of different features, so that the system can pay more attention to the key features in the risk voice.

[0065] As an optional implementation manner, after obtaining the self-attention matrix and the target feature, the target quantum feature can be obtained by weighting the target feature with the self-attention matrix. Among them, the fundamental frequency feature P can reflect the basic frequency of the signal and is helpful for capturing the pitch change of the speech. Therefore, the features in the value vector related to the fundamental frequency feature can be used as the target feature. The target quantum feature can be obtained in the following way: First, the product of the fundamental frequency feature and a preset third weight matrix is used as the value vector, and the self-attention matrix is used to perform weighted summation on the features in the value vector to obtain the target quantum feature.

[0066] Specifically, the value vector can be expressed as V = W V P, where W V is the preset third weight matrix, and the target quantum feature obtained by performing weighted summation on the features in the value vector can be expressed as: V ′ = softmax(A) · V.

[0067] In the above process, the query vector and the key vector are used to calculate the self-attention matrix A, while the value vector V is used for the final weighted summation. Since the fundamental frequency feature P is very crucial in identifying the pitch change of the speech, and the pitch change of the speech is an important feature of the risk voice, therefore, the product of the fundamental frequency feature P and the preset third weight matrix is introduced as the value vector. Although the spectral feature F and the energy feature E also participate in the calculation of the self-attention matrix, the final weighted summation only uses the fundamental frequency feature P because the fundamental frequency feature is more crucial in identifying the risk voice.

[0068] As an optional implementation manner, after obtaining the target quantum feature, the extraction of the intermediate feature representation can be achieved through a long short-term memory network. The specific implementation process is as follows:

[0069] Suppose a deep neural network model such as a long short-term memory network can be constructed to process the quantum feature vector and perform pattern recognition. Specifically, the weighted quantum feature vector V ′Input into the long short-term memory network, after multiple layers of processing, the intermediate feature representation H is obtained, that is, H = LSTM(V ′ ), which contains higher-level feature information: 1) Context information of the time series: The long short-term memory network can capture the context information of the speech signal in time, which helps to understand the continuity and changes of the speech; 2) Long-term dependence relationship: The long short-term memory network can capture the long-term dependence relationship in the speech signal through the gating mechanism, which helps to identify complex patterns in the speech; 3) Frequency and energy information: The intermediate feature representation H contains high-level information extracted from spectral features, energy features, and fundamental frequency features. These feature information helps the subsequent classification tasks.

[0070] After determining the intermediate feature representation, the speech can be further classified according to the intermediate feature representation.

[0071] As an optional implementation manner, the speech can be classified according to the intermediate feature representation in the following way to obtain the classification result: Determine the speech rate index, intonation index, and tone index respectively according to the intermediate feature representation; Use the preset weight coefficients to perform weighted summation on the speech rate index, intonation index, and tone index to obtain the comprehensive score of the voice call; When the comprehensive score is greater than the preset score threshold, determine that the voice call is a risk voice call; When the comprehensive score is not greater than the preset score threshold, determine that the voice call is a normal voice call.

[0072] Optionally, the speech rate index, intonation index, and tone index can be respectively represented by the following formulas:

[0073]

[0074] S tone = max(H) - min(H)

[0075]

[0076] Among them, S speed is the speech rate index, which reflects the change rate of the feature values; S tone is the intonation index, which reflects the fluctuation range of the feature values; S stability is the tone index, which reflects the stability of the feature values; H is the intermediate feature representation, n is the length of H, H i and H i+1 are respectively the i-th element and the (i + 1)-th element in H, max(H) and min(H) are respectively the maximum element and the minimum element in H, represents the average value of all elements in H.

[0077] When calculating S speed , |H i+1 - Hi | represents the difference between two adjacent eigenvalue in the intermediate feature representation H. The larger this difference is, the more drastic the change in the eigenvalue is, which may reflect a relatively fast speech rate. In addition, is the sum of the absolute differences between adjacent elements of the intermediate feature representation H, is to normalize the sum of the absolute differences so that feature vectors of different lengths are comparable when being compared; when calculating S tone max(H)-min(H) represents the fluctuation range of the eigenvalues in the intermediate feature representation H. The larger this range is, the more drastic the change in the eigenvalue is, which may reflect the degree of intonation exaggeration; when calculating S stability when, represents the sum of the absolute differences between each eigenvalue in the intermediate feature H and its average value. The larger it is, the more unstable the change in the eigenvalue is, which may reflect a high-pressure tone.

[0078] The speech rate index S speed , intonation index S tone and tone index S stability are calculated respectively. After that, weighted summation is performed on each index to obtain the final comprehensive score. The comprehensive score can be expressed by the following formula:

[0079] Score=W speed ·S speed +W tone ·S tone +W stability ·S stability

[0080] where, W speed , W tone and W stability represent the weights of the speech rate index, intonation index and tone index respectively, which can be adjusted according to actual needs and are not specifically limited here.

[0081] After obtaining the comprehensive score, the comprehensive score can be classified according to a preset threshold. When the comprehensive score is greater than the preset threshold, it is classified as a risky voice call. When the comprehensive score is not greater than the preset threshold, it is classified as a normal voice call. In actual output, a risky voice call can be output as 1 and a normal voice call can be output as 0.

[0082] In addition to the above method of processing the speech rate index, intonation index, and tone index through weighting to obtain a comprehensive score to determine whether it is a risky voice call, voice classification can also be performed based on intermediate features in the following way: normalize the intermediate feature representation, and use the quantum Fourier transform to convert the normalized intermediate feature representation into a fourth quantum state representation; use a quantum algorithm to classify the fourth quantum state representation to obtain a classification result.

[0083] Among them, normalizing the intermediate feature representation can ensure that the values of its respective elements are within the range of [0, 1], providing a basis for subsequent classification processing. After obtaining the normalized intermediate feature representation, it can be mapped to a quantum state representation and labeled as the fourth quantum state representation. Specifically, the normalized feature representation can be mapped to a quantum state representation through the quantum Fourier transform or other quantum encoding methods. After obtaining the quantum state representation, a quantum algorithm can be used to classify the quantum state representation to obtain the final classification result.

[0084] Optionally, the quantum algorithm includes but is not limited to: quantum support vector machine, quantum neural network, quantum nearest neighbor algorithm, quantum Bayesian classifier, Grover search algorithm.

[0085] Among them, when using a quantum algorithm such as the Grover search algorithm to classify the quantum state, the optimal solution can be found in a relatively short time, thus accelerating the classification speed.

[0086] Optionally, when using the Grover search algorithm to classify the fourth quantum state representation, it can be done in the following way: initialize the fourth quantum state representation as a superposition state; use a preset marking function to perform marking processing on each quantum state in the fourth quantum state representation, where the marking function is used to analyze the input quantum state from the dimensions of speech rate, intonation, and tone to determine whether the input quantum state corresponds to a risky voice call. If so, mark the input quantum state; if not, do not mark the input quantum state. The marking includes: quantum phase flip; perform diffusion processing on the quantum state space after the marking processing is completed, where the diffusion processing is used to enhance the probability amplitude of the marked quantum state and suppress the probability amplitude of the unmarked quantum state; repeatedly execute the marking processing and diffusion processing a preset number of times, and perform quantum measurement on the obtained quantum state representation, and determine the classification result based on the measurement result.

[0087] Among them, when the marking function analyzes the input quantum state from dimensions such as speech rate, intonation, and tone, the following rules can be referred to. The characteristics of a normal voice call are: generally moderate speech rate, without obvious haste or excessive speed; relatively stable intonation, without frequent pitch changes; natural tone, without elements of high pressure or threat. While the characteristics of a risky voice call are: fast speaking speed to increase a sense of urgency; exaggerated intonation to attract the attention of the called end user; high-pressure or threatening tone to force the called user to take action.

[0088] In the above process, the principle of action of the marking function is that when a quantum state representation (the fourth quantum state representation) matches the characteristics of a risky call, the function will perform an operation on this quantum state and mark it as a "target" or "non-target" quantum state. The marking operation usually uses quantum phase flipping, that is, changing the phase of the marked quantum state, which can be achieved in quantum computing by applying a conditional quantum gate (such as the Z gate), and this gate only affects the phase of the quantum state when specific conditions are met.

[0089] The purpose of quantum phase flipping is to enhance the probability amplitude of the target quantum state and suppress the probability amplitude of the non-target quantum state through quantum interference in subsequent diffusion processing. This is usually achieved through the Grover diffusion operation, which can be regarded as a reflection process in the quantum state space. It uses the average amplitude as a mirror and adjusts the probability amplitudes of all quantum states accordingly, so that the probability amplitude of the quantum state that matches the conditions of the marking function increases, and the probability amplitude of the quantum state that does not match decreases. The diffusion processing enables the quantum search algorithm to converge to the target quantum state faster, thereby improving the search efficiency.

[0090] Through a sufficient number of cycles, the frame rate amplitude of the target quantum state can be significantly enhanced until it becomes the most likely result in quantum state measurement. In each cycle, the marking function first identifies and marks the target quantum state, and then the diffusion processing adjusts the probability amplitudes of all quantum states. This process is repeated continuously, gradually making the target quantum state prominent.

[0091] Perform quantum measurement on the quantum state representation after marking and diffusion processing to determine the classification result. Quantum measurement is the process of reading the information of the quantum system, and the result is random, but the quantum state with a larger probability amplitude is more likely to be measured. If the measured quantum state representation is marked by the marking function, it is determined as a risky voice call; if it is not marked, it is determined as a normal voice call. In this way, by combining the results of multiple measurements, the accuracy and reliability of classification can be improved. A risky voice call can be output as 1, and a normal voice call can be output as 0.

[0092] Optionally, when the classification result indicates that the voice call is a risky voice call, a preset risk control action can be executed to prompt the called user of the terminal to pay attention to prevention, thereby avoiding losses to the called user. Among them, the risk control actions include, but are not limited to, one of the following: hanging up the voice call, sending a risk call prompt message to the user.

[0093] In the above steps, first, extracting key features from the quantum state as target features can achieve accurate feature extraction. Second, dynamically weighting the target features through the self-attention matrix can effectively highlight the voice features of risky voices. Combining the self-attention mechanism and quantum features enables the recognition algorithm to exhibit better anti-noise performance in a noisy environment. After that, a long short-term memory network is used to analyze the target quantum features to obtain an intermediate feature representation. Finally, techniques such as weighted scoring or quantum measurement are used to analyze the intermediate feature representation to achieve the classification of voice calls. Throughout the process, the parallel processing ability of quantum computing is utilized, which can significantly improve the speed and efficiency of voice recognition. Especially when dealing with large-scale data sets, it can achieve fast response and meet real-time requirements. Thus, the technical problem that traditional voice recognition algorithms are difficult to efficiently and accurately identify risky voice calls is solved.

[0094] Embodiment 2

[0095] According to an embodiment of the present application, there is also provided an apparatus for recognizing a risky voice call for implementing the method for recognizing a risky voice call in Embodiment 1, as Figure 2 shown. The apparatus for recognizing a risky voice call at least includes: an acquisition module 21, an extraction module 22, a weighting module 23, an analysis module 24, and a classification module 25, where:

[0096] The acquisition module 21 is configured to acquire the time-domain signal of the voice call and determine the first quantum state representation corresponding to the time-domain signal;

[0097] The extraction module 22 is configured to extract target features from the first quantum state representation and construct a self-attention matrix based on the target features. Among them, the target features at least include: spectral features, energy features, and fundamental frequency features;

[0098] The weighting module 23 is configured to weight the target features by using the self-attention matrix to obtain target quantum features;

[0099] The analysis module 24 is configured to analyze the target quantum features by using a long short-term memory network to obtain an intermediate feature representation;

[0100] The classification module 25 is configured to perform voice classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risky voice call.

[0101] The functions of each module of the risk voice call recognition device will be described below in combination with a specific implementation process.

[0102] As an alternative implementation, when the acquisition module acquires the time-domain signal of a voice call and determines the first quantum state representation corresponding to the time-domain signal, it can be implemented in the following manner: acquire the time-domain signal of the voice call, and perform preprocessing on the time-domain signal to obtain a target time-domain signal, where the preprocessing includes at least one of the following: noise reduction processing, normalization processing; use Fourier transform to convert the target time-domain signal into a target frequency-domain signal; use quantum Fourier transform to convert the target frequency-domain signal into a first quantum state representation.

[0103] After obtaining the first quantum state representation, the extraction module can extract target features in the following manner: perform quantum measurement on the first quantum state representation to obtain a spectral feature for reflecting the energy distribution state of the voice call in the frequency domain; calculate the first quantum state representation based on a preset energy operator to obtain a second quantum state representation, perform quantum measurement on the second quantum state representation to obtain an energy feature for reflecting the energy distribution state of the voice call in the time domain; use the quantum phase estimation algorithm to analyze the first quantum state representation to obtain a third quantum state representation, and perform quantum measurement on the third quantum state representation to obtain a fundamental frequency feature for reflecting the fundamental frequency change state of the voice call.

[0104] As an alternative implementation, after the extraction module extracts the target features, it can construct a self-attention matrix in the following manner: determine the product of the spectral feature and a preset first weight matrix as the query vector, and determine the product of the energy feature and a preset second weight matrix as the key vector;

[0105] Construct a self-attention matrix based on the query vector and the key vector, where the calculation formula for each element in the matrix is as follows:

[0106]

[0107] where, Q i is the i-th feature in the query vector Q, K j and K k are the j-th feature and the k-th feature in the key vector K respectively, d is the feature dimension of the vector to be weighted, and A ij is the attention weight of the j-th feature in the vector to be weighted to the i-th feature.

[0108] As an alternative implementation, the weighting module can obtain the target quantum feature by weighting the target feature through the self-attention matrix in the following way: determining the product of the fundamental frequency feature and a preset third weight matrix as a value vector; using the self-attention matrix to perform weighted summation on the features in the value vector to obtain the target quantum feature.

[0109] The analysis module uses a long short-term memory network to analyze the target quantum feature to obtain an intermediate feature representation. Further, the classification module can classify the voice call based on the intermediate feature representation in the following way: respectively determining a speech rate index, an intonation index, and a tone index based on the intermediate feature representation; using preset weight coefficients to perform weighted summation on the speech rate index, the intonation index, and the tone index to obtain a comprehensive score for the voice call; in the case where the comprehensive score is greater than a preset score threshold, determining the voice call as a risky voice call; in the case where the comprehensive score is not greater than the preset score threshold, determining the voice call as a normal voice call.

[0110] Among them, the speech rate index, the intonation index, and the tone index can be respectively represented by the following formulas:

[0111]

[0112] S tone = max(H) - min(H)

[0113]

[0114] Among them, S speed is the speech rate index, which reflects the change rate of the feature values; S tone is the intonation index, which reflects the fluctuation range of the feature values; S stability is the tone index, which reflects the stability of the feature values. H is the intermediate feature representation, n is the length of H, H i and H i+1 are respectively the i-th element and the (i + 1)-th element in H, max(H) and min(H) are respectively the maximum element and the minimum element in H, represents the average value of all elements in H.

[0115] In addition to the above method of processing the speech rate index, the intonation index, and the tone index through weighting to obtain a comprehensive score to determine whether it is a risky voice call, the embodiments of the present application also provide another alternative implementation to classify the monitored voice call: performing normalization processing on the intermediate feature representation, and using the quantum Fourier transform to convert the normalized intermediate feature representation into a fourth quantum state representation; using a quantum algorithm to classify the fourth quantum state representation to obtain a classification result.

[0116] Among them, the quantum algorithms include but are not limited to one of the following: quantum support vector machine, quantum neural network, quantum nearest neighbor algorithm, quantum Bayesian classifier, Grover search algorithm.

[0117] Optionally, when using a quantum algorithm such as the Grover search algorithm to classify quantum states, it can be specifically implemented in the following way: initialize the representation of the fourth quantum state as a superposition state; use a preset marking function to mark each quantum state in the representation of the fourth quantum state, where the marking function is used to analyze the input quantum state from the dimensions of speech rate, intonation, and tone, determine whether the input quantum state corresponds to a risk voice call, if so, mark the input quantum state, if not, do not mark the input quantum state, and the marking includes: quantum phase flip; perform a diffusion process on the quantum state space after the marking process is completed, where the diffusion process is used to enhance the probability amplitude of the marked quantum state and suppress the probability amplitude of the unmarked quantum state; repeatedly execute the marking process and the diffusion process a preset number of times, and perform a quantum measurement on the obtained quantum state representation, and determine the classification result based on the measurement result.

[0118] As an alternative implementation, the above-mentioned risk voice call recognition device may further include an execution module. When the classification result of the classification module indicates that the voice call is a risk voice call, the execution module may perform a preset risk control action, where the risk control action includes but is not limited to one of the following: hanging up the voice call, sending a risk call prompt message to the user.

[0119] It should be noted that each module in the risk voice call recognition device in the embodiments of the present application corresponds one by one to each implementation step of the risk voice call recognition method in Embodiment 1. Since detailed descriptions have been made in Embodiment 1, some details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0120] Embodiment 3

[0121] According to the embodiments of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the risk voice call recognition method in Embodiment 1.

[0122] According to the embodiments of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. The device where the non-volatile storage medium is located executes the risk voice call recognition method in Embodiment 1 by running the computer program.

[0123] According to the embodiments of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, it executes the risk voice call recognition method in Embodiment 1.

[0124] According to an embodiment of the present application, an electronic device is further provided, which includes: a memory and a processor. Among them, a computer program is stored in the memory, and the processor is configured to execute the risk voice call recognition method in Embodiment 1 through the computer program.

[0125] Specifically, when the computer program runs, it executes the following steps: obtaining the time-domain signal of the voice call and determining the first quantum state representation corresponding to the time-domain signal; extracting target features from the first quantum state representation and constructing a self-attention matrix based on the target features, where the target features at least include: spectral features, energy features, and fundamental frequency features; using the self-attention matrix to perform weighted processing on the target features to obtain target quantum features; using a long short-term memory network to analyze the target quantum features to obtain an intermediate feature representation; and performing voice classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risk voice call.

[0126] As an alternative implementation manner, the above-mentioned electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 3 The hardware structure block diagram of an electronic device for implementing the risk voice call recognition method is shown. As Figure 3 shown, the electronic device 30 may include one or more (shown as 302a, 302b,..., 302n in the figure) processors 302 (the processor 302 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 30 may further include more or fewer components than Figure 3 shown, or have a different configuration from Figure 3 shown.

[0127] It should be noted that the one or more processors 302 and / or other data processing circuits described above can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other components in the electronic device 30. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0128] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the risk voice call recognition method in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the vulnerability detection method of the above application program. The memory 304 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 304 can further include a memory remotely disposed relative to the processor 302, and these remote memories can be connected to the electronic device 30 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0129] The transmission device 306 is used to receive or send data via a network. Specific examples of the above network can include the wireless network provided by the communication provider of the electronic device 30. In one instance, the transmission device 306 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 306 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0130] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the electronic device 30.

[0131] The above embodiment numbers are only for description and do not represent the advantages or disadvantages of the embodiments.

[0132] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0134] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0137] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for identifying risky voice calls, characterized in that, including: Obtain the time-domain signal of a voice call and determine the first quantum state representation corresponding to the time-domain signal; Extract target features from the first quantum state representation and construct a self-attention matrix based on the target features, where at least the following are included in the target features: spectral features, energy features, and fundamental frequency features; Use the self-attention matrix to perform weighted processing on the target features to obtain target quantum features; Use a long short-term memory network to analyze the target quantum features to obtain an intermediate feature representation; Perform voice classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risky voice call.

2. The method according to claim 1, wherein Obtain the time-domain signal of a voice call and determine the first quantum state representation corresponding to the time-domain signal, including: Obtain the time-domain signal of a voice call and perform preprocessing on the time-domain signal to obtain a target time-domain signal, where the preprocessing includes at least one of the following: noise reduction processing, normalization processing; Use Fourier transform to convert the target time-domain signal into a target frequency-domain signal; Use quantum Fourier transform to convert the target frequency-domain signal into the first quantum state representation.

3. The method according to claim 1, wherein Extract target features from the first quantum state representation, including: Perform quantum measurement on the first quantum state representation to obtain spectral features for reflecting the energy distribution state of the voice call in the frequency domain; Perform calculation on the first quantum state representation based on a preset energy operator to obtain a second quantum state representation, perform quantum measurement on the second quantum state representation to obtain energy features for reflecting the energy distribution state of the voice call in the time domain; Use the quantum phase estimation algorithm to analyze the first quantum state representation to obtain a third quantum state representation, perform quantum measurement on the third quantum state representation to obtain fundamental frequency features for reflecting the fundamental frequency change state of the voice call.

4. The method according to claim 1, wherein Construct a self-attention matrix based on the target features, including: Determine the product of the spectral features and a preset first weight matrix as a query vector, and determine the product of the energy features and a preset second weight matrix as a key vector; Construct a self-attention matrix based on the query vector and the key vector, where the calculation formula for each element in the matrix is as follows: where Q i is the i-th feature in the query vector Q, K j and K k are the j-th feature and the k-th feature in the key vector K respectively, d is the feature dimension of the vector to be weighted, and A ij is the attention weight of the j-th feature in the vector to be weighted to the i-th feature.

5. The method according to claim 4, wherein Use the self-attention matrix to perform weighted processing on the target features to obtain target quantum features, including: Determine the product of the fundamental frequency features and a preset third weight matrix as a value vector; Use the self-attention matrix to perform weighted summation on the features in the value vector to obtain the target quantum features.

6. The method according to claim 1, characterized in that, Perform voice classification based on the intermediate feature representation to obtain a classification result, including: Determine a speech rate index, an intonation index, and a tone index respectively based on the intermediate feature representation; Use preset weight coefficients to perform weighted summation on the speech rate index, the intonation index, and the tone index to obtain a comprehensive score of the voice call; In the case where the comprehensive score is greater than a preset score threshold, determine that the voice call is a risky voice call; In the case where the comprehensive score is not greater than the preset score threshold, determine that the voice call is a normal voice call.

7. The method according to claim 6, characterized in that, Determine the speech rate index, intonation index, and tone index respectively according to the intermediate feature representation, including: Calculate the speech rate index, the intonation index, and the tone index respectively using the following formulas: S tone = max(H) - min(H) Wherein, S speed is the speech rate index, S tone is the intonation index, S stability is the tone index, H is the intermediate feature representation, n is the length of H, H i and H i+1 are the i-th element and the (i + 1)-th element in H respectively, max(H) and min(H) are the maximum element and the minimum element in H respectively, represents the average value of all elements in H.

8. The method according to claim 1, characterized in that Perform speech classification based on the intermediate feature representation to obtain a classification result, including: Perform normalization processing on the intermediate feature representation, and use the quantum Fourier transform to convert the normalized intermediate feature representation into a fourth quantum state representation; Use a quantum algorithm to classify the fourth quantum state representation to obtain the classification result, where the quantum algorithm includes at least one of the following: quantum support vector machine, quantum neural network, quantum nearest neighbor algorithm, quantum Bayesian classifier, Grover search algorithm.

9. The method according to claim 8, characterized in that, The quantum algorithm is the Grover search algorithm. Using the quantum algorithm to classify the fourth quantum state representation to obtain the classification result, including: Initialize the fourth quantum state representation as a superposition state; Use a preset marking function to perform marking processing on each quantum state in the fourth quantum state representation, where the marking function is used to analyze the input quantum state from the dimensions of speech rate, intonation, and tone, determine whether the input quantum state corresponds to a risky voice call, and if so, mark the input quantum state, otherwise do not mark the input quantum state. The marking includes: quantum phase flip; Perform diffusion processing on the quantum state space after the marking processing is completed, where the diffusion processing is used to enhance the probability amplitude of the marked quantum state and suppress the probability amplitude of the unmarked quantum state; Loop through the marking processing and diffusion processing a preset number of times, and perform quantum measurement on the obtained quantum state representation, and determine the classification result based on the measurement result.

10. The method according to claim 1, characterized in that, After obtaining the classification result, the method further includes: In the case where the classification result indicates that the voice call is a risky voice call, perform a preset risk control action, where the risk control action includes at least one of the following: hanging up the voice call, sending a risk call prompt message to the user.

11. An identification device for risk voice calls, characterized in that, Including: An acquisition module for acquiring the time-domain signal of the voice call and determining the first quantum state representation corresponding to the time-domain signal; An extraction module for extracting target features from the first quantum state representation and constructing a self-attention matrix based on the target features, where the target features at least include: spectral features, energy features, fundamental frequency features; A weighting module for weighting the target features using the self-attention matrix to obtain target quantum features; An analysis module for analyzing the target quantum features using a long short-term memory network to obtain an intermediate feature representation; A classification module for performing speech classification based on the intermediate feature representation to obtain a classification result, where the classification result is used to reflect whether the voice call is a risky voice call.

12. A computer program product, characterized in that, Including: A computer program, where the computer program, when executed by a processor, implements the method for identifying a risky voice call according to any one of claims 1 to 10.

13. An electronic device, characterized in that, Including: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the method for identifying a risky voice call according to any one of claims 1 to 10 through the computer program.

Citation Information

Cited By

  • Method, device and equipment for repairing kernel vulnerability of operating system and medium

    CN120781347A

  • Speech emotion similarity evaluation method and device, storage medium, product and equipment

    CN120853547A

  • Voice emotion similarity evaluation method and device, storage medium, product and equipment

    CN120853547B