Data processing method and device, electronic equipment and storage medium
By extracting frequency and time domain features from speech signals and combining them with anomaly recognition models and level determination, the problem of inaccurate anomaly detection in traditional audio violation detection methods is solved, achieving efficient and accurate speech signal processing.
Patent Information
- Application Number
- CN202511031711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional audio violation detection methods rely on speech recognition and natural language processing technologies, which cannot fully utilize audio information, resulting in inaccurate anomaly detection and a poor user experience.
By extracting frequency and time domain features from multiple frames of speech data, using a pre-set anomaly recognition model for anomaly recognition, and combining anomaly level determination and handling measures, a comprehensive analysis and processing of speech signals can be achieved.
It improves the accuracy and efficiency of anomaly identification, avoids over-processing or under-processing, ensures the timeliness and relevance of voice signal processing, and enhances the user experience.
Smart Images

Figure CN120913592A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a data processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the rapid development of financial technology, the customer service field is no longer limited to traditional counter services, but has expanded to telephone banking, online customer service, mobile applications and other channels. In particular, in this important link of telephone banking, customers can solve various user problems by calling hotlines and enjoy convenient service experience. However, there are audio violations in voice calls, so it is necessary to detect abnormal voice information in the voice call process and handle abnormal voice calls in a timely manner to improve user experience.
[0003] The traditional audio violation detection method mainly relies on speech recognition technology and natural language processing technology. The audio first needs to go through a complex preprocessing step, and then it is converted into a text format for further analysis and processing. Although this method performs well in some specific scenarios, its limitations cannot be ignored. It cannot fully utilize the voice information in the audio, and there is a problem of inaccurate abnormal detection, resulting in poor user experience. SUMMARY
[0004] The present application provides a data processing method, device, electronic equipment and storage medium to solve the problem of inaccurate voice information abnormal detection.
[0005] According to an aspect of the present application, a data processing method is provided, comprising:
[0006] determining the corresponding multi-frame voice data of the original voice signal to be processed, performing feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features;
[0007] performing abnormal recognition processing based on the frequency domain voice features and the time domain voice features through a preset abnormal recognition model to obtain an abnormal recognition result, wherein the abnormal recognition result includes at least one classification label and abnormal probability data corresponding to the classification label;
[0008] determining an abnormal level based on the abnormal recognition result to obtain an abnormal level determination result, determining an abnormal handling measure based on the abnormal level determination result, and processing the original voice signal through the abnormal handling measure.
[0009] Optionally, the multi-frame voice data is subjected to feature extraction processing to obtain frequency domain voice features and time domain voice features, including: performing feature extraction on the multi-frame voice data by a preset frequency domain feature extraction algorithm to obtain frequency domain voice features, wherein the frequency domain voice features include mel-frequency spectrum features; performing feature extraction on the multi-frame voice data by a preset time domain feature extraction algorithm to obtain multiple time domain features, and performing fusion processing on the multiple time domain features to obtain time domain voice features, wherein the time domain features include short-time energy features, zero-crossing rate features, amplitude envelope features, and instantaneous frequency derivative features.
[0010] Optionally, the preset abnormality recognition model includes a frequency domain feature processing module, a time domain feature processing module, a tone of voice feature extraction module, and a classification module; the abnormality recognition processing is performed on the frequency domain voice features and the time domain voice features by the preset abnormality recognition model to obtain an abnormality recognition result, including: performing feature processing on the frequency domain voice features by the frequency domain feature processing module to obtain a frequency domain feature processing result; performing feature processing on the time domain voice features by the time domain feature processing module to obtain a time domain feature processing result; splicing the frequency domain feature processing result and the time domain feature processing result to obtain feature splicing data, and performing feature extraction on the feature splicing data by the tone of voice feature extraction module to obtain a tone of voice feature; performing classification processing on the tone of voice feature by the classification module to obtain a classification result, and determining the classification result as the abnormality recognition result.
[0011] Optionally, the frequency domain feature processing module is constructed based on a convolutional neural network; the time domain feature processing module is constructed based on a unidirectional long short-term memory network; the tone of voice feature extraction module is constructed based on a bidirectional long short-term memory network; and the classification module is constructed based on a full connection layer and a Softmax layer.
[0012] Optionally, the abnormality level determination is performed based on the abnormality recognition result to obtain an abnormality level determination result, including: obtaining a mapping relationship between preset classification labels and level thresholds, matching at least one classification label in the abnormality recognition result with the mapping relationship to obtain a level threshold corresponding to each classification label; comparing the abnormality probability data corresponding to each classification label with the corresponding level threshold respectively to obtain a comparison result of each classification label; calling a preset abnormality level determination rule, matching the comparison result of each classification label with the preset abnormality level determination rule, and determining the abnormality level determination result according to a matching result.
[0013] Optionally, the abnormality disposal measure is determined based on the abnormality level determination result, and the original voice signal is processed through the abnormality disposal measure, including: generating corresponding abnormality prompt information according to the abnormality prompt information generation template based on the abnormality level determination result; and / or, obtaining corresponding disposal measures from the preset disposal measure set based on the abnormality level determination result; determining the abnormality disposal measure based on the abnormality prompt information and / or the disposal measure, and processing the original voice signal through the abnormality disposal measure.
[0014] Optionally, the multiple frames of voice data corresponding to the original voice signal to be processed are determined, including: obtaining the original voice signal in the authorized voice call process, and determining the original voice signal as the original voice signal to be processed; and performing a preprocessing operation on the original voice signal to obtain the multiple frames of voice data, wherein the preprocessing operation at least includes noise reduction processing and frame windowing processing.
[0015] According to another aspect of the present application, a data processing apparatus is provided, including:
[0016] The voice feature determination module is configured to determine the multiple frames of voice data corresponding to the original voice signal to be processed, and perform feature extraction processing on the multiple frames of voice data to obtain the frequency domain voice feature and the time domain voice feature.
[0017] The abnormality recognition result determination module is configured to perform abnormality recognition processing on the frequency domain voice feature and the time domain voice feature through a preset abnormality recognition model to obtain an abnormality recognition result, wherein the abnormality recognition result includes at least one classification label and abnormality probability data corresponding to the classification label.
[0018] The abnormality level determination and abnormality processing module is configured to perform abnormality level determination based on the abnormality recognition result to obtain an abnormality level determination result, determine an abnormality disposal measure based on the abnormality level determination result, and process the original voice signal through the abnormality disposal measure.
[0019] According to another aspect of the present application, an electronic device is provided, including:
[0020] at least one processor; and
[0021] a memory in communication with the at least one processor; wherein
[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data processing method of any one of the embodiments of the present application.
[0023] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for causing a processor to implement the data processing method of any of the embodiments of the present application when executed.
[0024] The technical scheme of the embodiment of the present application determines the multi-frame voice data corresponding to the original voice signal to be processed, performs feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features, performs double extraction of the time domain voice features and the frequency domain voice features on the original voice information, retains the instantaneous change information of the signal, and captures the essential characteristics in the frequency domain, thereby comprehensively depicting the multi-dimensional features of the voice signal, providing richer and more identifiable input basis for subsequent abnormality recognition and other tasks; the abnormality recognition model is used to perform abnormality recognition processing based on the frequency domain voice features and the time domain voice features, and an abnormality recognition result is obtained, wherein the abnormality recognition result includes at least one classification label and abnormality probability data corresponding to the classification label, the fusion of the frequency domain and time domain features enables the model to more comprehensively understand the voice signal characteristics, improves the accuracy of abnormality recognition, and can also recognize non-verbal violation sounds, and outputs the classification label and the corresponding abnormality probability data, which can not only determine the abnormality type, but also reflect the credibility of the recognition result through the probability value, thereby providing accurate and flexible basis for subsequent abnormality level determination and disposition measure formulation; the abnormality level determination result is obtained based on the abnormality level determination, and the abnormality disposition measure is determined based on the abnormality level determination result, and the original voice signal is processed through the abnormality disposition measure, the level determination is performed based on the quantized probability data and the clear abnormality type, the result is more objective and accurate, and the mechanism of matching the disposition measure according to the level enables closed-loop management from recognition to processing, which not only guarantees the timeliness and pertinence of processing, but also avoids excessive processing or insufficient processing, thereby effectively improving the efficiency and quality of voice signal processing. The present scheme can comprehensively capture abnormal information of the voice signal in frequency distribution and time change by synchronously extracting frequency domain and time domain voice features, avoid missed detection caused by single-dimensional features, and greatly improve the accuracy of abnormality recognition; the abnormality recognition result includes the classification label and the corresponding abnormality probability, thereby providing a quantitative basis for subsequent level determination, making the level division more accurate; the targeted disposition measure is matched based on the abnormality level, different strategies can be adopted according to the severity of the abnormality, the problem can be effectively solved, and damage to normal voice signals caused by excessive processing can be avoided, thereby realizing efficient and refined processing; the entire process forms a closed loop from feature extraction to disposition execution, and the comprehensiveness of recognition and the adaptability of processing are taken into account, which is suitable for various scenarios such as voice communication and voice recognition, and can significantly improve the quality and reliability of voice signals.
[0025] It is to be understood that the details set forth in the description contained herein do not limit the scope of the application. Other embodiments of the application will be readily apparent to those skilled in the art from the description herein. With reference to the drawings, embodiments of the application are herein described. BRIEF DESCRIPTION OF DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0027] Figure 1 is a flow chart of a data processing method provided by an embodiment of the present application;
[0028] Figure 2 is a flow chart of a data processing method provided by an embodiment of the present application;
[0029] Figure 3 is a flow chart of an abnormality recognition model processing suitable for an embodiment of the present application;
[0030] Figure 4 is a structural schematic diagram of a data processing device provided by an embodiment of the present application;
[0031] Figure 5 is a structural schematic diagram of an electronic device for implementing the data processing method of an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the technical personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of protection of the present application.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] Embodiment one
[0035] Figure 1 is a flowchart of a data processing method provided by the first embodiment of the present application. The present embodiment can be applicable to the case of detecting abnormality of voice information. The method can be executed by a data processing device, which can be realized in the form of hardware and / or software, and can be configured in a server, a computer or other electronic equipment. As shown in the figure, the method comprises: Figure 1
[0036] S110, determine the multi-frame voice data corresponding to the original voice signal to be processed, and perform feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features.
[0037] The original voice signal can be understood as an electrical signal obtained by converting continuous sound waves in the audio call process of a user, which is raw audio data without processing. It should be noted that the original voice signal to be processed in the present embodiment is collected under the authorization of the user. The multi-frame voice data can be understood as a plurality of voice segments obtained by processing continuous original voice signals according to a preset frame processing method. The preset frame processing method can be to divide a fixed time window (including overlapping parts) into a plurality of short-time segments, each frame representing the characteristics of the signal in a local time, which retains the time sequence correlation and facilitates local analysis. The time domain voice feature can be understood as a parameter reflecting the time dimension characteristics extracted from each frame of data, which can directly describe the amplitude change of the signal with time. The time domain voice feature includes but is not limited to short-time energy feature, zero-crossing rate feature, amplitude envelope feature and instantaneous frequency derivative feature. The frequency domain voice feature can be understood as a feature extracted after converting the time domain signal to the frequency domain by Fourier transform algorithm, including but not limited to Mel frequency cepstrum coefficient (MFCC, simulating the perceptual characteristics of human ear to frequency), which is used to describe the frequency distribution and energy distribution of the signal in different frequency bands.
[0038] Specifically, the original speech signal to be processed is first divided into continuous and overlapping multi-frame speech data according to a fixed time window to retain the time sequence correlation of the signal. For example, the original speech signal to be processed is divided into multi-frame speech data according to a preset time window length (such as 20-30 milliseconds) and an overlap rate (such as about 50%). Then, a time-domain speech feature extraction algorithm and a frequency-domain speech feature extraction algorithm are called to extract features from each frame of data in the multi-frame data to obtain frequency-domain speech features and time-domain speech features corresponding to the multi-frame speech data. For example, the time-domain speech features directly reflect the amplitude change and rhythm characteristics of the speech signal in the time dimension by calculating the short-time energy, zero-crossing rate, amplitude envelope feature, and instantaneous frequency derivative feature of each frame. The corresponding frequency spectrum is obtained by Fourier transform on the multi-frame speech data, which is further converted into a Mel frequency cepstral coefficient (MFCC) to characterize the frequency distribution characteristics and reflect the frequency distribution and energy distribution of the speech signal in different frequency bands.
[0039] In this embodiment, the continuous speech signal is converted into analyzable short-time segments through frame division processing, which not only adapts to the non-stationary characteristics of the speech signal, but also retains the inter-frame correlation through overlapping. In combination with the dual extraction of time-domain and frequency-domain features, the instantaneous change information of the signal is retained, and the essential characteristics of the frequency domain are captured, thereby comprehensively characterizing the multi-dimensional features of the speech signal, providing more abundant and more recognizable input basis for subsequent anomaly recognition tasks, improving the completeness and effectiveness of feature representation, providing more abundant and robust feature input for subsequent anomaly recognition tasks, and effectively improving the accuracy and reliability of subsequent processing.
[0040] Optionally, the multi-frame speech data corresponding to the original speech signal to be processed is determined by: obtaining the original speech signal in an authorized voice call process, determining the original speech signal as the original speech signal to be processed; and performing a preprocessing operation on the original speech signal to obtain the multi-frame speech data, wherein the preprocessing operation at least includes noise reduction processing and frame windowing processing.
[0041] Specifically, the original speech signal can be automatically triggered to be obtained from an authorized voice call scene and explicitly determined as the object to be processed under the condition of obtaining user authorization. Then, the original speech signal is subjected to preprocessing operation, i.e., noise reduction processing to weaken environmental noise, interference signals and other irrelevant components, and frame windowing processing, i.e., dividing the noise-reduced signal into multi-frame short-time segments according to a preset time interval window (including overlapping part), and adding a Hanning window or a Hamming window to each frame of data to reduce spectral leakage, to finally obtain the multi-frame speech data.
[0042] In some specific embodiments, noise estimation is first performed based on the original voice signal 0.5 seconds before the silence section, for suppressing the stationary noise, and the voice data X clean The determination formula of (f) is as follows:
[0043] |X clean (f)|=max(|X noisy (f)|-α·|N(f)|,β·|X noisy (f)|);
[0044] Wherein, X noisy (f) represents the original voice data before noise reduction, N(f) is the noise power spectrum estimated based on the audio 0.5 seconds before the silence, in this embodiment, it is considered that the 0.5 seconds before the silence is the stationary noise, such as air conditioner sound, background hum, etc., wherein, α, β are experimental optimization parameters, for example, α = 1.2, β = 0.1. Subsequently, the voice data after noise reduction can be further suppressed by Wiener filtering to suppress the transient noise, thereby completing the noise reduction processing. Through the above noise reduction processing, the background noise (such as keyboard clicking, environmental noise) commonly seen in voice communication scenarios can be well suppressed. After the noise reduction processing is completed, the voice data after noise reduction can be processed by frame windowing processing, the frame length can be set to 25 ms, the frame shift is 10 ms, and the window function is Hamming window. Windowing is to multiply each frame signal by a window function after framing, to reduce the boundary effect caused by framing, so that the voice signal of each frame is more smoothly transitioned to zero at the boundary, thereby avoiding problems such as spectral leakage when performing spectral analysis and the like.
[0045] In this embodiment, the voice signal is acquired in the authorized voice communication scenario, which can guarantee the legitimacy and pertinence of the data, and the noise reduction processing can significantly improve the signal-to-noise ratio, providing purer basic data for subsequent analysis; the frame windowing processing retains the local timing characteristics of the signal through short-time segmentation, and the application of the window function further optimizes the spectral analysis accuracy, so that the multi-frame data is suitable for local feature extraction and maintains the overall timing correlation, laying a high-quality foundation for subsequent processing links.
[0046] Optionally, the multi-frame voice data is subjected to feature extraction processing to obtain frequency domain voice features and time domain voice features, including: performing feature extraction on the multi-frame voice data by a pre-set frequency domain feature extraction algorithm to obtain frequency domain voice features, wherein the frequency domain voice features include mel spectrum features; performing feature extraction on the multi-frame voice data by a pre-set time domain feature extraction algorithm to obtain a plurality of time domain features, and performing fusion processing on the plurality of time domain features to obtain time domain voice features, wherein the time domain features include short-time energy features, zero-crossing rate features, amplitude envelope features and instantaneous frequency derivative features.
[0047] Specifically, a preset frequency domain feature extraction algorithm is applied to the multi-frame voice data, specifically, a fast Fourier transform (FFT) is used to convert a time domain signal into a frequency domain representation, and then a mel filter bank is applied to perform nonlinear compression on the frequency spectrum to extract mel spectrum features that can simulate human ear perception characteristics, thereby forming frequency domain voice features. Meanwhile, a preset time domain feature extraction algorithm is used to calculate the short-time energy feature, zero-crossing rate feature, amplitude envelope feature, and instantaneous frequency derivative feature of each frame, respectively, and then the multiple time domain features are processed by weighting fusion or feature concatenation, etc., to generate comprehensive time domain voice features. For example, in the time domain voice feature extraction process, the short-time energy feature, zero-crossing rate feature, amplitude envelope feature, and instantaneous frequency derivative feature of each frame of voice data are calculated to form a four-dimensional feature vector:
[0048] Short-time energy (STE) feature extraction: the sum of squares of each frame of voice data is calculated to reflect the volume change, and the calculation formula is as follows, where N is the frame length, which refers to the time length of each segment when the continuous audio signal is divided into short segments (frames), and can also correspond to the number of sampling points (number of sampling points = sampling rate × frame length time, for example, when the sampling rate is 16 kHz, 25 ms frame length corresponds to 400 sampling points):
[0049]
[0050] where x n is the voice data of the nth sampling point.
[0051] Zero-crossing rate (ZCR) feature extraction: the number of times each frame of voice data passes through zero, representing the high-frequency component, and the calculation formula is as follows, where sgn(x) is the sign function, which is 1 when x>0, -1 when x<0, and 0 when x=0:
[0052]
[0053] Amplitude envelope (AE) feature extraction: the mean of the absolute value of each frame of voice data, describing the dynamic range:
[0054]
[0055] Instantaneous frequency derivative (IFD) voice data: the difference of instantaneous frequency is calculated by Hilbert transform to capture fast-changing events, and the calculation process is as follows:
[0056] z n = x n +j·Hilbert(x n );
[0057]
[0058] For frequency domain feature extraction, the frequency domain speech features are obtained by the following way: first, the time domain signal is converted into frequency domain signal by fast Fourier transform, to obtain the following complex spectrum, wherein, N FFT is the number of sampling points used when performing fast Fourier transform (FFT) on each frame of time domain signal, N FFT ≥ frame length, if the frame length is insufficient, it needs to be zero-padded to the FFT point number (for example, frame length 400 points, FFT point number 512, then 112 zeros are padded), the larger the FFT point number, the higher the frequency domain resolution.
[0059]
[0060] The power spectrum is obtained by taking the modulus and squaring:
[0061] P[k]=|X[k]| 2 ;
[0062] The mel filter bank is set to 64 triangular filters, and the mel scale formula is as follows,
[0063]
[0064] The mel frequency points are distributed, the filter center frequencies are evenly distributed on the mel filter, covering 0 to the Nyquist frequency Each filter covers the area between two mel points in the frequency domain, and for the mth filter, the response function is as follows:
[0065]
[0066] The mel energy vector of each frame is calculated by applying the mel filter bank:
[0067]
[0068] Finally, the dynamic range is compressed to enhance the discrimination of low energy components, and a 64×T mel spectrum is obtained, T is the time step, which refers to the total number of frames obtained after the entire audio signal is framed, that is, the sequence length in the time dimension. The size of T is determined by the total duration of the audio and the frame shift (for example, 1 second of audio, frame shift 10 ms, T≈100), wherein, the correction parameter ε=0.000001 can be set to prevent taking the logarithm of zero, and the mel spectrum features are finally calculated by the following formula:
[0069] Mel[m,t]=log(E[m]+ε)。
[0070] In the embodiment, the mel-frequency spectrum feature in the frequency domain not only retains the frequency distribution characteristics of the speech, but also improves the feature representation ability for semantic information through mel scale compression; the fusion of the time domain multi-features comprehensively describes the dynamic characteristics of the speech from different dimensions, and enhances the complementarity of the features; the cooperative extraction of the two types of features provides more comprehensive and more identifiable feature representation for subsequent speech recognition, abnormality detection and other tasks, and effectively improves the understanding ability and processing precision of the model for the speech signal.
[0071] In S120, an abnormality recognition model is used to perform abnormality recognition processing based on the frequency domain speech feature and the time domain speech feature, to obtain an abnormality recognition result. The abnormality recognition result includes at least one classification label and abnormality probability data corresponding to the classification label.
[0072] The abnormality recognition result can be understood as a judgment output of whether there is an abnormality in the speech signal and the type of the abnormality. The abnormality recognition result includes at least one classification label and abnormality probability data corresponding to the classification label. The classification label explicitly indicates the specific type of the abnormality, including but not limited to noise surge, abnormal tone, sensitive information, threatening information, and non-verbal violation behavior. The abnormality probability data quantifies the credibility of the classification label. For example, if a label corresponds to an abnormality probability of 80%, the label can be determined as the type of the abnormality. The combination of the classification label and the abnormality probability data forms a complete description of the abnormal state of the speech signal. The classification label is a symbol or a text identifier used to identify the type of the abnormality. The classification label is a specific classification of the abnormality attribute of the speech signal based on the input feature of the abnormality recognition model, and can intuitively distinguish different types of abnormalities, thereby providing a clear direction for subsequent targeted processing.
[0073] Specifically, the extracted frequency domain speech feature and the time domain speech feature are spliced or weighted and fused to form a comprehensive feature vector containing multi-dimensional information. The vector is input into a pre-set abnormality recognition model. The model performs nonlinear transformation and classification decision on the comprehensive feature vector feature through the trained parameters, and outputs classification labels corresponding to different types of abnormalities. At the same time, the probability value corresponding to each classification label is calculated based on the output layer activation function of the model, as abnormality probability data, to form an abnormality recognition result together.
[0074] In the embodiment, the fusion of the frequency domain speech feature and the time domain speech feature fully utilizes the multi-dimensional characteristics of the speech signal, and improves the ability to distinguish different types of abnormalities. The probability data output by the model not only provides the confidence of the classification, but also realizes a flexible abnormality judgment strategy by setting a threshold. The quantitative abnormality recognition result provides an accurate basis for subsequent hierarchical disposal, and effectively improves the accuracy of abnormality detection and the robustness of the system in complex scenarios.
[0075] S130, based on the abnormality identification result, performing abnormality level determination to obtain an abnormality level determination result, determining an abnormality disposal measure based on the abnormality level determination result, and processing the original voice signal through the abnormality disposal measure.
[0076] The abnormality level determination result can be understood as a quantitative grading conclusion of the severity of the voice signal abnormality in combination with a preset abnormality level determination rule, for example, being divided into levels of slight, moderate, and severe. It directly reflects the influence of the abnormality on the voice quality or subsequent processing and provides clear level basis for subsequent disposal. The abnormality disposal measure can be understood as a targeted processing scheme matched according to the abnormality level determination result. Different levels correspond to different measures, such as simple noise reduction for slight abnormality and signal interruption and alarm for severe abnormality. The purpose is to eliminate or alleviate the influence of the abnormality through accurate and adaptive operation to ensure the effectiveness and stability of the voice signal.
[0077] Specifically, according to the classification labels and corresponding abnormality probability data in the abnormality identification result, the abnormality is divided into different levels in combination with a preset level determination rule to obtain an abnormality level determination result. Then, according to the disposal strategy library corresponding to each abnormality level, a specific abnormality disposal measure is matched, for example, basic noise reduction for slight level, signal repair for moderate level, and interruption transmission and alarm for severe level. Finally, the original voice signal is subjected to targeted processing through these measures.
[0078] In this embodiment, the level determination based on the quantitative probability and clear type makes the result more objective and accurate, avoiding subjective judgment bias. The binding mechanism of the level and the disposal measure ensures the timeliness and adaptability of the processing, which can reduce the interference of excessive processing on the signal through light measures and quickly suppress serious abnormality through heavy measures, effectively improving the efficiency and reliability of voice signal processing.
[0079] Optionally, based on the abnormality identification result, an abnormality level determination result is obtained, including: obtaining a mapping relationship between preset classification labels and level thresholds, matching at least one classification label in the abnormality identification result with the mapping relationship to obtain level thresholds corresponding to each classification label; comparing abnormality probability data corresponding to each classification label with the corresponding level threshold respectively to obtain comparison results of each classification label; calling a preset abnormality level determination rule, matching the comparison results of each classification label with the preset abnormality level determination rule, and determining an abnormality level determination result according to the matching result.
[0080] Specifically, a mapping relationship between preset classification labels and level thresholds can be obtained from configuration information or a storage space, for example, the threshold interval [0.3, 0.6, 0.9] corresponding to the abnormal label, respectively corresponding to the levels of slight, moderate, and severe; then the classification labels and corresponding abnormal probabilities are extracted from the abnormal identification result, and the level threshold of each label is matched according to the mapping relationship; then the abnormal probabilities of each label are compared with the corresponding threshold, for example, the probability 0.75 corresponding to the abnormal label exceeds the threshold 0.6, and the level tendency of a single label is obtained; finally, a preset abnormal level determination rule is called, and the final abnormal level determination result is determined by comprehensively determining the comparison results of all labels.
[0081] In the embodiment, the standardization of level determination is realized by the preset mapping relationship and threshold, avoiding subjective bias; the comprehensive determination of the comparison results of multiple labels and rules not only considers the severity of a single abnormality, but also takes into account the complex scene of multiple abnormalities coexisting, so that the level determination is more accurate and objective, and provides a reliable basis for matching subsequent disposal measures.
[0082] Optionally, an abnormal disposal measure is determined based on the abnormal level determination result, and the original voice signal is processed through the abnormal disposal measure, including: generating corresponding abnormal prompt information according to the abnormal prompt information generation template based on the abnormal level determination result; and / or, obtaining corresponding disposal measures from a preset disposal measure set based on the abnormal level determination result; determining the abnormal disposal measure based on the abnormal prompt information and / or the disposal measure, and processing the original voice signal through the abnormal disposal measure.
[0083] Specifically, according to the abnormal level determination result, the corresponding prompt information is automatically generated according to the preset abnormal prompt information generation template, for example, the moderate abnormality corresponds to the prompt text "there is an abnormal threatening voice, generate a risk report"; or the specific disposal operation corresponding to the abnormal level can also be matched and obtained from the preset disposal measure set; then the generated prompt information and / or obtained disposal measures are integrated to form a complete abnormal disposal measure, and the original voice signal is processed according to these measures, for example, the prompt information is output while starting the noise reduction algorithm.
[0084] In the embodiment, the abnormal state can be quickly fed back by generating prompt information through the template, and the interaction transparency is improved; the standardization and efficiency of the operation are ensured by matching the disposal measures from the preset set; and the flexible combination mechanism of "prompting + disposal" can meet the user's demand for knowing abnormal information, and can directly alleviate or eliminate the abnormality through targeted operation, which guarantees the processing efficiency while taking into account the scene adaptability, effectively improving the intelligence and user experience of voice signal processing.
[0085] The technical scheme of the embodiment determines the multi-frame voice data corresponding to the original voice signal to be processed, performs feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features, performs abnormality recognition processing on the frequency domain voice features and the time domain voice features based on a preset abnormality recognition model to obtain an abnormality recognition result, wherein the abnormality recognition result includes at least one classification label and abnormality probability data corresponding to the classification label, performs abnormality level determination based on the abnormality recognition result to obtain an abnormality level determination result, determines an abnormality disposal measure based on the abnormality level determination result, and processes the original voice signal through the abnormality disposal measure. The scheme can comprehensively capture abnormal information of the voice signal in frequency distribution and time variation by synchronously extracting frequency domain and time domain voice features, avoid missed detection caused by single dimension features, and greatly improve the accuracy of abnormality recognition. The abnormality recognition result includes classification labels and corresponding abnormality probabilities, which provides a quantitative basis for subsequent level determination, so that the level division is more accurate. The targeted disposal measure based on the abnormality level matching can take different strategies according to the severity of the abnormality, effectively solves the problem, avoids excessive processing to damage normal voice signals, and realizes efficient and refined processing. The whole process forms a closed loop from feature extraction to disposal execution, and takes into account the comprehensiveness of recognition and the adaptability of processing, which is suitable for various scenes such as voice communication and voice recognition, and can significantly improve the quality and reliability of voice signals.
[0086] Embodiment two
[0087] Figure 2 is a flowchart of a data processing method provided by the second embodiment of the application. The method of the embodiment is a further optimization of the method of the above-mentioned embodiment. Optionally, the preset abnormality recognition model includes a frequency domain feature processing module, a time domain feature processing module, a tone of voice feature extraction module, and a classification module. The frequency domain feature processing module is used to perform feature processing on the frequency domain voice features to obtain a frequency domain feature processing result. The time domain feature processing module is used to perform feature processing on the time domain voice features to obtain a time domain feature processing result. The frequency domain feature processing result and the time domain feature processing result are spliced to obtain feature splicing data. The feature splicing data is subjected to feature extraction by the tone of voice feature extraction module to obtain a tone of voice feature. The tone of voice feature is subjected to classification processing by the classification module to obtain a classification result, which is determined as the abnormality recognition result. As shown in Figure 2 , the method includes:
[0088] S210, determining multi-frame voice data corresponding to an original voice signal to be processed, performing feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features.
[0089] S220, performing feature processing on the frequency domain voice features by a frequency domain feature processing module to obtain a frequency domain feature processing result.
[0090] The frequency domain feature processing module can be specifically understood as a feature processing module constructed based on a convolutional neural network, which receives the frequency domain speech features extracted from the original speech signal, extracts local frequency domain features through multi-layer convolution and pooling operations, such as a convolution kernel with specific parameters, an activation function and normalization processing, and outputs a structured frequency domain feature processing result, such as a feature map. The core is to capture local details and spatial correlation in the frequency domain.
[0091] Specifically, the frequency domain feature processing module in the preset anomaly recognition model first receives the frequency domain speech features extracted from the original speech signal, and then performs local frequency domain feature extraction processing on the frequency domain speech features, and outputs the processed frequency domain feature processing result, which can be represented by a feature map, for example. Optionally, the frequency domain feature processing module is constructed based on a convolutional neural network. For example, the frequency domain feature processor gradually extracts local frequency domain features through 3 layers of convolution + pooling, the convolution kernel parameters are 3x3, the step is 1x1, the padding is Same Padding, and each layer is connected with ReLU activation and BatchNorm, and an output of 128x64xT feature map is obtained, and the frequency domain feature processing result is obtained.
[0092] In this embodiment, the frequency domain feature processing module constructed by the convolutional neural network extracts the local frequency domain features of the frequency domain speech features, so that the output feature map focuses on key information, improves the capture sensitivity and representation ability of local frequency domain anomalies, and lays a foundation for anomaly recognition accuracy.
[0093] S230, performing feature processing on the time domain speech features through a time domain feature processing module to obtain a time domain feature processing result.
[0094] The time domain feature processing module can be specifically understood as a feature extraction module constructed based on a one-way long short-term memory network. For time domain speech features, modeling is performed according to time sequence, one-way dependency in the time dimension is captured through a gating mechanism, time domain dynamic change rules are mined, and a time domain feature processing result reflecting time sequence correlation is output. The focus is on grasping the dynamic features of the speech signal in the time dimension.
[0095] Specifically, after receiving the time domain speech features extracted from the original speech signal through the time domain feature processing module in the preset abnormality recognition model, the one-way long short-term memory network (LSTM) models the features in the time sequence, selectively retains and transmits key time sequence information through the gating mechanism of the LSTM, captures the dynamic change rule of the speech signal in the time dimension, and finally outputs the time domain feature processing result that can reflect the time sequence correlation. For example, the time domain feature processing module receives the time domain speech features, processes the received time domain speech features through the one-way LSTM (64 units), captures the time domain dynamic pattern, outputs 64xT time sequence features, and then performs feature fusion, flattens the 128x64xT feature map output by the frequency domain feature processing module into 128xT, and concatenates the output of the LSTM according to the time step to form 192xT fusion features.
[0096] In this embodiment, the time domain speech features are processed by the time domain feature processing module, so that the time sequence dependence of the time domain features is accurately captured, the abnormal patterns of the speech signal in the time dimension are effectively mined, the problem of insufficient long-time dependence capture of the traditional time domain processing method is avoided, and the output features are more in line with the time sequence characteristics of the speech signal, providing more targeted time domain information support for the subsequent abnormality recognition model, and further improving the accuracy of the overall speech abnormality recognition.
[0097] S240, concatenating the frequency domain feature processing result and the time domain feature processing result to obtain feature concatenation data, performing feature extraction on the feature concatenation data through the tone feature extraction module to obtain the tone feature.
[0098] The tone feature extraction module can be specifically understood as a feature extraction module constructed based on a bidirectional long short-term memory network (bidirectional LSTM), which is used for feature processing of the frequency domain speech features, receives the feature concatenation data after the concatenation of the frequency domain and time domain features, learns the dynamic changes related to the tone (such as emotional fluctuations), the intonation (such as pitch and rhythm), and the non-verbal violation behavior from the forward and reverse directions, comprehensively captures the bidirectional time sequence features of the tone, and outputs the features accurately representing the tone characteristics.
[0099] Specifically, the structured frequency domain features output by the frequency domain feature processing module and the time-sequenced time domain features output by the time domain feature processing module are spliced according to a preset dimension to form feature splicing data that fuses frequency domain and time domain information; then the splicing data is input into a tone of voice feature extraction module constructed based on a bidirectional long short-term memory network (bidirectional LSTM), the bidirectional LSTM deeply mines the time sequence information in the splicing data from two directions of forward and reverse, focuses on capturing dynamic change features related to tone of voice (such as emotional ups and downs) and tone (such as pitch and rhythm), and finally outputs tone of voice features that can accurately represent the tone of voice characteristics.
[0100] In this embodiment, the preliminary fusion of frequency domain and time domain information is achieved through feature splicing, which provides more abundant basic data for the extraction of tone of voice features and avoids the limitation of insufficient information in a single feature dimension; the use of bidirectional LSTM fully considers the bidirectional correlation of tone of voice in time sequence, can more comprehensively and accurately capture tone of voice features, improves the recognition degree of the features, provides high-quality feature support for the judgment of tone of voice abnormalities involved in subsequent abnormality recognition, and further enhances the comprehensiveness and accuracy of the overall voice abnormality recognition.
[0101] S250, classifying the tone of voice features by a classification module to obtain a classification result, and determining the classification result as the abnormality recognition result.
[0102] The classification module can be specifically understood as a module for abnormality classification constructed based on a full connection layer and a Softmax layer.
[0103] Specifically, after receiving the tone of voice features output by the tone of voice feature extraction module, the classification module first performs nonlinear transformation and dimension integration on the features through the full connection layer, maps the high-dimensional features to a low-dimensional space corresponding to the abnormality types, then calculates the probability distribution of each type of abnormality (such as tone mutation and emotional abnormality) through the Softmax layer, outputs a classification result containing specific classification labels (such as specific abnormality types) and corresponding abnormality probabilities, and directly determines the result as the abnormality recognition result.
[0104] In this embodiment, the classification module accurately captures subtle abnormality features in the tone of voice, reduces the subjectivity and errors of manual judgment, the standardization of the classification process can ensure the consistency and repeatability of the recognition result, facilitates efficient application in different scenarios, and can improve the accuracy and robustness of abnormality recognition through continuous optimization of the classification model, thereby quickly and reliably completing the abnormality recognition task of the voice signal.
[0105] On the basis of the above embodiments, the frequency domain feature processing module is constructed based on a convolutional neural network; the time domain feature processing module is constructed based on a one-way long short-term memory network; the tone feature extraction module is constructed based on a bidirectional long short-term memory network; and the classification module is constructed based on a full connection layer and a Softmax layer.
[0106] Specifically, the frequency domain feature processing module adopts a convolutional neural network, extracts local key information in the frequency domain speech feature through multi-layer convolution and pooling operation, captures spatial correlation by using the sliding window characteristics of the convolution kernel, and outputs structured frequency domain features; the time domain feature processing module is based on a one-way long short-term memory network (LSTM), models the time domain speech feature according to the time sequence, focuses on capturing the one-way dependency relationship of the speech signal in the time dimension, and outputs the time-sequenced time domain feature; the tone feature extraction module uses bidirectional LSTM to learn the tone feature from both forward and reverse directions, comprehensively captures the bidirectional correlation of speech emotion and rhythm, and outputs the tone feature containing bidirectional time sequence information; and the classification module fuses the above three types of features, integrates multi-dimensional feature information through a full connection layer, calculates the probability distribution of each type of anomaly through a Softmax layer, and finally outputs the anomaly recognition result containing the classification label and the corresponding probability.
[0107] In one specific embodiment, the preset anomaly recognition model is a CNN-LSTM hybrid model that fuses the time domain and the frequency domain, including a frequency domain feature processing module, a time domain feature processing module, a tone feature extraction module, and a classification module; the frequency domain feature processing module (convolutional neural network) is mainly used for processing frequency domain features, and the time domain feature processing module (long short-term memory network) is mainly used for processing time domain features. By constructing such a hybrid model, different types of features are integrated together for modeling processing, so as to improve the performance and generalization ability of the model, such as Figure 3The flowchart of an abnormality recognition model processing is shown. First, the input layer is a time-frequency fusion feature matrix with a dimension of 68xT, wherein the frequency domain feature is a 64xT mel spectrum graph, and the time domain feature is 4xT (short-time energy, zero-crossing rate, amplitude envelope, and instantaneous frequency derivative). Then, it enters a double-branch structure. The frequency domain feature processing module gradually extracts local frequency domain features through 3 layers of convolution + pooling, the convolution kernel parameters are 3x3, the step is 1x1, the padding is Same Padding, and each layer is connected with ReLU activation and BatchNorm, and the output is a 128x64xT feature graph. The time domain feature processing module inputs the 4xT time domain feature, captures the time domain dynamic pattern through a one-way LSTM (64 units), and outputs a 64xT time sequence feature. Then, the feature fusion is performed, the 128x64xT feature graph output by the CNN is flattened into 128xT, and the output of the LSTM is spliced according to the time step to form a 192xT fusion feature. Then, the tone and intonation feature extraction module captures the long-time dependence relationship such as tone and intonation changes, outputs a 256xT feature, and the Dropout ratio is set to 0.5 (only activated during training) to prevent overfitting. The classification module maps the 256-dimensional feature to 4 class labels through a fully connected layer + Softmax layer, which are normal, vulgar, violent and political sensitive, respectively. When designing the loss function L, the weighted cross-entropy loss is used to reduce the influence of the class imbalance problem, and the loss function L formula is as follows, wherein N represents the batch size, c is the class index, w c is the class weight, which is set according to the proportion of the training set samples. Since the number of normal samples is much larger than that of the other three types of samples, a larger weight is given to the class with fewer samples, N total represents the total number of samples, C represents the total number of classes, N i represents the number of i-th class samples.
[0108]
[0109] The trainer of the preset abnormality recognition model adopts the Adam optimization algorithm, the initial learning rate is 0.0001, the learning rate is attenuated to 0.5 of the original value every 10 epochs, the training is terminated if the validation set loss does not decrease for 5 consecutive epochs, and the trained preset abnormality recognition model is obtained. Optionally, feedback optimization can be performed through incremental learning, the model is updated every month, 1000 new labeled samples are added, and 5 epochs are trained. The A / B testing method is adopted, 10% of the traffic is released in gray for the new model, and the model is put into full production after verification of the accuracy and false positive rate.
[0110] S260, based on the abnormality recognition result, an abnormality level determination result is obtained, an abnormality disposal measure is determined based on the abnormality level determination result, and the original voice signal is processed through the abnormality disposal measure.
[0111] The technical scheme of the embodiment determines the multi-frame voice data corresponding to the original voice signal to be processed, performs feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features, performs feature processing on the frequency domain voice features through a frequency domain feature processing module to obtain a frequency domain feature processing result, performs feature processing on the time domain voice features through a time domain feature processing module to obtain a time domain feature processing result, splices the frequency domain feature processing result and the time domain feature processing result to obtain feature splicing data, extracts features from the feature splicing data through a tone of voice feature extraction module to obtain a tone of voice feature, performs classification processing on the tone of voice feature through a classification module to obtain a classification result, determines the classification result as an abnormality recognition result, constructs the frequency domain feature processing module based on a convolutional neural network, constructs the time domain feature processing module based on a one-way long short-term memory network, constructs the tone of voice feature extraction module based on a bidirectional long short-term memory network, and constructs the classification module based on a full connection layer and a Softmax layer. Abnormality level judgment is performed based on the abnormality recognition result to obtain an abnormality level judgment result, abnormality disposal measures are determined based on the abnormality level judgment result, and the original voice signal is processed through the abnormality disposal measures. The scheme can comprehensively capture abnormal information of the voice signal in frequency distribution and time change by synchronously extracting frequency domain and time domain voice features, avoid missed detection caused by a single dimension feature, greatly improve the accuracy of abnormality recognition, and provide a quantitative basis for subsequent level judgment so that level division is more accurate. The targeted disposal measures based on the abnormality level matching can take different strategies according to the severity of the abnormality, effectively solve the problem, avoid damage to normal voice signals caused by excessive processing, and realize efficient and refined processing. The entire process forms a closed loop from feature extraction to disposal execution, realizes accurate recognition of voice abnormalities, can take corresponding measures according to level differences, significantly enhances the practicality and response efficiency of the system, balances the comprehensiveness of recognition and the adaptability of processing, is suitable for various scenes such as voice communication and voice recognition, and can significantly improve the quality and reliability of voice signals.
[0112] Embodiment three
[0113] Figure 4 is a structural schematic diagram of a data processing device provided by the embodiment three of the application. As shown in the figure, Figure 4 the device comprises:
[0114] The voice feature determination module 410 is configured to determine multi-frame voice data corresponding to an original voice signal to be processed, perform feature extraction processing on the multi-frame voice data to obtain frequency domain voice features and time domain voice features, perform feature processing on the frequency domain voice features through a frequency domain feature processing module to obtain a frequency domain feature processing result, perform feature processing on the time domain voice features through a time domain feature processing module to obtain a time domain feature processing result, splice the frequency domain feature processing result and the time domain feature processing result to obtain feature splicing data, extract features from the feature splicing data through a tone of voice feature extraction module to obtain a tone of voice feature, perform classification processing on the tone of voice feature through a classification module to obtain a classification result, and determine the classification result as an abnormality recognition result.
[0115] The abnormality recognition result determination module 420 is configured to perform abnormality recognition processing on the frequency domain voice feature and the time domain voice feature based on a preset abnormality recognition model to obtain an abnormality recognition result, wherein the abnormality recognition result comprises at least one classification label and abnormality probability data corresponding to the classification label.
[0116] The abnormality level determination and abnormality processing module 430 is configured to perform abnormality level determination based on the abnormality recognition result to obtain an abnormality level determination result, determine an abnormality disposal measure based on the abnormality level determination result, and perform processing on the original voice signal through the abnormality disposal measure.
[0117] The technical scheme of the embodiment determines the multiple frames of voice data corresponding to the original voice signal to be processed through the voice feature determination module, performs feature extraction processing on the multiple frames of voice data to obtain the frequency domain voice feature and the time domain voice feature, performs abnormality recognition processing on the frequency domain voice feature and the time domain voice feature based on the abnormality recognition result determination module to obtain an abnormality recognition result, wherein the abnormality recognition result comprises at least one classification label and abnormality probability data corresponding to the classification label, performs abnormality level determination based on the abnormality recognition result through the abnormality level determination and abnormality processing module to obtain an abnormality level determination result, determines an abnormality disposal measure based on the abnormality level determination result, and performs processing on the original voice signal through the abnormality disposal measure. The scheme can comprehensively capture abnormal information of the voice signal in frequency distribution and time variation by synchronously extracting frequency domain and time domain voice features, avoid missing detection caused by a single dimension feature, and greatly improve the accuracy of abnormality recognition. The abnormality recognition result contains a classification label and corresponding abnormality probability, which provides a quantitative basis for subsequent level determination, so that the level division is more accurate. The targeted disposal measure based on the abnormality level can take different strategies according to the severity of the abnormality, effectively solves the problem, avoids excessive processing that damages normal voice signals, and realizes efficient and fine processing. The entire process forms a closed loop from feature extraction to disposal execution, taking into account the comprehensiveness of recognition and the adaptability of processing, and is suitable for various scenarios such as voice communication and voice recognition, which can significantly improve the quality and reliability of voice signals.
[0118] On the basis of the above-mentioned embodiments, the voice feature determination module 410 can comprise a voice feature extraction unit. The voice feature extraction unit is configured to perform feature extraction on the multiple frames of voice data through a preset frequency domain feature extraction algorithm to obtain a frequency domain voice feature, wherein the frequency domain voice feature comprises a mel spectrum feature; perform feature extraction on the multiple frames of voice data through a preset time domain feature extraction algorithm to obtain multiple time domain features, and perform fusion processing on the multiple time domain features to obtain a time domain voice feature, wherein the time domain feature comprises a short-time energy feature, a zero-crossing rate feature, an amplitude envelope feature, and an instantaneous frequency derivative feature.
[0119] Optionally, the preset anomaly recognition model comprises a frequency domain feature processing module, a time domain feature processing module, a tone of voice feature extraction module, and a classification module; the anomaly recognition result determination module 420 is specifically configured to perform feature processing on the frequency domain voice features through the frequency domain feature processing module to obtain a frequency domain feature processing result; perform feature processing on the time domain voice features through the time domain feature processing module to obtain a time domain feature processing result; splice the frequency domain feature processing result and the time domain feature processing result to obtain feature splicing data, perform feature extraction on the feature splicing data through the tone of voice feature extraction module to obtain a tone of voice feature; perform classification processing on the tone of voice feature through the classification module to obtain a classification result, and determine the classification result as the anomaly recognition result.
[0120] Optionally, the frequency domain feature processing module is constructed based on a convolutional neural network; the time domain feature processing module is constructed based on a one-way long short-term memory network; the tone of voice feature extraction module is constructed based on a bidirectional long short-term memory network; and the classification module is constructed based on a full connection layer and a Softmax layer.
[0121] Optionally, the anomaly level determination and anomaly processing module 430 is specifically configured to obtain a mapping relationship of preset classification labels and level thresholds, match at least one classification label in the anomaly recognition result with the mapping relationship to obtain a level threshold corresponding to each classification label, compare the anomaly probability data corresponding to each classification label with the corresponding level threshold respectively to obtain a comparison result of each classification label, call a preset anomaly level determination rule, match the comparison result of each classification label with the preset anomaly level determination rule, and determine an anomaly level determination result according to the matching result.
[0122] Optionally, the anomaly level determination and anomaly processing module 430 is specifically configured to generate corresponding anomaly prompt information according to an anomaly prompt information generation template based on the anomaly level determination result; and / or, obtain corresponding disposal measures from a preset disposal measure set based on the anomaly level determination result; determine an anomaly disposal measure based on the anomaly prompt information and / or the disposal measure, and process the original voice signal through the anomaly disposal measure.
[0123] Optionally, the voice feature determination module 410 is specifically configured to obtain an original voice signal in an authorized voice call process, determine the original voice signal as a to-be-processed original voice signal, and perform a preprocessing operation on the original voice signal to obtain a plurality of frames of voice data, wherein the preprocessing operation at least comprises noise reduction processing and frame windowing processing.
[0124] The data processing apparatus provided in the embodiments of the present application can execute the data processing method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0125] Embodiment four
[0126] Figure 5 FIG. 1 is a structural diagram of an electronic device according to an embodiment of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0127] As shown in FIG. 1, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected in communication with the at least one processor 11, wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14. Figure 5
[0128] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0129] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as a data processing method.
[0130] In some embodiments, the data processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded onto and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. One or more steps of the data processing method described above can be performed when the computer program is loaded into RAM 13 and executed by processor 11. Alternatively, in other embodiments, processor 11 can be configured, by any suitable means (e.g., by means of firmware), to perform the data processing method.
[0131] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0132] Computer programs used to implement the data processing method of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / operations specified in the flowcharts and / or the block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or entirely on a remote machine or server.
[0133] Embodiment Five
[0134] Embodiment five of the present application also provides a computer readable storage medium, which stores computer instructions for causing a processor to execute a data processing method, the method comprising: determining a plurality of frames of voice data corresponding to an original voice signal to be processed, performing feature extraction processing on the plurality of frames of voice data to obtain frequency domain voice features and time domain voice features;
[0135] The abnormality recognition result includes at least one classification label and abnormal probability data corresponding to the classification label.
[0136] The abnormality level determination result is used to determine an abnormality disposal measure, and the original voice signal is processed through the abnormality disposal measure.
[0137] In the context of the present application, the computer-readable storage medium can be a tangible medium that can contain or store the computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more wires, portable computer disks, hard disk drives, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0138] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0139] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0140] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0141] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure can be achieved, and the present disclosure is not limited herein.
[0142] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A data processing method, characterized by, The method comprises the following steps: determining a plurality of frames of voice data corresponding to the original voice signal to be processed, performing feature extraction processing on the plurality of frames of voice data to obtain frequency domain voice features and time domain voice features; performing abnormality recognition processing on the frequency domain voice features and the time domain voice features based on a preset abnormality recognition model to obtain an abnormality recognition result, wherein the abnormality recognition result comprises at least one classification label and abnormality probability data corresponding to the classification label; performing abnormality level determination based on the abnormality recognition result to obtain an abnormality level determination result, determining an abnormality disposal measure based on the abnormality level determination result, and processing the original voice signal through the abnormality disposal measure.
2. The method of claim 1, wherein, The method comprises the following steps: performing feature extraction on the plurality of frames of voice data through a preset frequency domain feature extraction algorithm to obtain the frequency domain voice features, wherein the frequency domain voice features comprise mel spectrum features; performing feature extraction on the plurality of frames of voice data through a preset time domain feature extraction algorithm to obtain a plurality of time domain features, performing fusion processing on the plurality of time domain features to obtain the time domain voice features, wherein the time domain features comprise short-time energy features, zero-crossing rate features, amplitude envelope features, and instantaneous frequency derivative features.
3. The method of claim 1, wherein, The preset abnormality recognition model comprises a frequency domain feature processing module, a time domain feature processing module, a tone of voice feature extraction module, and a classification module. The method comprises the following steps: performing feature processing on the frequency domain voice features through the frequency domain feature processing module to obtain a frequency domain feature processing result; performing feature processing on the time domain voice features through the time domain feature processing module to obtain a time domain feature processing result; performing feature extraction on the frequency domain feature processing result and the time domain feature processing result to obtain tone of voice features; performing classification processing on the tone of voice features through the classification module to obtain a classification result, and determining the classification result as the abnormality recognition result.
4. The method of claim 3, wherein, The frequency domain feature processing module is constructed based on a convolutional neural network; the time domain feature processing module is constructed based on a unidirectional long short-term memory network; the tone of voice feature extraction module is constructed based on a bidirectional long short-term memory network; and the classification module is constructed based on a fully connected layer and a Softmax layer.
5. The method of claim 1, wherein, The method comprises the following steps: obtaining a mapping relationship between preset classification labels and level thresholds, matching at least one classification label in the abnormality recognition result with the mapping relationship to obtain a level threshold corresponding to each classification label; comparing the abnormality probability data corresponding to each classification label with the corresponding level threshold respectively to obtain a comparison result of each classification label; and performing abnormality level determination based on the comparison result of each classification label to obtain an abnormality level determination result. The preset abnormal level determination rule is called to match the comparison results of the classification labels with the preset abnormal level determination rule, and the abnormal level determination result is determined according to the matching result.
6. The method of claim 1, wherein, The abnormal treatment measure is determined based on the abnormal level determination result, and the original voice signal is processed through the abnormal treatment measure, including: According to the abnormal level determination result, the corresponding abnormal prompt information is generated according to the abnormal prompt information generation template; and / or, According to the abnormal level determination result, the corresponding treatment measure is obtained from the preset treatment measure set; According to the abnormal prompt information and / or the treatment measure, the abnormal treatment measure is determined, and the original voice signal is processed through the abnormal treatment measure.
7. The method of claim 1, wherein, The multiple frames of voice data corresponding to the original voice signal to be processed are determined, including: The original voice signal in the authorized voice call process is obtained, and the original voice signal is determined as the original voice signal to be processed; The original voice signal is preprocessed to obtain multiple frames of voice data, wherein the preprocessing operation at least includes noise reduction processing and frame windowing processing.
8. A data processing apparatus, characterized by, It includes: The voice feature determination module is used to determine the multiple frames of voice data corresponding to the original voice signal to be processed, and the feature extraction processing is performed on the multiple frames of voice data to obtain the frequency domain voice feature and the time domain voice feature; The abnormal recognition result determination module is used to perform abnormal recognition processing on the frequency domain voice feature and the time domain voice feature through the preset abnormal recognition model to obtain an abnormal recognition result, wherein the abnormal recognition result includes at least one classification label and abnormal probability data corresponding to the classification label; The abnormal level determination and abnormal processing module is used to perform abnormal level determination based on the abnormal recognition result to obtain an abnormal level determination result, determine an abnormal treatment measure based on the abnormal level determination result, and process the original voice signal through the abnormal treatment measure.
9. An electronic device, comprising: The electronic device includes: At least one processor; and The memory is in communication connection with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data processing method in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to enable the processor to execute the data processing method in any one of claims 1-7.