Audio security monitoring method and device, electronic equipment and computer program product
By segmenting and multi-granularity review of audio data streams, combined with feature extraction and speech recognition, the shortcomings of traditional audio security review methods in terms of real-time performance and contextual analysis are addressed, enabling accurate identification and timely intervention in complex contexts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional audio security auditing methods are insufficient in terms of real-time performance, contextual analysis capabilities, and automation adaptability, making it difficult to timely and comprehensively identify hidden violations in complex contexts under the high-speed data transmission scenarios of 5G.
By segmenting the audio data stream, segmented audio data of different acquisition durations are generated, and multi-granularity review is performed based on preset rules. Combining feature extraction, speech recognition, and text classification, real-time security monitoring of the audio data stream is achieved, including feature vector matching and contextual semantic analysis.
It enables accurate identification of potential risks in complex contexts, ensuring the comprehensiveness and accuracy of review results, and can trigger automated intervention measures in a timely manner when unsafe content is detected.
Smart Images

Figure CN121768367A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and in particular to an audio security monitoring method, apparatus, electronic device, and computer program product. Background Technology
[0002] In today's digital communication era, the security review of media resources has become a crucial step in ensuring information compliance and security. However, traditional audio security review methods have the following shortcomings: 1) Insufficient real-time performance: Security audits relying on technologies such as keyword matching and spectrum analysis can only identify explicit inappropriate content in audio and cannot handle hidden violations in complex contexts. Especially in scenarios with high-speed data transmission using 5G technology (e.g., multi-party calls, complex interactions), timely and comprehensive audits are not possible, making it difficult to meet high real-time requirements.
[0003] 2) Limited contextual analysis capabilities: Traditional solutions tend to overlook the textual context and relationships when processing audio content. Even with the addition of a pattern recognition system, they can only identify specific audio patterns and lack a deep understanding of the semantic coherence and contextual relationships of long-term audio. Accuracy and efficiency drop significantly when reviewing audio with multiple themes and frequent theme changes.
[0004] 3) Insufficient automation and adaptability: Security audits that rely on large amounts of labeled data struggle to quickly adapt to emerging threats of illegal content. Furthermore, while manual audits are highly accurate, they are inefficient and cannot handle the demands of large-scale, real-time audio data processing.
[0005] In summary, there is an urgent need to propose an audio security auditing method to solve the above problems. Summary of the Invention
[0006] This disclosure is made in view of the above-mentioned problems. This disclosure provides an audio security monitoring method, apparatus, electronic device, and computer program product.
[0007] According to one aspect of this disclosure, an audio security monitoring method is provided, the method comprising: Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method includes: acquiring an audio data stream in real time; segmenting the audio data stream to generate multiple segments of audio data corresponding to multiple acquisition durations; and performing real-time security audits on each of the multiple segments of audio data corresponding to multiple acquisition durations based on preset rules to determine whether the audio data stream is secure.
[0008] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the audio data stream is segmented to generate multiple segments of audio data corresponding to multiple acquisition durations, including: segmenting the audio data stream to generate a first segment of audio data and a second segment of audio data, wherein the first acquisition duration of the first segment of audio data is less than the second acquisition duration of the second segment of audio data.
[0009] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method involves performing security audits on each of multiple audio segments corresponding to multiple acquisition durations based on preset rules to determine whether the audio data stream is secure. This includes: performing security audits on the first segment audio data and the second segment audio data respectively based on a first preset rule to generate a first segment audit result and a second segment audit result; and determining whether the audio data stream is secure based on a second preset rule, the first segment audit result, and / or the second segment audit result.
[0010] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, a security audit is performed on a first segment of audio data and a second segment of audio data based on a first preset rule, generating a first segment audit result and a second segment audit result, including: extracting features from the first segment of audio data to generate a first audio feature vector; determining the similarity between the preset harmful feature vector and the first audio feature vector, and using the similarity as the first segment audit result.
[0011] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, wherein, based on a first preset rule, security audits are performed on the first segment audio data and the second segment audio data respectively to generate the first segment audit result and the second segment audit result, the method further includes: performing speech recognition on the second segment audio data based on a preset recognition model to generate the second audio recognition text; determining the category and probability value of the second audio recognition text based on a preset classification model, and taking the highest probability value in the abnormal category as the second segment audit result.
[0012] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method for extracting features from the first segment of audio data to generate a first audio feature vector includes: performing frame-by-frame processing on the first segment of audio data to generate multiple frames of audio data; performing frequency domain transformation on each of the multiple frames of audio data to generate a spectral distribution; generating a feature vector for each frame of audio data based on the spectral distribution; and concatenating each feature vector to generate a first audio feature vector corresponding to the first segment of audio data.
[0013] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method involves performing speech recognition on second segment audio data based on a preset recognition model to generate second audio recognition text, including: training an initial recognition model based on a labeled dataset corresponding to a historical audio data stream to generate a preset recognition model; and performing speech recognition on second segment audio data based on the preset recognition model to generate second audio recognition text.
[0014] Furthermore, according to one aspect of the audio security monitoring method of this disclosure, determining whether an audio data stream is secure based on a second preset rule, a first segment review result, and / or a second segment review result includes: determining that the audio data stream is insecure when the first segment review result meets a first threshold or the second segment review result meets a second threshold; when, within a second collection period, each first segment review result fails to meet the first threshold and the second segment review result fails to meet the second threshold, calculating a comprehensive segment review result based on each first segment review result and the second segment review result; and if the comprehensive segment review result meets a third threshold, determining that the audio data stream is insecure.
[0015] Furthermore, according to one aspect of the audio security monitoring method of this disclosure, the method further includes: performing a preset intervention when the audio data stream is determined to be insecure, wherein the preset intervention includes one or more of the following: sending a warning message, blocking the transmission of insecure segmented audio data without interrupting the call, and interrupting the call.
[0016] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method further includes: determining a preset harmful feature vector based on historical unsafe audio data streams.
[0017] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method further includes: training an initial classification model based on a labeled dataset corresponding to sample text to generate a preset classification model.
[0018] Furthermore, according to one aspect of the audio security monitoring method disclosed herein, the method further includes: processing first segment audio data and second segment audio data, wherein the processing includes one or more of the following: denoising, enhancement, and format conversion.
[0019] According to another aspect of this disclosure, an audio security monitoring device is provided, comprising: an acquisition module for acquiring an audio data stream in real time; a processing module for segmenting the audio data stream to generate multiple segments of audio data corresponding to multiple acquisition durations; and an auditing module for performing real-time security audits on each of the multiple segments of audio data corresponding to multiple acquisition durations based on preset rules to determine whether the audio data stream is secure.
[0020] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the processing module includes: a segmentation processing submodule, used to segment the audio data stream to generate first segmented audio data and second segmented audio data, wherein the first acquisition duration of the first segmented audio data is less than the second acquisition duration of the second segmented audio data.
[0021] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the audit module includes: a first audit submodule, configured to perform security audits on the first segment audio data and the second segment audio data respectively based on a first preset rule, and generate a first segment audit result and a second segment audit result; and a second audit submodule, configured to determine whether the audio data stream is secure based on the second preset rule, the first segment audit result and / or the second segment audit result.
[0022] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the first review submodule includes: an extraction unit, used to extract features from the first segmented audio data to generate a first audio feature vector; and a matching unit, used to determine the similarity between a preset harmful feature vector and the first audio feature vector, and to use the similarity as the first segment review result.
[0023] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the first review submodule further includes: a recognition unit, used to perform speech recognition on the second segment audio data based on a preset recognition model to generate second audio recognition text; and a classification unit, used to determine the category and probability value of the second audio recognition text based on a preset classification model, and to take the highest probability value in the abnormal category as the second segment review result.
[0024] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the extraction unit includes: a framing subunit for performing framing processing on the first segmented audio data to generate multiple frames of audio data; a conversion subunit for performing frequency domain conversion on each of the multiple frames of audio data to generate a spectral distribution; a calculation subunit for generating a feature vector for each frame of audio data based on the spectral distribution; and a splicing subunit for splicing each feature vector to generate a first audio feature vector corresponding to the first segmented audio data.
[0025] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the recognition unit includes: a training subunit, used to train an initial recognition model based on a labeled dataset corresponding to a historical audio data stream, and generate a preset recognition model; and a recognition subunit, used to perform speech recognition on second segmented audio data based on the preset recognition model, and generate second audio recognition text.
[0026] Furthermore, according to one aspect of the audio security monitoring device of this disclosure, the second audit submodule includes: a separate audit unit, used to determine that the audio data stream is insecure when the first segment audit result meets a first threshold or the second segment audit result meets a second threshold; and a comprehensive audit unit, used to calculate a comprehensive segment audit result based on each first segment audit result and the second segment audit result when, within a second acquisition duration, each first segment audit result does not meet the first threshold and the second segment audit result does not meet the second threshold, and if the comprehensive segment audit result meets a third threshold, then the audio data stream is determined to be insecure.
[0027] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the device further includes: an intervention module for performing a preset intervention when the audio data stream is determined to be insecure, wherein the preset intervention includes one or more of the following: sending a warning message, blocking the transmission of insecure segmented audio data without interrupting the call, and interrupting the call.
[0028] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the device further includes: a pre-configuration module for determining a preset harmful feature vector based on historical unsafe audio data streams.
[0029] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the device further includes: a pre-training module for training an initial classification model based on a labeled dataset corresponding to sample text, thereby generating a preset classification model.
[0030] Furthermore, according to one aspect of the audio security monitoring device disclosed herein, the device further includes: a preprocessing module for processing first segment audio data and second segment audio data, wherein the processing includes one or more of the following: noise reduction, enhancement, and format conversion.
[0031] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the electronic device to perform the audio security monitoring method as described above.
[0032] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, it implements the audio security monitoring method as described above.
[0033] As will be described in detail below, the audio security monitoring method according to embodiments of this disclosure achieves a multi-granularity review mechanism by segmenting the audio data stream based on multiple acquisition durations, while taking into account both real-time security monitoring and in-depth analysis, and more accurately identifying potential risks in complex contexts; and through high-precision speech recognition and understanding of contextual information, it achieves accurate classification of the audio data stream, ensuring the comprehensiveness and accuracy of the review results; at the same time, it introduces a real-time automated review and intervention mechanism, which can immediately trigger automated intervention measures when unsafe content is detected.
[0034] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description
[0035] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0036] Figure 1 This is a schematic diagram illustrating an application scenario of the audio security monitoring method according to an embodiment of the present disclosure.
[0037] Figure 2 This is a flowchart illustrating an audio security monitoring method according to an embodiment of the present disclosure.
[0038] Figure 3 This is an overall flowchart illustrating an audio security monitoring method according to an embodiment of the present disclosure.
[0039] Figure 4 This is a schematic diagram of an audio security monitoring device according to an embodiment of the present disclosure.
[0040] Figure 5 This is a hardware block diagram illustrating an electronic device according to an embodiment of the present disclosure.
[0041] Figure 6 This is a schematic diagram illustrating a computer program product according to an embodiment of the present disclosure. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0043] First, refer to Figure 1 Overview of application scenarios according to embodiments of this disclosure.
[0044] Figure 1 This is a schematic diagram illustrating an application scenario of the audio security monitoring method according to an embodiment of this disclosure. For example... Figure 1 As shown, the application scenario includes at least: an audio security monitoring device 10, a terminal device 21, and a terminal device 22, wherein the terminal device 21 and the terminal device 22 are the two parties in a call, and the audio security monitoring device 10 performs security monitoring on the audio of the call between the two parties.
[0045] The audio security monitoring device 10 can be a single physical server, a server cluster or distributed system consisting of at least two physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Considering the real-time requirements of the application scenario, the audio security monitoring device 10 needs to have high recognition, calculation, and processing capabilities.
[0046] Terminal devices 21 and 22 can be devices capable of making calls. Examples include mobile phones, tablets, laptops, desktop computers, and smart TVs. It should be noted that this disclosure does not limit the number or type of terminal devices; that is, the scope of protection of this disclosure covers both two-party and multi-party calls. However, for ease of description, the following example uses a two-party call between terminal devices 21 and 22. Furthermore, terminal devices can also be any other devices with call functionality.
[0047] Figure 2 This is a flowchart illustrating an audio security monitoring method according to an embodiment of the present disclosure. Figure 2 As shown, the audio security monitoring method applied to the audio security monitoring device 10 in this embodiment of the present disclosure may include at least the following steps.
[0048] In step S201, the audio data stream is acquired in real time. As described above, the audio security monitoring device 10 can capture the audio data stream during the call between terminal device 21 and terminal device 22 in real time.
[0049] It should be noted that the captured audio data stream may include all participants in the call. In step S202, the audio data stream is segmented to generate segmented audio data corresponding to multiple acquisition durations. That is, after capturing the real-time audio data stream in step S201, the audio security monitoring device 10 can first segment the audio data stream at different frequencies (i.e., different acquisition durations) to provide a data foundation for establishing a multi-granularity audit mechanism.
[0050] In one embodiment of this disclosure, the audio data stream can be segmented based on two different acquisition durations to generate a first segment audio data and a second segment audio data. The first acquisition duration of the first segment audio data is shorter than the second acquisition duration of the second segment audio data. In other words, the first segment audio data can be understood as a speech segment obtained through high-frequency sampling (hereinafter referred to as a high-frequency sampled speech segment), and the second segment audio data can be understood as a speech segment obtained through low-frequency sampling (hereinafter referred to as a low-frequency sampled speech segment).
[0051] In step S203, based on preset rules, each segment of audio data corresponding to multiple acquisition durations undergoes real-time security verification to determine whether the audio data stream is secure. In other words, the audio security monitoring device 10 performs multi-granular verification on the high-frequency and low-frequency sampled speech segments obtained in step S202. This allows for rapid analysis of the high-frequency sampled speech segments while simultaneously achieving a deep understanding of the contextual semantic information of the low-frequency sampled speech segments, balancing real-time performance and accuracy. Specifically, this will be discussed in detail later. Figure 3 Further explanation is needed.
[0052] Figure 3 This is an overall flowchart illustrating an audio security monitoring method according to an embodiment of the present disclosure. Figure 3 As shown, the overall process of the audio security monitoring method in this embodiment includes the following steps.
[0053] In step S1, the audio data stream is acquired in real time. This step is the same as step S201 described above, and will not be repeated here.
[0054] In step S2, the audio data stream is segmented to generate multiple audio segments corresponding to the acquisition duration (e.g., the first audio segment and the second audio segment).
[0055] In one example embodiment of this disclosure, the first acquisition duration of the high-frequency sampling voice segment can be 30 seconds, and the second acquisition duration of the low-frequency sampling voice segment can be 3 minutes. That is, the audio security monitoring device 10 generates a high-frequency sampling voice segment every 30 seconds and a low-frequency sampling voice segment every 3 minutes based on the call audio data stream between terminal device 21 and terminal device 22, and the two are performed simultaneously without interference.
[0056] It is understood that this step is the same as step S202 mentioned above. In this step, high-frequency sampled speech segments can help to quickly capture instantaneous changes in audio, while low-frequency sampled speech segments can help to analyze speech features over a longer time range. In other words, the segmentation processing in this step provides a data foundation for the multi-granularity review in the subsequent steps. The audio security monitoring device 10 can then perform security reviews on the first segmented audio data and the second segmented audio data based on the first preset rule, generating the first segment review result and the second segment review result. For details, please refer to steps S3-S6.
[0057] It should be noted that after the audio security monitoring device 10 completes the segmentation process, it can further process the high-frequency sampled speech segments and the low-frequency sampled speech segments, including but not limited to noise reduction, enhancement, and format conversion, in order to optimize their quality and reduce noise interference.
[0058] In step S3, feature extraction is performed on the high-frequency sampled speech segment obtained in step S2 to generate a first audio feature vector. Specifically, the feature extraction method may include: S3.1 Perform frame segmentation on the first segment of audio data to generate multiple frames of audio data. Specifically, the audio data is a continuous time series. In order to capture its local features, the audio usually needs to be segmented into frames to divide it into several short frames. Each frame is treated as an independent unit for subsequent feature extraction.
[0059] In one example embodiment of this disclosure, the length of each frame can be between 20 and 40 milliseconds, which can cover the pitch period of the speech while ensuring that the signal characteristics within the frame are relatively stable.
[0060] S3.2 Perform frequency domain transformation on each of the multiple frames of audio data to generate a spectral distribution.
[0061] In one embodiment of this disclosure, each frame of audio data can be frequency domain converted using the following formula to convert the time-domain audio data to the frequency domain, and then a spectrum diagram can be obtained to show the frequency components at different time points.
[0062]
[0063] in, It is time and frequency The short-time Fourier transform result on the time domain, x[n] is the time-domain audio data. It is a window function, where n is continuous time and j is the imaginary unit in a complex number.
[0064] S3.3. Based on the spectral distribution, generate the feature vector of each frame of audio data.
[0065] In one embodiment of this disclosure, the spectrogram obtained in step S3.2 is processed by a Mel filter bank and its logarithm is taken to obtain a log-Mel spectrum. Then, the log-Mel spectrum is subjected to a Discrete Cosine Transform (DCT) to compress the data and retain the first few coefficients (e.g., 12-13). These coefficients can accurately represent the timbre and pitch characteristics of the audio (e.g., the coefficient distribution of abusive language differs significantly from that of normal conversation, while eliminating interference from differences in pitch and speech rate, ensuring that the same offensive words spoken by different speakers can be classified into the same feature category).
[0066] The process of obtaining the above coefficients can be shown by the following formula:
[0067] in, This is the spectrum after passing through the Mel filter, where K is the number of coefficients retained. It represents the number of Mel filters.
[0068] S3.4. Concatenate each feature vector to generate the first audio feature vector corresponding to the first segment of audio data. Specifically, concatenate the coefficients of all frames within the same high-frequency sampled speech segment to form a high-dimensional audio feature vector fn. This fn can capture the features of the audio data in multiple dimensions. In other words, fn integrates the instantaneous features of the entire 30-second audio segment, preserving the local violation information of each frame while forming a standardized mathematical expression, which can be used as input data for the subsequent step S4 to achieve quantitative comparison with the preset harmful feature vector.
[0069] In step S4, similarity matching is performed on the feature vector fn from step S3. Specifically, this can be based on a preset harmful feature vector to determine the similarity between the first audio feature vector and the preset harmful feature vector, and the similarity is used as the first segmentation review result.
[0070] In one embodiment of this disclosure, the audio security monitoring device 10 may have a pre-set harmful feature vector knowledge base. This knowledge base is built upon historical unsafe audio data streams, which may include: a large number of collected publicly available harmful audio datasets, historical violation audio records, typical external violation cases, etc., covering various known categories of violation and inappropriate content. The knowledge base includes a large number of labeled harmful feature vectors.
[0071] The feature vector fn obtained in step S3 is matched with the knowledge base to quickly determine whether there is a security risk in the high-frequency sampled speech segment.
[0072] In one embodiment of this disclosure, the similarity between the feature vector fn and the preset harmful features can be determined by the following formula. The higher the similarity, the closer the current high-frequency sampled speech segment is to the features of the known illegal audio, the higher the risk of violation, and the less safe it is. Conversely, the higher the similarity, the more distant the current high-frequency sampled speech segment is from the features of the known illegal audio, the lower the risk of violation, and the safer it is.
[0073]
[0074] in, This represents the inner product of the input feature vector fn and the preset harmful feature vector. Let fn and fn represent the norms of the input feature vector and the preset harmful feature vector, respectively. This represents the cosine similarity between the input feature vector fn and the preset harmful feature vector.
[0075] It should be noted that, in one embodiment of this disclosure, when performing knowledge base matching, an approximate nearest neighbor algorithm can be used to quickly retrieve the preset harmful feature vector that is closest to the input feature vector fn in the knowledge base, and the cosine similarity between the preset harmful feature vector and the predefined harmful feature vector is used as its risk score in the high-frequency sampled speech segment, i.e., the first segment review result (denoted as fn). This is to allow for subsequent assessment of the audio data stream (i.e., the call) based on the results of the first segment review and the second preset rule, in order to determine whether the audio data stream (i.e., the call) is secure.
[0076] It is understandable that steps S3-S4 above are security audits of the high-frequency sampled speech segments obtained in step S2. At the same time, security audits of the low-frequency sampled speech segments obtained in step S2 are also required, i.e., steps S5-S6. In other words, steps S3-S4 and steps S5-S6 are performed in parallel, and the speech segments audited are different and do not interfere with each other.
[0077] In step S5, the low-frequency sampled speech segment obtained in step S2 is subjected to speech recognition to generate a second audio recognition text. Specifically, speech recognition can be performed based on a preset recognition model.
[0078] In one embodiment of this disclosure, the preset recognition model can be obtained by training the initial recognition model based on the labeled dataset corresponding to the historical audio data stream. That is, the audio security monitoring device 10 integrates the initial recognition model for processing audio data captured in real time during a call.
[0079] In one example embodiment, the initial recognition model can be OpenAI's Whisper model. This model, with its superior speech recognition capabilities and multilingual support, can efficiently convert audio to text and adapt to the high-speed, complex communication requirements of a 5G environment. When the audio security monitoring device 10 interfaces with the Whisper model, it can use a dedicated Application Programming Interface (API) to call the interface, enabling batch input and parallel processing of audio data to improve processing speed and conversion efficiency.
[0080] To enable the Whisper model to better adapt to the dialects and local language features of different regions, the audio security monitoring device 10 trained the model based on a large amount of local data. This training process included: labeling real call data from various regions and using the labeled data to train the model, ensuring that the model could accurately identify and understand the subtle differences in dialects from different regions.
[0081] The training process utilizes transfer learning technology to transfer the Whisper model's recognition capabilities to local data. Through multiple rounds of iterative training and optimization, the model's performance in handling local accents, language habits, and speech rate variations is significantly improved. This enables the audio security monitoring device 10 to maintain high-precision speech recognition capabilities across various language environments nationwide. The training process can be specifically represented as follows: The local dataset contains multiple audio-text pairs. (i.e., the labeled dataset corresponding to the historical audio data stream), where This indicates an audio signal with a timestamp. This represents the correct text sequence it corresponds to. Through transfer learning, the initial parameters of the Whisper model... After fine-tuning and training on the local dataset, it was updated to The formula is as follows:
[0082]
[0083]
[0084] in, It represents the number of audio-text pairs contained in the native dataset; This represents the parameters after fine-tuning and training, i.e., the optimal parameters; Indicates that it has initial parameters The Whisper model; Indicates that it has initial parameters The Whisper model is a predicted text sequence obtained by performing speech recognition based on the input audio signal; This represents the loss during the fine-tuning training of the Whisper model; Represents the correct text sequence No. One-hot encoded vectors of 1 word, Indicates the prediction of the first text sequence. The probability distribution of each word. The length of the text sequence.
[0085] In other words, by integrating the Whisper model into the audio security monitoring device 10 and fine-tuning it with local data, the audio security monitoring device 10 can achieve high-precision speech recognition and obtain accurate and reliable recognized text. This not only ensures accurate recognition of various dialects and variant languages, but also ensures efficient conversion speed and output quality in the 5G communication environment, laying a solid foundation for text analysis and audio content security review.
[0086] Then, based on the preset recognition model, speech recognition is performed on the second segment of audio data to generate the second audio recognition text. It can be understood that the trained preset recognition model can not only handle Mandarin, but also effectively deal with dialects and variant languages in different regions. Whether it is the soft Wu dialect in the south or the fast and connected speech in the north, the Whisper model can accurately recognize the speech and obtain accurate and reliable recognition text (i.e., the second audio recognition text).
[0087] In step S6, the identified text from step S5 is classified. Specifically, after the speech recognition of the low-frequency sampled speech segment is completed, the audio security monitoring device 10 needs to perform deep semantic analysis on the generated identified text so that subsequent security audits can more accurately identify and classify risk information in the text by combining contextual semantic information. The identification text embedding and classification process mainly includes preliminary parsing of the identified text and classification judgment based on a language model, as follows: S6.1. Based on the labeled dataset corresponding to the sample text, train the initial classification model to generate a preset classification model.
[0088] In one embodiment of this disclosure, the initial classification model can be a BERT (Bidirectional Encoder Representations from Transformers) pre-trained language model, which can understand the contextual relationships of text through bidirectional encoding. Specifically, the audio security monitoring device 10 can pre-collect a large amount of labeled data corresponding to sample texts categorized as fraud, intimidation, obscene information, violence, and compliance and security. Then, based on this labeled data, supervised learning is performed on the initial classification model, and the model parameters are optimized through multiple iterations to enable it to accurately capture the semantic features in the text and perform effective classification judgments.
[0089] Understandably, during the training process, the audio security monitoring device 10 pays special attention to the model's ability to identify various types of risky content in order to ensure the accuracy of the classification results (e.g., being able to distinguish between threatening and violent speech, and being able to identify complex texts that mix multiple categories), and uses the trained model as the preset classification model.
[0090] S6.2. Based on the preset classification model, determine the category and probability value of the second audio recognition text, and take the highest probability value in the abnormal category as the second segment review result.
[0091] Based on the pre-trained classification model in S6.1, the recognized text in step S5 is reviewed to achieve text classification of low-frequency sampled speech segments.
[0092] Specifically, the text to be recognized is first embedded using a pre-defined classification model to convert it into a high-dimensional embedding vector. Then, a classifier is used to classify the text and output the probability distribution of the category of the recognized text corresponding to each low-frequency sampled speech segment. This includes the category to which the recognized text belongs and the probability of each category.
[0093] In one embodiment of this disclosure, the preset classification model pre-defines N+1 categories, including: 1 normal category and N abnormal categories. The classification result may include one or more of these categories, as well as the probability value corresponding to each of these categories. For example, the classification result of a certain low-frequency sampled voice segment is "fraud 90%, compliance 10%". Then, based on the above classification result, the category with the highest probability is determined as its category. If the category is one of the N abnormal categories, such as the fraud category mentioned above, then its corresponding probability (i.e., 90%) is used as the risk score of this low-frequency sampled voice segment, that is, the second segment review result (denoted as...). ).
[0094] This concludes the introduction to the security audit process for the first and second audio data segments based on the first preset rule. The following section describes the specific process for determining the security of the audio data stream based on the second preset rule, the audit results of the first segment, and / or the audit results of the second segment. In other words, to enable real-time assessment of the security of call content and immediate intervention when necessary, the audio security monitoring device 10 is designed with multi-band, multi-level assessment to ensure the accuracy of the audit and the timeliness of the intervention. Specifically: In step S7, it is determined whether the review result of the first segment meets the first threshold.
[0095] In one example embodiment of this disclosure, the first threshold The value is set between 0.7 and 0.9. That is, when the first segment review result... When the cosine similarity is high, the audio security monitoring device 10 can directly mark the passage as unsafe and immediately execute the intervention mechanism of step S10.
[0096] In step S8, it is determined whether the second segment review result meets the second threshold.
[0097] In one example embodiment of this disclosure, the second threshold The value is set between 0.6 and 0.8. That is, when the second stage review results... When the probability of an abnormal classification is high, the audio security monitoring device 10 can directly mark the passage as unsafe and immediately execute the intervention mechanism of step S10.
[0098] In step S9, it is determined whether the comprehensive segment review result meets the third threshold. Specifically, if within a second collection duration, each first segment review result does not meet the first threshold, and the second segment review result does not meet the second threshold, a comprehensive segment review result is calculated based on each first segment review result and the second segment review result. If the comprehensive segment review result meets the third threshold, the audio data stream is determined to be insecure.
[0099] In other words, when neither the high-frequency sampled speech segment nor the low-frequency sampled speech segment meets its respective threshold, i.e. neither triggers the intervention mechanism, the audio security monitoring device 10 will perform a comprehensive score (i.e., comprehensive segment review result) on the call to conduct a comprehensive risk assessment.
[0100] In one example embodiment of this disclosure, the formula for calculating the comprehensive score is as follows:
[0101] It should be noted that since the sampling durations of high-frequency and low-frequency audio segments are different, the shorter high-frequency segment can be used as the standard when calculating the comprehensive score. For example, if there is one low-frequency audio segment and six high-frequency audio segments within three minutes, the risk score for the corresponding low-frequency audio segment can be the same when calculating the comprehensive score for these six high-frequency audio segments.
[0102] After calculating the comprehensive segmented review results Then, compare it with the third threshold. To make a comparison, if If the audio security monitoring device 10 can directly mark the passage as unsafe and immediately execute the intervention mechanism of step S10.
[0103] In one example embodiment of this disclosure, the third threshold It can be set between 0.6 and 0.9.
[0104] Understandably, the specific threshold settings (first threshold, second threshold, and third threshold) can be fine-tuned according to the specific business scenario. For example, the first threshold corresponding to the high-frequency range... In high-risk scenarios (i.e., scenarios with stricter voice call detection), the threshold can be set to 0.7 to improve detection sensitivity; in low-risk scenarios (i.e., scenarios with more lenient voice call detection), the threshold can be increased to 0.85 or 0.9 to reduce false alarms and improve user experience.
[0105] It should be noted that step S9 combines the evaluation results of the high-frequency and low-frequency sampled speech segments, assessing the risk by taking the maximum value. This design ensures that the audio security monitoring device 10 can respond to potential risks in high frequencies in real time, while also not ignoring deeper issues in low frequencies. Combined, the audio security monitoring device 10 can comprehensively and quickly identify and address potential security threats.
[0106] In step S10, if the audio data stream is determined to be insecure, a preset intervention is executed, wherein the preset intervention includes one or more of the following: sending a warning message, blocking the transmission of insecure segmented audio data without interrupting the call, and interrupting the call.
[0107] As described above, if the audio data stream is marked as insecure at any step (S7-S9) during monitoring, this step will be executed immediately. Specifically, the intervention mechanism includes, but is not limited to: ① Warning: The audio security monitoring device 10 generates a warning message, which includes a detailed description of the inappropriate / unsafe content, the category detected, the timestamp of the violation, and the action to be taken. This warning message will be sent to the user through various means, including but not limited to SMS, email, and instant messaging, to ensure that the user receives the notification in a timely manner.
[0108] The specific description of unsafe content can be generated automatically using a large language model, such as the Spark Big Model, based on the specific categories output by the preset classification model or the categories corresponding to the harmful feature vectors matched by the knowledge base.
[0109] ② Block transmission: Directly blocks the transmission of insecure audio frames during the call, ensuring that unauthorized content is not received by the other party. This intervention measure can guarantee the security of the call content without interrupting the call.
[0110] ③ Call interruption: For serious violations, the audio security monitoring device 10 can choose to interrupt the current call and notify the user of the specific reason to protect the user from potentially harmful content.
[0111] ④ Reporting Management: For recurring violations, the audio security monitoring device 10 can report the relevant records and logs to the management department or service provider for further review and processing.
[0112] Furthermore, the intervention mechanism can also include recording and tracking functions. When unsafe content is detected and corresponding intervention measures are taken, the audio security monitoring device 10 records relevant information in the logs table, including the detection time, the category of the violation content, and the intervention measures taken. Through these records, the audio security monitoring device 10 can comprehensively monitor and track the review and intervention process, which helps to continuously optimize the review strategy, thereby improving security and reliability.
[0113] Figure 4 This is a schematic diagram of an audio security monitoring device according to an embodiment of the present disclosure. Figure 4 As shown, the audio security monitoring device 10 may include at least the following modules.
[0114] The acquisition module 101 is used to acquire audio data streams in real time.
[0115] The processing module 102 is used to segment the audio data stream and generate segmented audio data corresponding to multiple acquisition durations.
[0116] The audit module 103 is used to perform real-time security audits on each segment of audio data corresponding to multiple acquisition durations based on preset rules, and to determine whether the audio data stream is secure.
[0117] The processing module 102 may further include: The segmentation processing submodule 1021 is used to segment the audio data stream to generate first segment audio data and second segment audio data, wherein the first acquisition duration of the first segment audio data is less than the second acquisition duration of the second segment audio data.
[0118] The audit module 103 may further include: The first audit submodule 1031 is used to perform security audits on the first segment audio data and the second segment audio data based on the first preset rules, and generate the first segment audit result and the second segment audit result. The second review submodule 1032 is used to determine whether the audio data stream is secure based on the second preset rules, the first segment review result and / or the second segment review result.
[0119] Furthermore, the first audit submodule 1031 may include: Extraction unit 10311 is used to extract features from the first segmented audio data and generate a first audio feature vector; The matching unit 10312 is used to determine the similarity between a preset harmful feature vector and a first audio feature vector, and to use the similarity as the first segmentation review result.
[0120] The recognition unit 10313 is used to perform speech recognition on the second segmented audio data based on a preset recognition model and generate second audio recognition text. The classification unit 10314 is used to determine the category and probability value of the second audio recognition text based on a preset classification model, and to take the highest probability value in the abnormal category as the second segment review result.
[0121] The extraction unit 10311 can be further subdivided into: The framing subunit 103111 is used to perform framing processing on the first segmented audio data to generate multiple frames of audio data. The conversion subunit 103112 is used to perform frequency domain conversion on each of multiple frames of audio data to generate a spectral distribution; The computational subunit 103113 is used to generate a feature vector for each frame of audio data based on the spectral distribution. The splicing subunit 103114 is used to splice each feature vector to generate the first audio feature vector corresponding to the first segment of audio data.
[0122] The recognition unit 10313 can be further subdivided into: Training subunit 103131 is used to train the initial recognition model based on the labeled dataset corresponding to the historical audio data stream, and generate the preset recognition model. The recognition subunit 103132 is used to perform speech recognition on the second segmented audio data based on a preset recognition model, and generate the second audio recognition text.
[0123] Furthermore, the second audit submodule 1032 may include: The separate review unit 10321 is used to determine that the audio data stream is insecure when the review result of the first segment meets the first threshold or the review result of the second segment meets the second threshold. The comprehensive review unit 10322 is used to calculate a comprehensive segment review result based on each first segment review result and the second segment review result when, within a second collection period, each first segment review result does not meet the first threshold and the second segment review result does not meet the second threshold. If the comprehensive segment review result meets the third threshold, the audio data stream is determined to be insecure.
[0124] In addition, the audio security monitoring device 10 may also include: Intervention module 104 is used to perform preset interventions when the audio data stream is determined to be insecure. The preset interventions include one or more of the following: sending a warning message, blocking the transmission of insecure segmented audio data without interrupting the call, and interrupting the call.
[0125] The pre-configuration module 105 is used to determine a preset harmful feature vector based on historical unsafe audio data streams.
[0126] The pre-training module 106 is used to train the initial classification model based on the labeled dataset corresponding to the sample text, and generate a preset classification model.
[0127] The preprocessing module 107 is used to process the first segment audio data and the second segment audio data, wherein the processing includes one or more of the following: noise reduction, enhancement, and format conversion.
[0128] Figure 5 This is a hardware block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device according to an embodiment of the present disclosure includes at least a processor and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor performs the audio security monitoring method as described above.
[0129] Figure 5The illustrated electronic device 500 specifically includes a central processing unit (CPU) 501, a graphics processing unit (GPU) 502, and a memory 503. These units are interconnected via a bus 504. The CPU 501 and / or GPU 502 can function as the aforementioned processor, and the main memory 503 can function as the aforementioned memory storing computer-readable instructions. Furthermore, the electronic device 500 may also include a communication unit 505, a storage unit 506, an output unit 507, an input unit 508, and an external device 509, all of which are also connected to the bus 504.
[0130] Figure 6 This is a schematic diagram illustrating a computer program product according to an embodiment of the present disclosure. Figure 6 As shown, a computer program product 600 according to an embodiment of this disclosure stores a computer program 601. When the computer program 601 is executed by a processor, it performs the audio security monitoring method described with reference to the above figures. The computer program product includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0131] The audio security monitoring method, apparatus, electronic device, and computer program product according to embodiments of the present disclosure have been described above with reference to the accompanying drawings. The audio security monitoring method according to embodiments of the present disclosure achieves a multi-granularity review mechanism by segmenting the audio data stream based on multiple acquisition durations, while simultaneously considering both real-time security monitoring and in-depth analysis, more accurately identifying potential risks in complex contexts. Furthermore, through high-precision speech recognition and understanding of contextual information, it achieves accurate classification of the audio data stream, ensuring the comprehensiveness and accuracy of the review results. At the same time, a real-time automated review and intervention mechanism is introduced, which can immediately trigger automated intervention measures when unsafe content is detected.
[0132] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0133] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0134] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0135] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.
[0136] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0137] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0138] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0139] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. An audio security monitoring method, characterized by, The method comprises: real-time acquisition of an audio data stream; segmenting the audio data stream to generate segmented audio data corresponding to a plurality of collection time lengths; based on a preset rule, real-time security audit of each of the segmented audio data corresponding to the plurality of collection time lengths to determine whether the audio data stream is secure.
2. The audio security monitoring method of claim 1, wherein, The segmented audio data corresponding to a plurality of collection time lengths generated by segmenting the audio data stream comprises: segmenting the audio data stream to generate first segmented audio data and second segmented audio data, wherein the first collection time length of the first segmented audio data is less than the second collection time length of the second segmented audio data.
3. The audio security monitoring method of claim 2, wherein, The security audit of each of the segmented audio data corresponding to the plurality of collection time lengths based on a preset rule to determine whether the audio data stream is secure comprises: based on a first preset rule, security audit of the first segmented audio data and the second segmented audio data to generate first segmented audit results and second segmented audit results; based on a second preset rule, the first segmented audit results and / or the second segmented audit results, determining whether the audio data stream is secure.
4. The audio security monitoring method of claim 3, wherein, The security audit of the first segmented audio data and the second segmented audio data based on a first preset rule to generate first segmented audit results and second segmented audit results comprises: feature extraction of the first segmented audio data to generate a first audio feature vector; based on a preset harmful feature vector, determining the similarity thereof with the first audio feature vector, and taking the similarity as the first segmented audit result.
5. The audio security monitoring method of claim 3, wherein, The security audit of the first segmented audio data and the second segmented audio data based on a first preset rule to generate first segmented audit results and second segmented audit results further comprises: based on a preset recognition model, voice recognition of the second segmented audio data to generate a second audio recognition text; based on a preset classification model, determining the category and probability value of the second audio recognition text, and taking the maximum probability value in the abnormal category as the second segmented audit result.
6. The audio security monitoring method of claim 4, wherein, The feature extraction of the first segmented audio data to generate a first audio feature vector comprises: frame processing of the first segmented audio data to generate a plurality of frame audio data; frequency domain conversion of each of the plurality of frame audio data to generate a frequency spectrum distribution; based on the frequency spectrum distribution, generating a feature vector of each of the frame audio data; splicing of each of the feature vectors to generate the first audio feature vector corresponding to the first segmented audio data.
7. The audio security monitoring method of claim 5, wherein, The voice recognition of the second segmented audio data based on a preset recognition model to generate a second audio recognition text comprises: training of an initial recognition model based on a labeled data set corresponding to historical audio data streams to generate the preset recognition model; based on the preset recognition model, voice recognition of the second segmented audio data to generate the second audio recognition text.
8. The audio security monitoring method of claim 4, wherein, The determining, based on the second preset rule, the first segment review result, and / or the second segment review result, whether the audio data stream is safe includes: When the first segment review result meets a first threshold or the second segment review result meets a second threshold, determining that the audio data stream is unsafe; When, in one of the second collection time lengths, each of the first segment review results does not meet the first threshold and the second segment review result does not meet the second threshold, calculating a comprehensive segment review result based on each of the first segment review results and the second segment review result, and if the comprehensive segment review result meets a third threshold, determining that the audio data stream is unsafe.
9. The audio security monitoring method of claim 1, wherein, The method further includes: In a case where the audio data stream is determined to be unsafe, performing a preset intervention, wherein the preset intervention includes one or more of the following: sending a warning message, preventing transmission of unsafe segment audio data without interrupting the call, and interrupting the call.
10. The audio security monitoring method of claim 4, wherein, The method further includes: Determining the preset harmful feature vector based on historical unsafe audio data streams.
11. The audio security monitoring method of claim 5, wherein, The method further includes: Training an initial classification model based on a labeled data set corresponding to a sample text to generate the preset classification model.
12. The audio security monitoring method of claim 2, wherein, The method further includes: Processing the first segment audio data and the second segment audio data, wherein the processing includes one or more of the following: denoising, enhancing, and format conversion.
13. An audio security monitoring device, characterized by The device includes: An acquisition module configured to acquire an audio data stream in real time; A processing module configured to segment the audio data stream to generate segment audio data corresponding to a plurality of collection time lengths; An audit module configured to determine whether the audio data stream is safe based on a preset rule by performing real-time safety auditing on each of the segment audio data corresponding to the plurality of collection time lengths.
14. The audio security monitoring device of claim 13, wherein, The processing module includes: A segment processing submodule configured to segment the audio data stream to generate first segment audio data and second segment audio data, wherein a first collection time length of the first segment audio data is less than a second collection time length of the second segment audio data.
15. The audio security monitoring device of claim 14, wherein, The audit module includes: A first audit submodule configured to perform safety auditing on the first segment audio data and the second segment audio data based on a first preset rule to generate a first segment review result and a second segment review result; A second audit submodule configured to determine whether the audio data stream is safe based on a second preset rule, the first segment review result, and / or the second segment review result.
16. The audio security monitoring device of claim 15, wherein, The first audit submodule includes: An extraction unit configured to perform feature extraction on the first segment audio data to generate a first audio feature vector; A matching unit configured to determine a similarity between a preset harmful feature vector and the first audio feature vector based on the preset harmful feature vector, and use the similarity as the first segment review result.
17. The audio security monitoring device of claim 15, wherein, The first audit submodule further includes: An identification unit configured to perform speech recognition on the second segment audio data based on a preset identification model to generate a second audio recognition text; The classification unit is configured to determine a category and a probability value of the second audio recognition text based on a preset classification model, and take a maximum probability value in an abnormal category as the second segment review result.
18. The audio security monitoring device of claim 16, wherein, The extraction unit includes: The frame dividing subunit is configured to perform frame dividing processing on the first segmented audio data to generate a plurality of frame audio data; The conversion subunit is configured to perform frequency domain conversion on each of the plurality of frame audio data to generate a spectrum distribution; The calculation subunit is configured to generate a feature vector of each of the frame audio data based on the spectrum distribution; The splicing subunit is configured to splice each of the feature vectors to generate the first audio feature vector corresponding to the first segmented audio data.
19. The audio security monitoring device of claim 17, wherein, The recognition unit includes: The training subunit is configured to train an initial recognition model based on a labeled data set corresponding to a historical audio data stream to generate the preset recognition model; The recognition subunit is configured to perform speech recognition on the second segmented audio data based on the preset recognition model to generate the second audio recognition text.
20. The audio security monitoring device of claim 16, wherein, The second review sub-module includes: The individual review unit is configured to determine that the audio data stream is unsafe when the first segment review result meets a first threshold or the second segment review result meets a second threshold; The comprehensive review unit is configured to calculate a comprehensive segment review result based on each of the first segment review result and the second segment review result when each of the first segment review result in a second collection time length does not meet the first threshold and the second segment review result does not meet the second threshold, and determine that the audio data stream is unsafe if the comprehensive segment review result meets a third threshold.
21. The audio security monitoring device of claim 13, wherein, The device further includes: The intervention module is configured to perform a preset intervention when the audio data stream is determined to be unsafe, The preset intervention includes one or more of the following: sending a warning message, preventing transmission of unsafe segmented audio data without interrupting the call, and interrupting the call.
22. The audio security monitoring device of claim 16, wherein, The device further includes: The pre-configuration module is configured to determine the preset harmful feature vector based on a historical unsafe audio data stream.
23. The audio surveillance device of claim 17, wherein, The device further includes: The pre-training module is configured to train an initial classification model based on a labeled data set corresponding to a sample text to generate the preset classification model.
24. The audio security monitoring device of claim 14, wherein, The device further includes: The pre-processing module is configured to process the first segmented audio data and the second segmented audio data, wherein the processing includes one or more of the following: denoising, enhancing, and format conversion.
25. An electronic device, comprising: The device further includes: The memory is configured to store computer readable instructions; and The processor is configured to run the computer readable instructions to enable the electronic device to perform the audio safety monitoring method of any one of claims 1 to 12. The computer program is executed by the processor to implement the audio safety monitoring method of any one of claims 1 to 12.
26. A computer program product comprising a computer program, characterised in that, The computer program is executed by the processor to implement the audio safety monitoring method of any one of claims 1 to 12.