Method, gateway and medium for managing multi-channel content based on audio restoration

CN122621406BActive Publication Date: 2026-09-22HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611038977.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-22
Estimated Expiration
2046-07-13

AI Technical Summary

Benefits of technology

本申请实施例提供的对于音频数据的内容管控方法,创造性的在广播场景中设计流量旁路部署,在不干扰原有广播链路的基础上,采集用于网关进行异常检测的音频数据。同时,考虑到音频数据本身包含多种时间尺度的信息,本申请提供的内容管控方法还设计了一种包含快速通道和慢速通道的异常检测模型,对不同采样率、不同特征内容进行特征提取和处理,关注毫秒级短期变化和分钟级的长期变化,形成覆盖多时间尺度的特征提取网络,从而扩展特征深度,进而提高整个异常检测模型的异常检测的准确性,进一步提升整个广播系统的可靠性。进一步,在广播场景下单“音频源”多“播放源”架构下,音源设备到每个功放设备的音频数据可能相同,通过对冗余会话筛选后写入能够降低网关的数据交互压力,同时面对网络抖动、数据包丢失产生的音频数据恢复不完整问题,通过完整性位图实现对不完整的内容缺失会话进行补全,进一步提高音频数据传播过程的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621406B_ABST
    Figure CN122621406B_ABST
Patent Text Reader

Abstract

The application provides a multi-channel content management and control method based on audio restoration, a gateway and a medium, and relates to the technical field of network security. The method comprises the following steps: receiving mirror traffic data, wherein the mirror traffic data is generated by mirroring the original traffic data broadcast by a broadcast integrated machine based on a switch, and is transmitted to the gateway through a bypass deployment; writing the mirror traffic data after removing duplicate content in the mirror traffic data based on a content hash table; restoring the session in the mirror traffic data to generate audio data to be detected, including completing the restoration of the content missing session based on an integrity bitmap; inputting the audio data to be detected into an anomaly detection model to output an anomaly detection result, wherein the anomaly detection model extracts time domain features through a fast channel and extracts frequency domain features through a slow channel to generate target fusion features for anomaly content identification, and then outputs the anomaly detection result; and triggering a corresponding content management and control operation. The accuracy of anomaly content detection in the audio transmission process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security, and in particular to a multi-channel content control method, gateway, and medium based on audio restoration. Background Technology

[0002] Content control refers to the management and control of various forms of information content (including text, images, audio, video, etc.) to maintain network order and ensure network security. Especially in broadcasting scenarios, a single broadcast control unit typically manages dozens or even hundreds of amplifiers, transmitting specific audio data via network signal propagation. Current technologies primarily perform content detection at the "audio source" of the broadcast control unit, halting transmission upon detecting anomalies. However, attackers can still tamper with and forge the transmission link through man-in-the-middle attacks, session hijacking, and other methods; the security of the audio transmission process cannot be guaranteed.

[0003] Meanwhile, since audio content often includes information at multiple time scales, typically ranging from slow background changes to rapid text descriptions, existing content control methods for audio broadcasts cannot capture the unique rhythmic information of audio, making it difficult to accurately identify whether the audio content contains abnormal information, resulting in low accuracy in identifying abnormal information.

[0004] Therefore, there is an urgent need for a method to improve the accuracy of abnormal content identification during audio transmission in order to solve the above-mentioned technical problems. Summary of the Invention

[0005] The purpose of this application is to provide a multi-channel content control method, gateway, and medium based on audio restoration, so as to improve the accuracy of abnormal audio content identification and thus enhance the security of the entire broadcast system during audio transmission, further improving the overall security of the broadcast system. The specific technical solution is as follows: In a first aspect, embodiments of this application provide a multi-channel content management method based on audio restoration, applied to a gateway within a broadcast system, the method comprising: Receive mirrored traffic data, which is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one machine, and transmitted to the gateway through bypass deployment. The mirrored traffic data includes multiple sessions. After filtering out duplicate content in the mirror traffic data based on the content hash table, it is written into the mirror traffic data. The content hash table is used to maintain the content hash value corresponding to the mirror traffic data already written in the gateway. Restore sessions in mirrored traffic data to generate audio data to be detected, including restoring sessions with missing content in mirrored traffic data based on an integrity bitmap, which is used to maintain the content integrity of each session; Input the audio data to be detected into the anomaly detection model to output the anomaly detection result. The anomaly detection model extracts the temporal features from the audio data to be detected through the fast channel and extracts the frequency features from the audio data to be detected through the slow channel. Based on the temporal and frequency features at a preset scale, feature fusion is performed to generate the target fused feature. Based on the target fused feature, anomaly content is identified to output the anomaly detection result. Based on the anomaly detection results, corresponding content control operations are triggered.

[0006] Optionally, based on temporal and frequency domain features at a preset scale, feature fusion is performed to generate target fusion features, including: Interpolation and channel projection operations are performed on the frequency domain features to generate the first intermediate feature, wherein the first intermediate feature is aligned with the time domain feature in both the time and channel dimensions. Based on the first intermediate feature and the temporal feature, a first fusion feature is generated; Perform the first downsampling operation on the time-domain features and the frequency-domain features respectively to obtain the 1 / 2 scale time-domain features and the 1 / 2 scale frequency-domain features; Interpolation alignment and channel projection operations are performed on the 1 / 2 scale frequency domain features to generate a second intermediate feature, wherein the second intermediate feature is aligned with the 1 / 2 scale time domain features in both the time and channel dimensions. A second fusion feature is generated based on the second intermediate feature and the 1 / 2 scale time domain feature; Perform an interpolation alignment operation on the second fused feature to obtain a third fused feature that is aligned with the first fused feature in the time dimension; The target fusion feature is generated based on the third fusion feature and the first fusion feature.

[0007] Optionally, after generating the second fused feature based on the second intermediate feature and the 1 / 2 scale temporal feature, the method further includes updating the second fused feature: A second downsampling operation is performed on the time-domain features and the frequency-domain features respectively to obtain 1 / 4-scale time-domain features and 1 / 4-scale frequency-domain features; Interpolation alignment and channel projection operations are performed on the 1 / 4-scale frequency domain features to generate a third intermediate feature, wherein the third intermediate feature is aligned with the 1 / 4-scale time domain features in both the time and channel dimensions. A fourth fusion feature is generated based on the third intermediate feature and the 1 / 4 scale time-domain feature; Perform an interpolation alignment operation on the fourth fusion feature to obtain a fifth fusion feature that is aligned with the second fusion feature in the time dimension; The elements in each channel of the second and fifth fusion features are added together to generate a new second fusion feature.

[0008] Optionally, a first fused feature is generated based on the first intermediate feature and the temporal feature, including: The first intermediate feature and the temporal feature are concatenated to generate a feature sequence; The input feature sequence is fed into the bidirectional cross-attention mechanism module to obtain enhanced frequency domain features and enhanced temporal features. The enhanced temporal features include a query vector whose elements are derived from the temporal features, a key vector whose elements are derived from the first intermediate feature, and a value vector whose elements are derived from the temporal features. The enhanced frequency domain features include a query vector whose elements are derived from the first intermediate feature, a key vector whose elements are derived from the temporal features, and a value vector whose elements are derived from the temporal features. The fused enhanced frequency domain features and enhanced time domain features are input into the feedforward network to generate the first fused feature.

[0009] Optionally, the session includes at least one data packet, and after filtering duplicate content in the mirrored traffic data based on a content hash table, the data written to the mirrored traffic data includes: In response to receiving mirrored traffic data containing at least one data packet, the payload of the data packet is used as a data block describing the data packet, and a content hash value matching the data packet is generated based on the data block; Check whether the content hash value is included in the content hash table, which records the content hash values ​​of multiple data blocks contained in the mirrored traffic data that has been written to the main buffer; In response to the detection that the content hash value is not included in the content hash table, the data block is written to the specified position of the session in the main buffer based on the position marker in the packet; In response to the detection that the content hash value is already included in the content hash table, the write count of the data block in the redundancy backup table is updated, where the redundancy backup table is used to record the write count of each data block.

[0010] Optionally, the audio data to be detected includes multiple audio files. The system reconstructs missing sessions from the mirrored traffic data based on an integrity bitmap, including: The integrity bitmap is used to identify the target sessions to be supplemented and the missing data blocks that match the sessions with missing content. The integrity bitmap records the status of data blocks in different positions of each session. Traverse the integrity bitmap to obtain candidate sessions containing missing data blocks; If the number of candidate sessions is one, then the candidate session is determined as the target candidate session; If the number of candidate sessions is not one, the target candidate session is determined based on the number of missing data blocks and the activity level of the candidate session. Based on the candidate data blocks in the target candidate sessions that match the missing data blocks, supplement the target sessions to be supplemented; Restore the target session to be supplemented to generate the audio file to be detected.

[0011] Optionally, the gateway also maintains a session replenishment time window, which is used to constrain the replenishment time for sessions with missing content. After traversing the integrity bitmap to obtain candidate sessions containing missing data blocks, the method also includes: In response to the absence of a candidate session in the session supplement window, the missing session is forcibly restored to generate an audio file to be detected based on the data blocks currently contained in the missing session.

[0012] Optionally, before inputting the audio data to be detected into the anomaly detection model to output the anomaly detection result, the method further includes: Retrieve the historical file hash values ​​of the detected audio files; Record the current file hash values ​​of multiple audio files to be detected contained in the audio data to be detected; If a target audio file exists whose current file hash value matches a historical file hash value, stop inputting the target audio file into the anomaly detection model. Find the anomaly detection results of the detected audio files that match the target audio file to be detected, and output them as the anomaly detection results of the target audio file to be detected.

[0013] Secondly, embodiments of this application provide a gateway, including: The receiving module receives mirrored traffic data, which is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one device, and is transmitted to the gateway through a bypass deployment. The mirrored traffic data includes multiple sessions. The restore module is used to filter duplicate content in the mirror traffic data based on the content hash table and then write it into the mirror traffic data. The content hash table is used to maintain the content hash value corresponding to the mirror traffic data already written in the gateway. The restoration module is also used to restore the sessions in the mirrored traffic data to generate audio data to be detected, including restoring the sessions with missing content contained in the mirrored traffic data based on the integrity bitmap, which is used to maintain the content integrity of each session. The detection module inputs the audio data to be detected into the anomaly detection model and outputs the anomaly detection result. The anomaly detection model extracts the temporal features from the audio data to be detected through the fast channel and extracts the frequency features from the audio data to be detected through the slow channel. Based on the temporal and frequency features at a preset scale, it performs feature fusion to generate target fusion features and performs anomaly content recognition based on the target fusion features to output the anomaly detection result. The feedback module is used to trigger corresponding content control operations based on the anomaly detection results.

[0014] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multi-channel content management method based on audio restoration disclosed in the first aspect.

[0015] Beneficial effects of the embodiments in this application: The content management method for audio data provided in this application creatively designs a traffic bypass deployment in broadcast scenarios, collecting audio data for gateway anomaly detection without interfering with the original broadcast link. Simultaneously, considering that audio data itself contains information at multiple time scales, the content management method provided in this application also designs an anomaly detection model including fast and slow channels. It extracts and processes features for different sampling rates and different characteristic content, focusing on short-term changes at the millisecond level and long-term changes at the minute level, forming a feature extraction network covering multiple time scales. This expands the feature depth, thereby improving the accuracy of anomaly detection in the entire anomaly detection model and further enhancing the reliability of the entire broadcast system. Furthermore, in a single "audio source" multi-"playback source" architecture in a broadcast scenario, the audio data from the audio source device to each power amplifier device may be the same. By filtering redundant sessions before writing, the data interaction pressure on the gateway can be reduced. At the same time, facing the problem of incomplete audio data recovery caused by network jitter and packet loss, an integrity bitmap is used to complete incomplete content missing sessions, further improving the stability of the audio data propagation process.

[0016] Of course, implementing any product or method of this application does not necessarily require achieving all of the above advantages at the same time. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0018] Figure 1 A schematic diagram of the structure of a broadcast system provided in an embodiment of this application; Figure 2 A flowchart illustrating a multi-channel content management method based on audio restoration, provided for an embodiment of this application; Figure 3 A schematic diagram illustrating the workflow of an anomaly detection model provided in an embodiment of this application; Figure 4 This is a schematic diagram of a feature fusion method provided in an embodiment of this application; Figure 5This is a schematic diagram of another feature fusion method provided in the embodiments of this application; Figure 6 This is a schematic diagram of a feature fusion method after feature enhancement provided in an embodiment of this application; Figure 7 A schematic diagram of another feature fusion method after feature enhancement provided in this application embodiment; Figure 8 This is a schematic diagram of the structure of a gateway provided in an embodiment of this application; Figure 9 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0020] As disclosed in the background technology, existing content control methods for audio data are still at the stage of recognition models based on single frame rate sampling, which cannot capture the rhythmic information unique to audio data. The model analysis has a single dimension, resulting in unsatisfactory final detection results. At the same time, it cannot solve the problem of incomplete audio data recovery caused by network jitter and packet loss. Based on this, in order to improve the accuracy of anomaly detection in audio data, this application provides a multi-channel content control method, gateway, and medium based on audio restoration.

[0021] The following section introduces the multi-channel content management method based on audio restoration provided in this application.

[0022] The multi-channel content control method based on audio restoration provided in this application is applied to a gateway within a broadcast system. This gateway plays the role of content security detection. By monitoring whether the audio data to be broadcast contains abnormal content and promptly controlling and processing abnormal audio, the security of the entire broadcast system is improved. For example, Figure 1As shown, a broadcast system typically includes audio source devices, switches, broadcast all-in-one units, gateways, and power amplifiers. Of course, in other scenarios, broadcast systems may also include other adaptively added functional devices. This application only describes a typical system architecture for illustration and has no restrictive effect. The gateway can be a standalone server, an embedded industrial computer, or a network hardware device integrating specific functions; this application does not limit the specific hardware type of the gateway. The content control method disclosed in this application adds content security detection scenarios without interfering with the original broadcast link, requiring no modification or adaptation of the broadcast all-in-one unit, power amplifier, or other devices, and has a wide range of applications.

[0023] This application provides a multi-channel content management method based on audio restoration, such as... Figure 2 As shown, it includes the following steps: S100: Receive mirrored traffic data, which is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one machine, and transmitted to the gateway through bypass deployment.

[0024] Specifically, the broadcast all-in-one unit communicates with at least one audio source device and one power amplifier device via a network. The audio source device is any device that generates raw traffic data, such as a microphone, DVD player, network speaker, or media host. The raw traffic data includes analog audio source data and network audio source data. The audio content (i.e., raw traffic data) that the audio source device needs to play must be sent to each power amplifier device via the broadcast all-in-one unit. It should be noted that in some scenarios, the broadcast all-in-one unit itself can also act as an audio source device, generating and broadcasting raw traffic data that needs to be sent to the power amplifier devices for playback. This application deploys a traffic mirroring function on the switch that forwards the raw traffic data to be broadcast by the broadcast all-in-one unit. While normally receiving and sending the raw traffic data generated by the audio source device, the broadcast all-in-one unit can also act as an audio source device to send traffic to the playback device. The switch performs mirroring processing on the raw traffic data to generate mirrored traffic data. In this bypass deployment method, the switch configures the mirrored traffic data of each service port to the mirror port through mirrored traffic configuration. The gateway is directly connected to the mirror port of the switch via a network cable to receive the traffic data. After the gateway opens its network port in promiscuous mode, it can monitor mirrored traffic data. It is understood that, in this embodiment, the aforementioned mirrored traffic data is transmitted in the form of network traffic, and the mirrored traffic data includes multiple sessions.

[0025] S200: After removing duplicate content from the mirror traffic data based on the content hash table, write the mirror traffic data.

[0026] In a specific implementation scenario, step S200 above includes the following: Step S210: In response to receiving mirrored traffic data containing at least one data packet, the payload of the data packet is used as a data block describing the data packet, and a content hash value matching the data packet is generated based on the data block. The aforementioned payload is the audio data segment actually contained in the data packet, and the content hash value is used to distinguish the content in other sessions.

[0027] Step S220: Detect whether the content hash value is included in the content hash table maintained by the gateway. The content hash table is used to maintain the content hash values ​​corresponding to the mirrored traffic data already written to the gateway. Specifically, the content hash table records and stores the content hash values ​​corresponding to multiple data blocks contained in the mirrored traffic data already written to the main buffer. Step S230: In response to the detection that the content hash value is not included in the content hash table, the data block is written to the specified position of the session already written in the main buffer based on the position marker in the data packet. The fact that the content hash value is not included in the content hash table indicates that the content of this data packet is being written to the gateway for the first time. At this time, based on the position marker recorded in the data packet, the data block is written to the specified position of the session in the main buffer. The position marker includes a TCP session sequence number or a custom timestamp, used to mark the logical position of the data block in the session stream; that is, the data block is stored in the main buffer at the logical position corresponding to the position marker.

[0028] Step S240: In response to detecting that the content hash value is already included in the hash table, update the write count of the data block in the redundancy backup table, where the redundancy backup table is used to record the write count of each data block. The fact that the content hash value is already included in the content hash table indicates that the same content as that data block has already been written to the main cache within the gateway. In this case, only the write count of the data block in the redundancy backup table is updated. It is understood that, to more clearly describe the source of each data block, the redundancy backup table can also maintain the corresponding session identifier.

[0029] In this embodiment, when multiple sessions transmit the same content, only the first valid data is written to the main buffer. Subsequent sessions only update the "redundant backup table" for the same data blocks, without repeatedly writing to the disk, which greatly reduces the read and write interaction pressure inside the gateway.

[0030] S300, restore the sessions in the mirrored traffic data to generate audio data to be detected.

[0031] In specific implementation scenarios, the aforementioned mirrored traffic data is traffic data transmitted based on the TCP protocol. It can be understood that this mirrored traffic data includes data from at least one network session, and each session includes at least one data packet. The gateway reassembles and restores the mirrored traffic data based on the TCP session sequence by capturing multiple data packets from the mirrored traffic data. Then, it obtains the audio data to be detected from the restored session, mainly including: static file restoration and real-time session stream restoration. Static files refer to audio files present in the session. Supported audio file formats include MP3 (Moving Picture Experts Group Audio Layer III) and WAV (Waveform Audio File Format), etc.; and the following video file formats containing audio are also supported: AVI (Audio Video Interleaved), FLV (FLASHVIDEO), MP4 (MPEG-4 Part 14), MPG (Moving Picture Experts Group), MOV (QuickTime Movie), etc. The complete audio file is only available after the session ends, and sessions are relatively short. Real-time session streams refer to streams where each data packet contains a segment of audio data. The application layer transmits these packets via protocols such as HTTP, ISAPI, and SIP. The audio data in each data packet can be analyzed, and sessions are longer. After the above processing, the audio data to be detected can be obtained. In specific implementation scenarios, the audio data to be detected can be set to WAV format, which contains specific audio files. Each specific audio file has the same duration, such as 5s, 10s, etc. This application does not require the duration of the audio files.

[0032] Understandably, due to network jitter, the mirrored traffic data contains both complete and incomplete sessions. As mentioned earlier, the high buffer stores the written sessions, and the data blocks in each session are sorted according to their corresponding position tags. For complete sessions, the entire session content can be reconstructed directly from all the data blocks recorded in that session. That is, according to the encoding format matching the transmission protocol, the payload corresponding to the data blocks is decoded, and the data is restored from the traffic format to the audio data format.

[0033] In specific implementation scenarios, step S300 above also includes restoring the missing content sessions contained in the mirrored traffic data based on the integrity bitmap: In response to the detection of sessions with unrecoverable missing content, a target session to be supplemented and its missing data blocks are determined based on an integrity bitmap. A session with missing content is one that has lost a portion of its data blocks. The integrity bitmap maintains the content integrity of each session, specifically recording the data block status of each session at different locations. The integrity bitmap uses location markers as the horizontal axis (same as previously disclosed, which can be TCP session sequence numbers, custom timestamps, file offsets, etc.) and session identifiers as the vertical axis. Specifically, data block status can be marked with colors; for example, green indicates that the data block at that location is available, and red indicates that the data block at that location is lost for all sessions. By traversing the integrity bitmap, the session with the fewest missing data blocks among multiple sessions transmitting the same content is identified as the target session to be supplemented.

[0034] It is understandable that when the gateway responds to mirrored traffic data, it may detect multiple sessions with missing content, and correspondingly, identify multiple target sessions to be supplemented. For each target session to be supplemented, this application proposes the following: traversing the integrity bitmap to obtain candidate sessions containing missing data blocks; if the number of candidate sessions is one, then the candidate session is determined as the target candidate session; if the number of candidate sessions is not one, the target candidate session is determined based on the number of missing data blocks and the activity of the candidate sessions; supplementing the target session to be supplemented based on the candidate data blocks in the target candidate session that match the missing data blocks; restoring the supplemented target session to be supplemented to generate the audio file to be detected.

[0035] The process of determining the target candidate session based on the number of missing data blocks and the status indicators of candidate sessions includes the following steps: If a first candidate session containing all missing data blocks is detected, and the number of first candidate sessions is one, then this first candidate session is directly determined as the target candidate session; if the number of first candidate sessions is not one, then the first candidate session with the highest activity level is selected as the target candidate session based on its activity ranking. If no first candidate session containing all missing data blocks is found, then the candidate session containing the largest number of corresponding missing data blocks is determined as the second candidate session. Similarly, if the number of second candidate sessions is one, then this second candidate session is directly determined as the target candidate session; if the number of second candidate sessions is not one, then the second candidate session with the highest activity level is selected as the target candidate session based on its activity ranking.

[0036] In specific implementation scenarios, the gateway also maintains a sliding session replenishment time window, such as 5 seconds or 10 seconds, during which it waits for the best data block that can be used for splicing. After traversing the integrity bitmap to obtain candidate sessions containing missing data blocks, this application also proposes: In response to the absence of a candidate session within the session supplementation window, the missing session is forcibly restored to generate the audio file to be detected, based on the data blocks currently contained in the missing session. That is, if no candidate session is detected within the time window, the session content is forcibly restored using the data blocks currently contained in the target session to be supplemented, ensuring real-time performance; if a missing data block suitable for splicing a target candidate session is found within the time window, the missing data block and the target session to be spliced ​​are immediately pushed to the buffer, restored, and then output to the anomaly detection model. This application allows data blocks to arrive in the buffer out of order.

[0037] Understandably, traditional solutions treat each network session as an independent island. If session A loses 10% of its packets and session B loses 15%, and the packet loss locations are different, traditional solutions will generate two incomplete audio files, failing to utilize the complementarity between sessions. For example, a session needs to transmit data blocks [1 2 3 4 5 6 7 8 9 10]. Session A might restore data blocks [1 2 3 7 8 9 10], while session B might restore data blocks [1 2 4 5 6 7 8 9 10]. Using the method disclosed above in the embodiments of this application, all data blocks from session A and data blocks 4, 5, and 6 from session B can be combined to completely recover the content that the session needs to transmit.

[0038] S400: Input the audio data to be detected into the anomaly detection model to output the anomaly detection results.

[0039] The anomaly detection model extracts temporal features from the audio data to be detected through a fast channel and frequency domain features through a slow channel. Based on these temporal and frequency features at a preset scale, feature fusion is performed to generate a target fused feature. Anomaly detection results are then output based on this target fused feature. In other words, features are extracted from the audio data to be detected from multiple channels for anomaly identification and subsequent content control. This anomaly detection model is pre-trained and deployed in a gateway. In the anomaly detection model disclosed in this application, a uniform sampling rate is not used. Instead, the fast channel maintains a high sampling rate, focusing on "fast" changes (transient); the slow channel maintains a low sampling rate, focusing on "slow" changes (steady-state), which is more consistent with the physical characteristics of audio.

[0040] In this application, the aforementioned anomaly detection model extracts temporal features from the audio data to be detected through a fast channel. Specifically, this involves sampling the audio data to be detected input to the model at a high sampling rate (e.g., 44.1kHz / 48kHz) and performing feature processing through a vector convolutional network to obtain temporal features. In a specific implementation scenario, the aforementioned vector convolutional network specifically includes two vector convolutional layers. First, a convolutional process is performed with 1 channel, 64 kernels, a kernel size of 5, and a stride of 2; after normalization, ReLU (Rectified Linear Unit) processing is applied. Then, a second convolutional process is performed with 64 channels, 128 kernels, a kernel size of 3, and a stride of 2; after normalization, ReLU processing is applied to extract temporal features. Of course, to improve model processing efficiency, this application can set a short time window, such as 10ms, to extract features from segments of audio files input within the short time window to obtain temporal features.

[0041] Understandably, time-domain features analyze the amplitude of an audio signal directly over time, reflecting the signal's intuitive characteristics on the time axis. Time-domain features are fast to calculate, suitable for rapid response and precise time positioning, and can extract contextual information. The core objective of extracting time-domain features is to capture transient changes, energy spikes, and waveform distortions in audio data. These time-domain features include, but are not limited to, amplitude features and short-time zero-crossing rate (i.e., the number of times the curve crosses the zero axis). Amplitude features include short-time energy, mean amplitude, amplitude variance, and peak amplitude.

[0042] In this application, the aforementioned anomaly detection model extracts frequency domain features from the audio data to be detected through a slow channel. Specifically, this involves downsampling the audio data to be detected input to the model to a lower sampling rate (e.g., 16kHz) and then extracting frequency domain features using a tensor convolutional network. Tensor convolutional networks are an important variant of convolutional neural networks; they employ various tensor decomposition algorithms to decompose the massive convolutional kernels into low-rank forms for storage, thus significantly compressing the network and making it more efficient. In a specific implementation scenario, the aforementioned tensor convolutional network contains two tensor convolutional layers and one pooling layer. First, a convolutional process is performed with 1 channel, 32 kernels, a kernel size of 3, and a stride of 1. After normalization, ReLU processing is performed, followed by pooling. Then, another convolutional process is performed with 32 channels, 64 kernels, a kernel size of 3, and a stride of 1. Finally, after normalization, ReLU processing is performed to extract frequency domain features. Similarly, to improve model processing efficiency, a long window such as 100ms can be set, and the model extracts frequency domain features from the input audio data to be detected at intervals of the long window.

[0043] It is understandable that frequency domain features are obtained by converting audio signals from the time domain to the frequency domain using Fourier transform, and are used to analyze the distribution characteristics of the signal at different frequency components. The core objective of extracting frequency domain features is to capture timbre structure, harmonic distribution, and spectral texture. Frequency domain features include, but are not limited to, spectral shape features, spectral distribution features, harmonic features, and cepstral features. Among them, spectral shape features include the spectral centroid, spectral roll-off point, and spectral bandwidth; harmonic features include the fundamental frequency and harmonics; and cepstral features include MFCC features (Mel-Frequency Cepstral Coefficients) and LFCC features (Linear Frequency Cepstral Coefficients). Frequency domain features contain rich semantic information and are high-level features suitable for deep model learning, providing deep information such as timbre and content, thereby ensuring the reliability of the detection results of anomaly detection models.

[0044] like Figure 3 As shown, the anomaly detection model disclosed in this application extracts corresponding time-domain and frequency-domain features from the input audio data to be detected through two channels, then performs feature fusion and inputs it into the corresponding task head to finally output the anomaly detection result. It is understandable that, to adapt to classification and target recognition tasks in specific scenarios, different task heads can be set in the above anomaly detection model for feature processing. For classification tasks, a classification head is set, which performs global pooling after obtaining the target fusion features, followed by a fully connected layer and a softmax layer (normalization exponential layer). The output anomaly detection result includes the category to which the audio data to be detected belongs, such as normal, obscene, drug-related, gambling-related, and reactionary, etc. For target recognition tasks, a detection head is set, which connects to a fully connected layer after obtaining the target fusion features, and outputs the target labels included in the anomaly detection result, such as howling, electromagnetic interference, stuttering, etc. Of course, in specific implementation scenarios, the above anomaly detection model can simultaneously set two task heads, that is, simultaneously set a classification head and a detection head to perform classification and recognition tasks simultaneously. In implementation scenarios where only recognition tasks are required, the above anomaly detection model can be configured to output anomaly detection results containing the target label only; in implementation scenarios where only classification tasks are required, the above anomaly detection model can be configured to output anomaly detection results containing the category to which the audio data to be detected belongs only.

[0045] In specific implementation scenarios, the above-mentioned features are based on temporal and frequency domain features at a preset scale and are fused to generate target fused features, such as... Figure 4 As shown, it specifically includes: Interpolation and channel projection operations are performed on the frequency domain features to generate a first intermediate feature, wherein the first intermediate feature is aligned with the time domain feature in both the time and channel dimensions. Based on the first intermediate feature and the time domain feature, a first fused feature is generated. A first downsampling operation is performed on the time domain feature and the frequency domain feature respectively to obtain a 1 / 2-scale time domain feature and a 1 / 2-scale frequency domain feature. Interpolation and channel projection operations are performed on the 1 / 2-scale frequency domain feature to generate a second intermediate feature, wherein the second intermediate feature is aligned with the 1 / 2-scale time domain feature in both the time and channel dimensions. Based on the second intermediate feature and the 1 / 2-scale time domain feature, a second fused feature is generated. Interpolation and alignment operations are performed on the second fused feature to obtain a third fused feature aligned with the first fused feature in the time dimension. Based on the third fused feature and the first fused feature, a target fused feature is generated.

[0046] In specific implementation scenarios, the aforementioned first fused feature can be generated by adding the channel elements of the first intermediate feature and the temporal feature, or by weighted summation using a learnable weight. The interpolation alignment operation can be implemented using nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation. The interpolation alignment operation upsamples the high-dimensional frequency domain features in the slow channel, ensuring that the frequency domain features extracted by the fast channel in the low frame rate are consistent with the temporal features extracted in the high frame rate in the time dimension. The channel projection operation expands the number of channels in the frequency domain features. For example, in the previously disclosed implementation scenario, the frequency domain features have 32 channels, which can be expanded to 64 channels through channel projection. Specifically, a 1*1 convolutional kernel can be used to project 32 channels to 64 channels, thus ensuring that the number of channels in the frequency domain features and the temporal features are the same.

[0047] The aforementioned downsampling operation can be understood as extraction. In this application, it refers to sampling the current feature sequence at intervals to generate a new feature sequence. The first downsampling operation is to sample half of the original scale features at a certain interval to form a new feature sequence. The second downsampling operation disclosed below is to sample one-quarter of the original scale features at a certain interval to form a new feature sequence. The 1 / 2 scale and the 1 / 4 scale mentioned below refer to the scales used for time-division. For example, if the original scale of the frequency domain feature is 100ms, the 1 / 2 scale means the frequency domain feature extracted within 50ms, and the 1 / 4 scale means the frequency domain feature extracted within 25ms. The 1 / 2 scale time domain feature is the time domain feature within 1 / 2 of the original scale in the time dimension, and the 1 / 2 scale frequency domain feature is the frequency domain feature within 1 / 2 of the original scale in the time dimension. The 1 / 4 scale time domain feature disclosed below is the time domain feature within 1 / 4 of the original scale in the time dimension, and the 1 / 4 scale frequency domain feature is the frequency domain feature within 1 / 4 of the original scale in the time dimension.

[0048] In specific implementation scenarios, the fusion method for the aforementioned target fusion features can be achieved by adding the elements within the third fusion feature and the first fusion feature; it can also be achieved by concatenating the first fusion feature and the third fusion feature and then fusing them through a linear layer; or it can be achieved by weighted summation using a learnable weight; this application does not limit this method. By fusing frequency domain features and time domain features at different scales, this application ensures that the model can recognize impact sounds as short as a few milliseconds as well as understand musical styles or speech intonations lasting several seconds, further enhancing the accuracy of the entire model in detection.

[0049] S500: Based on the anomaly detection results, trigger the corresponding content control operation.

[0050] After the gateway determines that there is an anomaly in the audio data in the traffic based on the anomaly detection results, it supports multiple methods for content control processing, including but not limited to the following methods, which are for illustrative purposes only: In terms of alarm notification methods, the gateway supports sending alarms to users of the broadcast all-in-one machine via SMS, WeChat, alarm devices, and other linkage methods to prompt users to interrupt the broadcast all-in-one machine from broadcasting traffic.

[0051] In the active blocking method, the gateway interrupts the link between the broadcast unit and the power amplifier via TCP-Reset (a mechanism in the TCP protocol used to forcibly and immediately terminate a connection), or directly uses ARP (Address Resolution Protocol) spoofing to tamper with the IP (Internet Protocol Address) - MAC (Media Access Control Address) mapping relationship of the broadcast unit in the network, causing data packets to be unable to be sent and received normally, thus blocking the broadcast of abnormal audio data. ARP spoofing is an attack technique targeting the Ethernet Address Resolution Protocol, which deceives the gateway MAC address of visitors within the local area network, causing network outages.

[0052] The linkage blocking method involves the gateway blocking the broadcast of abnormal data by linking with a broadcast all-in-one machine or switch.

[0053] The multi-channel content management method based on audio restoration provided in this application creatively designs a traffic bypass deployment in broadcast scenarios. Without interfering with the original broadcast link, it collects audio data for gateway anomaly detection. Furthermore, considering that audio data itself contains information at multiple time scales, the content management method also designs an anomaly detection model including fast and slow channels. It extracts and processes features for different sampling rates and different characteristic content, focusing on short-term changes at the millisecond level and long-term changes at the minute level, forming a feature extraction network covering multiple time scales. This expands the feature depth and improves the accuracy of anomaly detection in the entire anomaly detection model, further enhancing the reliability of the entire broadcast system. Moreover, in a single "audio source" multi-"playback source" architecture in a broadcast scenario, the audio data from the audio source device to each power amplifier device may be the same. By filtering redundant sessions before writing, the data interaction pressure on the gateway can be reduced. Simultaneously, to address the problem of incomplete audio data recovery caused by network jitter and packet loss, an integrity bitmap is used to complete incomplete content missing sessions, further improving the stability of the audio data propagation process.

[0054] Optionally, in another embodiment of this application, in order to improve the semantic richness of the second fused feature, this application also proposes to update the second fused feature after performing the step of generating the second fused feature based on the second intermediate feature and the 1 / 2 scale temporal domain feature, such as... Figure 5 As shown, it specifically includes: A second downsampling operation is performed on both the time-domain and frequency-domain features to obtain quarter-scale time-domain and frequency-domain features, respectively. The definitions of the quarter-scale time-domain and frequency-domain features have been described previously and will not be repeated here. Interpolation alignment and channel projection operations are performed on the quarter-scale frequency-domain features to generate a third intermediate feature, which is aligned with the quarter-scale time-domain feature in both the time and channel dimensions. Based on the third intermediate feature and the quarter-scale time-domain feature, a fourth fused feature is generated. Interpolation alignment is performed on the fourth fused feature to obtain a fifth fused feature aligned with the second fused feature in the time dimension. The elements within each channel of the second and fifth fused features are summed to generate a new second fused feature.

[0055] Optionally, in another embodiment of this application, in order to improve the accuracy and efficiency of model detection, based on the two embodiments disclosed above, such as... Figure 6 and Figure 7 The diagram shown illustrates the feature fusion method after feature enhancement. The first fused feature, generated based on the first intermediate feature and the temporal feature, can also be achieved through the following steps: The first intermediate feature and the time-domain feature are spliced ​​together to generate a feature sequence. Taking the parameters of the aforementioned specific anomaly detection model as an example, the features of 64 channels each from the fast channel and slow channel branches are spliced ​​together to obtain a high-dimensional feature sequence of 128 channels. The first 64 channels of the high-dimensional feature sequence are frequency domain features, and the last 64 channels are time domain features.

[0056] Input the feature sequence into the bidirectional cross-attention mechanism module to obtain enhanced frequency domain features and enhanced time domain features.

[0057] The enhancement of temporal features includes query vectors (elements derived from temporal features), key vectors (elements derived from the first intermediate feature), and value vectors (elements derived from the first intermediate feature). This enhances the understanding of details in the original frequency domain features, such as enhancing the "heavy hit node" in a "drum sound clip." In other words, it leverages frequency domain features to enhance temporal features. For each desired temporal feature, it determines which information should be obtained from the frequency domain features; therefore, it uses query vectors derived from temporal features. Since the information originates from the frequency domain features, it uses key vectors and value vectors derived from the first intermediate feature.

[0058] The enhanced frequency domain features include query vectors with elements derived from the first intermediate feature, key vectors with elements derived from the temporal feature, and value vectors to correct the boundaries of the original temporal feature description, such as "drum sound segment" starting from "this moment". That is, frequency domain features are enhanced using temporal features. For each feature in the desired frequency domain features, the information to be obtained from the temporal features is determined. Since the information comes from the frequency domain features, query vectors with elements derived from the first intermediate frequency domain feature are used. Since the information comes from the temporal features, key vectors and value vectors with elements derived from the temporal features are used. By employing cross-channel attention, the problem of traditional stitching merely "piling up" information is solved. The frequency domain semantics of the slow channel can act like a "mask," telling the dense temporal features acquired by the fast channel where the important transients are, and vice versa.

[0059] The enhanced frequency domain features and enhanced time domain features after fusion are input into the feedforward network to generate the first fused feature. Compared with the sequence features input into the cross-attention mechanism module, the first fused feature has its dimensionality reduced to half of the original number of channels after passing through the feedforward network. This eliminates the need for extremely high-dimensional deep computation at all levels, achieving a balance between accuracy and efficiency.

[0060] In this embodiment, the present application obtains a first fused feature by performing serial adaptive feature projection and cross-channel attention fusion on features at the original scale; at the same time, the features are processed at multiple scales and fused with the first fused feature to obtain the final target fused feature, which is used for subsequent classification and target detection to improve the accuracy of the entire model in identifying abnormal content.

[0061] Optionally, in another embodiment of this application, in order to save the computing power of the anomaly detection model, this application further proposes the following before inputting the audio data to be detected into the anomaly detection model to output the anomaly detection result: Obtain the historical file hash values ​​of the detected audio files; record the current file hash values ​​of multiple audio files to be detected contained in the audio data to be detected, that is, the gateway records the file hash value of the file while restoring the audio file. If there is a target audio file to be detected whose current file hash value matches the historical file hash value, stop inputting the target audio file to be detected into the anomaly detection model; find the anomaly detection results of the detected audio files that match the target audio file to be detected, and output them as the anomaly detection results of the target audio file to be detected.

[0062] That is, based on the historical file hash value of the already output audio files, it is determined whether the currently collected audio file to be detected still needs to be output to the model for detection. If it has already been output, the model will not be detected again, and the detection result of the already detected audio file corresponding to the historical file hash value will be directly used as the detection result of the audio file. This greatly saves the computing power of the gateway and speeds up the efficiency of the entire content control process.

[0063] Based on the above method embodiments, this application also provides a gateway 800, such as... Figure 8 As shown, it includes: The receiving module 810 receives mirrored traffic data, which is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one machine, and is transmitted to the gateway through a bypass deployment. The mirrored traffic data includes multiple sessions. The restore module 820 is used to filter duplicate content in the mirror traffic data based on the content hash table and then write it into the mirror traffic data. The content hash table is used to maintain the content hash value corresponding to the mirror traffic data already written in the gateway. The restoration module 820 is also used to restore the sessions in the mirrored traffic data to generate audio data to be detected, including restoring the sessions with missing content contained in the mirrored traffic data based on the integrity bitmap, which is used to maintain the content integrity of each session. The detection module 830 inputs the audio data to be detected to the anomaly detection model and outputs the anomaly detection result. The anomaly detection model extracts the temporal features in the audio data to be detected through the fast channel and extracts the frequency features in the audio data to be detected through the slow channel. Based on the temporal and frequency features at a preset scale, it performs feature fusion to generate target fusion features and performs anomaly content recognition based on the target fusion features to output the anomaly detection result. Feedback module 840 is used to trigger corresponding content control operations based on anomaly detection results.

[0064] Optionally, the detection module 830 is further configured to: perform interpolation alignment and channel projection operations on the frequency domain features to generate a first intermediate feature, wherein the first intermediate feature is aligned with the time domain feature in both the time and channel dimensions; generate a first fused feature based on the first intermediate feature and the time domain feature; perform a first downsampling operation on the time domain feature and the frequency domain feature respectively to obtain a 1 / 2-scale time domain feature and a 1 / 2-scale frequency domain feature; perform interpolation alignment and channel projection operations on the 1 / 2-scale frequency domain feature to generate a second intermediate feature, wherein the second intermediate feature is aligned with the 1 / 2-scale time domain feature in both the time and channel dimensions; generate a second fused feature based on the second intermediate feature and the 1 / 2-scale time domain feature; perform an interpolation alignment operation on the second fused feature to obtain a third fused feature aligned with the first fused feature in the time dimension; and generate a target fused feature based on the third fused feature and the first fused feature.

[0065] Optionally, the detection module 830 is further configured to: perform a second downsampling operation on the time-domain features and the frequency-domain features respectively to obtain 1 / 4-scale time-domain features and 1 / 4-scale frequency-domain features; perform interpolation alignment operation and channel projection operation on the 1 / 4-scale frequency-domain features to generate a third intermediate feature, wherein the third intermediate feature is aligned with the 1 / 4-scale time-domain features in both the time and channel dimensions; generate a fourth fusion feature based on the third intermediate feature and the 1 / 4-scale time-domain features; perform an interpolation alignment operation on the fourth fusion feature to obtain a fifth fusion feature aligned with the second fusion feature in the time dimension; and add the elements in each channel of the second fusion feature and the fifth fusion feature to generate a new second fusion feature.

[0066] Optionally, the detection module 830 is further configured to: concatenate the first intermediate feature and the temporal feature to generate a feature sequence; input the feature sequence into the bidirectional cross-attention mechanism module to obtain enhanced frequency domain features and enhanced temporal features, wherein the enhanced temporal features include a query vector whose elements are derived from the temporal feature, a key vector whose elements are derived from the first intermediate feature, and a value vector whose elements are derived from the first intermediate feature; the enhanced frequency domain features include a query vector whose elements are derived from the first intermediate feature, a key vector whose elements are derived from the temporal feature, and a value vector; and input the fused enhanced frequency domain features and enhanced temporal features into the feedforward network to generate the first fused feature.

[0067] Optionally, the restoration module 820 is further configured to: in response to receiving mirrored traffic data containing at least one data packet, use the payload of the data packet as a data block describing the data packet, and generate a content hash value matching the data packet based on the data block; detect whether the content hash value is included in a content hash table, the content hash table recording the content hash values ​​of multiple data blocks contained in the mirrored traffic data that have been written to the main buffer; in response to detecting that the content hash value is not included in the content hash table, write the data block to a specified position of the session in the main buffer based on the position marker in the data packet; and in response to detecting that the content hash value is included in the content hash table, update the write count of the data block in the redundancy backup table, wherein the redundancy backup table is used to record the write count of each data block.

[0068] Optionally, the restoration module 820 is further configured to: determine the target session to be supplemented and the missing data blocks that match the session with missing content based on the integrity bitmap, wherein the integrity bitmap records the data block status of each session at different positions; traverse the integrity bitmap to obtain candidate sessions containing missing data blocks; if the number of candidate sessions is one, then determine the candidate session as the target candidate session; if the number of candidate sessions is not one, determine the target candidate session based on the number of missing data blocks and the activity of the candidate sessions; supplement the target session to be supplemented based on the candidate data blocks in the target candidate session that match the missing data blocks; and restore the supplemented target session to be supplemented to generate an audio file to be detected.

[0069] Optionally, the gateway also maintains a session supplementation time window, which is used to constrain the supplementation time of the content-missing session; after traversing the integrity bitmap to obtain candidate sessions containing the missing data blocks, the above-mentioned restoration module is further configured to: in response to the absence of the candidate session in the session supplementation window, force the restoration of the content-missing session to generate the audio file to be detected based on the data blocks currently contained in the content-missing session.

[0070] Optionally, the restoration module 820 is further configured to: obtain the historical file hash values ​​of the detected audio files; record the current file hash values ​​of multiple audio files to be detected contained in the audio data to be detected; if there is a target audio file to be detected whose current file hash value matches the historical file hash value, stop inputting the target audio file to be detected into the anomaly detection model; find the anomaly detection result of the detected audio file that matches the target audio file to be detected, and output it as the anomaly detection result of the target audio file to be detected.

[0071] In this application embodiment, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the multi-channel content management method based on audio restoration disclosed in any of the above embodiments.

[0072] In another embodiment provided in this application, an electronic device is also provided, such as... Figure 9 As shown, it includes: Memory 901 is used to store computer programs; The processor 902, when executing the program stored in the memory 901, implements the multi-channel content management method based on audio restoration disclosed in any embodiment.

[0073] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 902, communication interface, and memory 901 communicating with each other via the communication bus.

[0074] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0075] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0076] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0077] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0078] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state disk (SSD), etc.

[0079] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0080] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0081] The above are merely preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A multi-channel content management method based on audio restoration, applied to a gateway within a broadcast system, the broadcast system further including a switch and a broadcast unit, characterized in that, The method includes: Receive mirrored traffic data, wherein the mirrored traffic data is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one machine, and is transmitted to the gateway through a bypass deployment method. The mirrored traffic data includes multiple sessions. After removing duplicate content from the mirror traffic data based on the content hash table, the content hash table is written into the mirror traffic data. The content hash table is used to maintain the content hash value corresponding to the mirror traffic data already written in the gateway. Restore the sessions in the mirrored traffic data to generate audio data to be detected, including restoring the sessions with missing content contained in the mirrored traffic data based on an integrity bitmap, wherein the integrity bitmap is used to maintain the content integrity of each session; The audio data to be detected is input into the anomaly detection model to output anomaly detection results. The anomaly detection model extracts the temporal features from the audio data to be detected through a fast channel and extracts the frequency features from the audio data to be detected through a slow channel. Based on the temporal features and the frequency features at a preset scale, feature fusion is performed to generate target fusion features. Based on the target fusion features, anomaly content is identified to output anomaly detection results. Based on the anomaly detection results, the corresponding content control operation is triggered.

2. The method according to claim 1, characterized in that, The process of fusing the time-domain features and frequency-domain features at a preset scale to generate target fused features includes: The frequency domain features are interpolated and aligned, and channel projection is performed to generate a first intermediate feature, wherein the first intermediate feature is aligned with the time domain features in both the time and channel dimensions. Based on the first intermediate feature and the temporal feature, a first fusion feature is generated; A first downsampling operation is performed on the time-domain features and the frequency-domain features respectively to obtain 1 / 2 scale time-domain features and 1 / 2 scale frequency-domain features; Interpolation and channel projection operations are performed on the 1 / 2 scale frequency domain features to generate a second intermediate feature, wherein the second intermediate feature is aligned with the 1 / 2 scale time domain features in both the time and channel dimensions. Based on the second intermediate feature and the 1 / 2 scale time-domain feature, a second fusion feature is generated; Perform an interpolation alignment operation on the second fusion feature to obtain a third fusion feature that is aligned with the first fusion feature in the time dimension; Based on the third fusion feature and the first fusion feature, a target fusion feature is generated.

3. The method according to claim 2, characterized in that, After generating the second fused feature based on the second intermediate feature and the 1 / 2 scale time-domain feature, the method further includes updating the second fused feature: A second downsampling operation is performed on the time-domain features and the frequency-domain features respectively to obtain 1 / 4-scale time-domain features and 1 / 4-scale frequency-domain features; Interpolation and channel projection operations are performed on the 1 / 4-scale frequency domain features to generate a third intermediate feature, wherein the third intermediate feature is aligned with the 1 / 4-scale time domain features in both the time and channel dimensions. Based on the third intermediate feature and the 1 / 4 scale time-domain feature, a fourth fusion feature is generated; Perform an interpolation alignment operation on the fourth fusion feature to obtain a fifth fusion feature that is aligned with the second fusion feature in the time dimension; The elements in each channel of the second fusion feature and the fifth fusion feature are added together to generate a new second fusion feature.

4. The method according to claim 2 or 3, characterized in that, The step of generating a first fused feature based on the first intermediate feature and the temporal feature includes: The first intermediate feature and the temporal feature are concatenated to generate a feature sequence; The feature sequence is input into the bidirectional cross-attention mechanism module to obtain enhanced frequency domain features and enhanced time domain features. The enhanced time domain features include a query vector whose elements are derived from the time domain features, a key vector whose elements are derived from the first intermediate feature, and a value vector whose elements are derived from the time domain features. The enhanced frequency domain features include a query vector whose elements are derived from the first intermediate feature, a key vector whose elements are derived from the time domain features, and a value vector whose elements are derived from the time domain features. The fused enhanced frequency domain features and enhanced time domain features are input into the feedforward network to generate the first fused feature.

5. The method according to claim 1, characterized in that, The session includes at least one data packet, and the step of removing duplicate content from the mirrored traffic data based on a content hash table and then writing it into the mirrored traffic data includes: In response to receiving the mirrored traffic data containing at least one of the data packets, the payload of the data packets is used as a data block describing the data packets, and a content hash value matching the data packets is generated based on the data blocks; Detect whether the content hash value is included in the content hash table, which records the content hash values ​​of multiple data blocks contained in the mirror traffic data that has been written to the main buffer; In response to the detection that the content hash value is not included in the content hash table, the data block is written to the specified position of the session in the main buffer based on the position marker in the data packet; In response to detecting that the content hash value is already included in the content hash table, the write count of the data block in the redundancy backup table is updated, wherein the redundancy backup table is used to record the write count of each data block.

6. The method according to claim 5, characterized in that, The audio data to be detected includes multiple audio files to be detected. The process of restoring the missing content sessions contained in the mirrored traffic data based on the integrity bitmap includes: The target session to be supplemented and the missing data block are determined based on the integrity bitmap that matches the session with missing content, wherein the integrity bitmap records the data block status of each session at different locations; Traverse the integrity bitmap to obtain candidate sessions containing the missing data blocks; If the number of candidate sessions is one, then the candidate session is determined to be the target candidate session; If the number of candidate sessions is not one, the target candidate session is determined based on the number of missing data blocks and the activity level of the candidate sessions; Based on the candidate data blocks in the target candidate sessions that match the missing data blocks, supplement the target sessions to be supplemented; The target session to be supplemented is restored to generate an audio file to be detected.

7. The method according to claim 6, characterized in that, The gateway also maintains a session replenishment time window, which is used to constrain the replenishment time for sessions with missing content. After traversing the integrity bitmap to obtain candidate sessions containing the missing data blocks, the method further includes: In response to the absence of the candidate session in the session supplement window, the missing session is forcibly restored to generate the audio file to be detected based on the data blocks currently contained in the missing session.

8. The method according to claim 6, characterized in that, Before inputting the audio data to be detected into the anomaly detection model to output the anomaly detection result, the method further includes: Retrieve the historical file hash values ​​of the detected audio files; Record the current file hash values ​​of the multiple audio files to be detected contained in the audio data to be detected; If a target audio file exists whose current file hash value matches a historical file hash value, stop inputting the target audio file into the anomaly detection model; Find the anomaly detection results of the detected audio files that match the target audio file to be detected, and output them as the anomaly detection results of the target audio file to be detected.

9. A gateway, characterized in that, The gateway includes: The receiving module receives mirrored traffic data, wherein the mirrored traffic data is generated by the switch after mirroring the original traffic data broadcast by the broadcast all-in-one device, and is transmitted to the gateway through a bypass deployment. The mirrored traffic data includes multiple sessions. The restore module is used to remove duplicate content from the mirror traffic data based on the content hash table and then write it back into the mirror traffic data, wherein the content hash table is used to maintain the content hash value corresponding to the mirror traffic data already written in the gateway; The restoration module is also used to restore the sessions in the mirrored traffic data to generate audio data to be detected, including restoring the sessions with missing content contained in the mirrored traffic data based on the integrity bitmap, wherein the integrity bitmap is used to maintain the content integrity of each session; The detection module inputs the audio data to be detected into the anomaly detection model to output anomaly detection results. The anomaly detection model extracts the temporal features from the audio data to be detected through a fast channel and extracts the frequency features from the audio data to be detected through a slow channel. Based on the temporal features and the frequency features at a preset scale, it performs feature fusion to generate target fusion features and performs anomaly content recognition based on the target fusion features to output anomaly detection results. The feedback module is used to trigger corresponding content control operations based on the anomaly detection results.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Audio tampering detection method and device, server and storage medium

    CN113808603A

  • Video and audio signal monitoring method of switch node port mirror image

    CN121907668A