Audio audit method, electronic equipment, storage medium and program product

By using multimodal audio feature fusion and dynamic trust assessment algorithms, the problem of insufficient accuracy in anomaly detection and real-time protection capabilities in bastion host audio auditing is solved. This achieves highly accurate intelligent auditing of operation and maintenance audio, improves the accuracy of anomaly detection and real-time protection capabilities of the system, and protects privacy and resources.

CN121789718APending Publication Date: 2026-04-03BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing bastion hosts suffer from low accuracy in anomaly detection and insufficient real-time protection capabilities in monitoring and auditing audio channels, failing to effectively utilize the rich security information contained in audio data.

Method used

By extracting multimodal audio features and combining them with a dynamic trust assessment algorithm to process audio feature data, multimodal data fusion is achieved. Differential privacy algorithm is used to protect privacy, semi-homomorphic encryption technology is used to encrypt audio audit results, and a differentiated audio audit strategy is implemented based on trust scores.

Benefits of technology

It improves the accuracy of anomaly detection and real-time protection capabilities, optimizes resource allocation, enhances the system's adaptability and privacy protection capabilities, and meets the requirements of regulations such as GDPR.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789718A_ABST
    Figure CN121789718A_ABST
Patent Text Reader

Abstract

The invention provides an audio auditing method, electronic equipment, a storage medium and a program product, and relates to the technical field of data security. The audio auditing method comprises the following steps: acquiring initial audio data to be audited; extracting multi-mode audio features according to the initial audio data to obtain audio feature data; processing the audio feature data according to a dynamic trust evaluation algorithm to obtain trust integral data; obtaining an audio audit strategy according to a trust integral-audit strategy library and the trust integral data; and auditing the audio feature data according to the audio auditing strategy to obtain an audio auditing result. According to the audio auditing method, the technical effect of improving the anomaly detection accuracy and the real-time protection capability can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and more specifically, to an audio auditing method, electronic device, storage medium, and program product. Background Technology

[0002] Currently, bastion hosts, as core security management and auditing tools for operations and maintenance, have evolved from basic command logs to screen recording auditing, and then to comprehensive operational monitoring. However, in terms of audio channel monitoring and auditing, technological development has lagged behind, with significant technical bottlenecks.

[0003] In existing technologies, traditional bastion hosts primarily focus on monitoring graphical interface operations and command-line input, achieving operational auditing through screen recording and keystroke logging. However, with the increasing complexity of remote operational scenarios and rising security requirements, the role of audio communication in the operational process is becoming increasingly prominent: including voice-guided collaboration, emergency response dialogues, and monitoring of equipment operating sounds in industrial environments. This audio data contains rich security information, yet it has long been overlooked by traditional auditing systems; and commonly used audio auditing methods suffer from low accuracy in anomaly detection and insufficient real-time protection capabilities. Summary of the Invention

[0004] The purpose of this application is to provide an audio auditing method, electronic device, storage medium, and program product that can achieve the technical effect of improving the accuracy of anomaly detection and real-time protection capabilities.

[0005] Firstly, this application provides an audio auditing method, including: Obtain the initial audio data to be audited; Multimodal audio features are extracted from the initial audio data to obtain audio feature data; The audio feature data is processed according to the dynamic trust assessment algorithm to obtain trust score data; An audio audit strategy is obtained based on the trust score-audit strategy library and the trust score data; The audio feature data is audited according to the audio audit strategy to obtain audio audit results.

[0006] In the above implementation process, multimodal audio features are extracted from the initial audio data to obtain audio feature data, thus giving the audio feature data multidimensional features. This audio feature data is then processed by a dynamic trust assessment algorithm to achieve multimodal data fusion and obtain a trust score for the real-time risk level of the current audio. This not only improves the accuracy of anomaly detection and context awareness but also establishes a dynamic trust assessment mechanism, enabling differentiated audio auditing strategies based on real-time trust levels. Therefore, this audio auditing method can achieve the technical effect of improving the accuracy of anomaly detection and real-time protection capabilities.

[0007] Further, the step of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data includes: Based on the audio feature data, obtain one or more of the following data: identity credibility factor, behavioral compliance factor, environmental risk factor, and session context factor; The trust score data is obtained by processing one or more of the data from the identity credibility factor, the behavior compliance factor, the environmental risk factor, and the session context factor according to the adaptive weight model. The adaptive weight model applies weighted penalties to each factor based on preset risk events.

[0008] In the above implementation process, multiple dimensions of factors are continuously acquired based on audio feature data. Through the fusion of data from multiple dimensions such as identity credibility factor, behavioral compliance factor, environmental risk factor, and conversation context factor, a trust score representing the real-time risk level of the current audio is calculated. A dynamic trust assessment algorithm is implemented through an adaptive weight model, which enables adaptive adjustment of the weights of each factor. Furthermore, each factor can be weighted and penalized according to preset risk events to improve the accuracy of anomaly detection.

[0009] Further, the step of obtaining an audio auditing policy based on the trust score-audit policy library and the trust score data includes: Based on the trust score data, the trust level data is obtained; Based on the trust level data, the trust score-audit strategy library is matched to obtain an audio audit strategy, which includes a segment capture mode, a sampling audit mode, and a full audit mode. If the audio audit strategy matches the full audit mode, then preset GPU resources are allocated to analyze and audit the initial audio data.

[0010] In the above implementation process, risk is classified based on trust score data. After the risk classification, the corresponding audio audit strategy is selected. High trust only needs to execute the segment capture mode, that is, only when a suspected abnormal event is detected, the audio before and after a few seconds is recorded and analyzed. Medium trust needs to execute the sampling audit mode, which randomly records audio segments at a certain proportion and performs standard analysis. Low trust needs to execute the full audit mode, which records all audio completely and starts the deep analysis mode of all modalities. In addition, when executing the full audit mode, preset GPU resources are allocated to analyze the initial audio data of the audit to ensure effective use of resources.

[0011] Furthermore, before the step of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data, the method further includes: The audio feature data is processed using a differential privacy algorithm and preset noise data to obtain privacy-preserving audio feature data. The steps of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data include: The privacy audio feature data is processed according to a dynamic trust assessment algorithm to obtain trust score data.

[0012] In the above implementation process, a differential privacy algorithm is used to add precisely calculated statistical noise to the feature vector after audio feature extraction and before the dynamic trust assessment algorithm processes the audio feature data. This ensures that the aggregated trust assessment result is basically accurate while protecting the privacy of individual speakers.

[0013] Further, after the step of auditing the audio feature data according to the audio auditing strategy to obtain the audio audit results, the method further includes: The audio audit results are encrypted using a semi-homomorphic encryption algorithm to obtain encrypted audio audit results. The encrypted audio audit results are subject to trust-adaptive access control based on the trust level of the accessing user, and can be retrieved and processed before decryption based on preset audit tags.

[0014] In the above implementation process, a semi-homomorphic encryption algorithm is used to encrypt the audio audit results, allowing the system to retrieve and calculate certain audit tags in the audio audit results in ciphertext without decryption; at the same time, combined with trust-adaptive access control, administrators with low trust levels cannot decrypt and access highly sensitive raw audio data, but can only view the de-identified analysis report.

[0015] Further, the step of auditing the audio feature data according to the audio auditing strategy to obtain audio audit results includes: Perform cross-modal correlation analysis and early warning on the audio feature data to obtain the audio audit results; After the step of auditing the audio feature data according to the audio audit strategy and obtaining the audio audit result, the method further includes: If the audio audit results include alarm information, a preset response strategy is executed based on the alarm information.

[0016] In the above implementation process, based on cross-modal correlation analysis of audio feature data, multi-dimensional feature correlation analysis and cross-validation of audio feature data can be realized, which can effectively eliminate false alarms of single-modal analysis; if there is an alarm message in the audio audit results, it indicates that there is an anomaly in the audio feature data, and a preset response strategy is executed according to the alarm message.

[0017] Furthermore, after executing a preset response strategy based on the alarm information, the method further includes: Obtain the response and processing results of the alarm information; The alarm feedback category of the alarm information is marked according to the response processing result. The alarm feedback category includes correct alarm, false alarm, and missed alarm.

[0018] In the above implementation process, alarm events and their corresponding processing results are recorded, and alarm information is marked as correct alarm, false alarm, or missed alarm. Thus, by analyzing the alarm information of false alarm or missed alarm, the audio auditing method can be dynamically adjusted to improve the accuracy of audio auditing.

[0019] Furthermore, after the step of labeling the alarm feedback category of the alarm information based on the processing result, the method further includes: Based on the alarm feedback category, obtain model adjustment data; The dynamic trust assessment algorithm or the audio audit strategy is adjusted based on the model adjustment data.

[0020] In the above implementation process, the dynamic trust assessment algorithm or audio audit strategy is adjusted according to the model adjustment data, so that the dynamic trust assessment and multimodal fusion model can be continuously optimized to adapt to new operation and maintenance scenarios and attack methods. For example, for false alarm cases, the detection threshold or trust factor weight of the relevant modality is fine-tuned; for missed alarms or newly discovered attack patterns, the relevant voiceprint, audio or semantic patterns are added to the feature library of the audio audit strategy for updating the model.

[0021] Further, the step of extracting multimodal audio features from the initial audio data to obtain audio feature data includes: Extract one or more of the following from the initial audio data: voiceprint biometric features, semantic features, machine health status features, and emotional stress features. Audio feature data is obtained based on one or more of the following features: voiceprint biometrics, semantic features, machine health status features, and emotional stress features.

[0022] In the above implementation process, features from multiple dimensions such as voiceprint biometrics, semantic features, machine health status features, and emotional stress features are extracted in real time, providing rich and structured data for subsequent analysis.

[0023] Secondly, this application provides an audio auditing system, including: The audio acquisition module is used to acquire the initial audio data to be audited. The audio feature module is used to extract multimodal audio features based on the initial audio data to obtain audio feature data; The trust score module is used to process the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data. The audit strategy module is used to obtain audio audit strategies based on the trust score-audit strategy library and the trust score data; The audit results module is used to audit the audio feature data according to the audio audit strategy and obtain audio audit results.

[0024] Furthermore, the trust score module is also used to: obtain one or more data from identity credibility factor, behavior compliance factor, environmental risk factor, and session context factor based on the audio feature data; process one or more data from identity credibility factor, behavior compliance factor, environmental risk factor, and session context factor according to an adaptive weight model to obtain the trust score data, wherein the adaptive weight model applies weighted penalties to each factor based on preset risk events.

[0025] Furthermore, the audit strategy module is also used to: obtain trust level data based on the trust score data; match the trust score-audit strategy library based on the trust level data to obtain an audio audit strategy, wherein the audio audit strategy includes a segment capture mode, a sampling audit mode, and a full audit mode; if the audio audit strategy matches the full audit mode, then allocate preset GPU resources to analyze and audit the initial audio data.

[0026] Furthermore, the audio auditing system also includes a privacy processing module, used to: process the audio feature data based on a differential privacy algorithm and preset noise data to obtain privacy audio feature data; The trust score module is also used to: process the privacy audio feature data according to the dynamic trust assessment algorithm to obtain trust score data.

[0027] Furthermore, the audio auditing system also includes an encryption module, used to: encrypt the audio audit results based on a semi-homomorphic encryption algorithm to obtain encrypted audio audit results, wherein the encrypted audio audit results perform trust-adaptive access control according to the trust level of the accessing user, and can be retrieved and processed based on preset audit tags when not decrypted.

[0028] Furthermore, the audit result module is also used to: perform cross-modal correlation analysis and early warning on the audio feature data, and obtain the audio audit result; The audio auditing system also includes an alarm module, which is used to: if the audio auditing result includes alarm information, execute a preset response strategy based on the alarm information.

[0029] Furthermore, the alarm module is also used to: obtain the response processing result of the alarm information; and mark the alarm feedback category of the alarm information according to the response processing result, wherein the alarm feedback category includes correct alarm, false alarm, and missed alarm.

[0030] Furthermore, the audio auditing system also includes a model adjustment module, used to: obtain model adjustment data according to the alarm feedback category; and adjust the dynamic trust assessment algorithm or the audio auditing strategy according to the model adjustment data.

[0031] Furthermore, the audio feature module is also used to: extract one or more of voiceprint biometric features, semantic features, machine health status features, and emotional stress features based on the initial audio data; and obtain audio feature data based on one or more of the voiceprint biometric features, the semantic features, the machine health status features, and the emotional stress features.

[0032] Thirdly, this application provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any of the first aspects.

[0033] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in any of the first aspects.

[0034] Fifthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the method described in any of the first aspects.

[0035] Other features and advantages disclosed in this application will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the above-described technology disclosed in this application.

[0036] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0037] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating an audio auditing method provided in an embodiment of this application; Figure 2 A flowchart illustrating another audio auditing method provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the process of privacy-processing audio feature data and encrypted audio audit results provided in an embodiment of this application; Figure 4 A structural block diagram of the audio auditing system provided in the embodiments of this application; Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0039] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0040] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0041] Generally speaking, bastion hosts, as core security management and auditing tools for operations and maintenance, have evolved from basic command logs to screen recording auditing, and then to comprehensive operational monitoring. However, in terms of audio channel monitoring and auditing, technological development has lagged behind, with significant technical bottlenecks.

[0042] In existing technologies, traditional bastion hosts primarily focus on monitoring graphical interface operations and command-line input, achieving operational auditing through screen recording and keystroke logging. However, with the increasing complexity of remote operational scenarios and rising security requirements, the role of audio communication in the operational process is becoming increasingly prominent: including voice-guided collaboration, emergency response dialogues, and monitoring of equipment operating sounds in industrial environments. This audio data contains rich security information, yet it has long been overlooked by traditional auditing systems; and commonly used audio auditing methods suffer from low accuracy in anomaly detection and insufficient real-time protection capabilities.

[0043] To address the aforementioned technical problems, embodiments of this application provide an audio auditing method, electronic device, storage medium, and program product. This audio auditing method extracts multimodal audio features from initial audio data, resulting in audio feature data with multiple dimensions. This data is then processed using a dynamic trust assessment algorithm. Through multimodal data fusion, the accuracy of anomaly detection and context awareness are improved. Furthermore, a dynamic trust assessment mechanism is established to implement differentiated audio auditing strategies based on real-time trust levels. Therefore, this audio auditing method can achieve the technical effects of improving anomaly detection accuracy and real-time protection capabilities.

[0044] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a bastion host audio auditing system and method based on dynamic trust assessment and multimodal fusion, achieving adaptive, fine-grained, and highly accurate intelligent auditing of operation and maintenance audio. Specific objectives include: (1) Establish a dynamic trust assessment mechanism to realize a differentiated audio audit strategy based on real-time trust levels; (2) Improve the accuracy and context awareness of anomaly detection through multimodal data fusion; (3) Introduce lightweight privacy protection technology to reduce the exposure of privacy data while ensuring the audit effect; (4) Optimize the allocation of audit resources to improve system efficiency and usability.

[0045] Specifically, the beneficial effects of the present invention are reflected in: Improve anomaly detection accuracy: By using multimodal feature fusion and cross-validation, the limitations of single-dimensional analysis are overcome, and the false positive and false negative rates are significantly reduced.

[0046] Optimize resource allocation: The adaptive auditing strategy based on dynamic trust assessment avoids indiscriminate full recording, significantly reducing storage space and computing resource consumption.

[0047] Enhance real-time protection capabilities: Extend auditing from post-event analysis to in-event intervention, realize risk response based on real-time trust assessment, and achieve second-level detection and alarm for high-risk operations.

[0048] Balancing security and privacy: By using privacy protection technologies, we minimize the exposure of raw privacy data without affecting the audit results, in compliance with regulations such as GDPR and the Personal Information Protection Act.

[0049] Improve system adaptability: The dynamic trust assessment model can adapt to different operation and maintenance scenarios and new attack patterns, ensuring the long-term effectiveness of the system.

[0050] Please see Figure 1 , Figure 1 This is a flowchart illustrating an audio auditing method provided in an embodiment of this application. The audio auditing method includes the following steps: S100: Obtain the initial audio data to be audited; S200: Extract multimodal audio features from the initial audio data to obtain audio feature data; S300: Processes audio feature data according to the dynamic trust assessment algorithm to obtain trust score data; S400: Obtain audio auditing strategies based on the trust score-audit strategy library and trust score data; S500: Audit audio feature data according to the audio audit strategy to obtain audio audit results.

[0051] For example, the initial audio data can be the raw audio stream obtained by the operation and maintenance system based on audio recording technology, including voice guidance collaboration, fault emergency response dialogue, and monitoring of equipment operation sounds in industrial environments, etc. This is only an example and not a limitation.

[0052] For example, multimodal audio features are extracted from the initial audio data to obtain audio feature data, which may include: Voiceprint biometric features are continuously extracted, and the current speaker's identity ID and real-time voiceprint matching score are output for subsequent trust assessment. Semantic feature extraction outputs a structured semantic risk report, including transcribed text, identified risk instructions, and context-based classification of operational intent. Extract abnormal machine sound features and output machine health status features, including severity and timestamp of occurrence. Among them, abnormal types include disk bad sector noise, fan failure, server beeping alarm, etc. Emotional stress features are extracted from the initial audio data, including the speaker's pitch jitter, speech rate, energy changes, and other paralinguistic features. A pre-trained classification model is used to determine the speaker's emotional state (calm, tense, angry) and stress level. The stress index feature (0-1) is output as an auxiliary soft signal for trust assessment. Optionally, each feature in the audio feature data refers to a feature vector; where a feature vector is a fixed-length list of numbers condensed from a short segment of audio signal (e.g., 20-40 milliseconds) through a series of mathematical transformations and calculations. Each number in this list (called a "feature" or "dimension") describes a specific attribute of the audio segment, such as energy, frequency distribution, spectral shape, etc.

[0053] For example, multimodal data fusion can be achieved by extracting multimodal features from the initial audio data, which can effectively improve context awareness. Context awareness refers to a system (such as software, devices, applications, etc.) being able to automatically collect, interpret, and utilize its "context" information to provide the most relevant and appropriate services or information for a specific person, place, object, or event. The core idea of ​​context awareness is to understand "who," "when," "where," "what is being done," and "what the surrounding environment is like," and to intelligently adjust its own behavior accordingly.

[0054] For example, the dynamic trust assessment algorithm is a security mechanism that continuously, in real time, and in multiple dimensions assesses and adjusts the trust level of users, devices, applications, or sessions. By processing the audio feature data through the dynamic trust assessment algorithm, the multi-dimensional features of the audio feature data are fused together to obtain the corresponding trust score data, and the trust score representing the real-time risk level of the current audio is calculated, thereby realizing a differentiated audio auditing strategy based on the real-time trust level.

[0055] For example, based on trust score data, a corresponding audio audit strategy is selected to achieve dynamic decision-making on the audit strategy; then, the audio feature data is audited according to the corresponding audio audit strategy, and computing resources can be dynamically allocated according to the strategy to ensure resource utilization efficiency.

[0056] The audio auditing method provided in this application extracts multimodal audio features from initial audio data to obtain audio feature data, thereby enabling the audio feature data to possess multi-dimensional feature data. This audio feature data is then processed using a dynamic trust assessment algorithm to achieve multimodal data fusion and obtain a trust score for the real-time risk level of the current audio. This not only improves the accuracy of anomaly detection and context awareness but also establishes a dynamic trust assessment mechanism, enabling differentiated audio auditing strategies based on real-time trust levels. Therefore, this audio auditing method can achieve the technical effect of improving the accuracy of anomaly detection and real-time protection capabilities.

[0057] Please see Figure 2 , Figure 2 This is a flowchart illustrating another audio auditing method provided in an embodiment of this application.

[0058] In some implementations, S300: the step of processing audio feature data according to a dynamic trust assessment algorithm to obtain trust score data includes: S310: Obtain one or more data from the following factors based on audio feature data: identity credibility factor, behavioral compliance factor, environmental risk factor, and conversation context factor; S320: The adaptive weight model processes one or more of the data from identity credibility factor, behavioral compliance factor, environmental risk factor, and conversation context factor to obtain trust score data. The adaptive weight model applies weighted penalties to each factor based on preset risk events.

[0059] For example, multiple factors are continuously acquired based on audio feature data. By fusing data from multiple dimensions such as identity credibility factor, behavioral compliance factor, environmental risk factor, and conversation context factor, a trust score representing the real-time risk level of the current audio is calculated. An adaptive weight model is used to implement a dynamic trust assessment algorithm, which enables adaptive adjustment of the weights of each factor. Furthermore, each factor can be weighted and penalized based on preset risk events to improve the accuracy of anomaly detection.

[0060] For example, for each factor of the audio feature data: Identity credibility: The core is the voiceprint matching score, combined with multi-factor authentication status; Behavioral compliance: Based on the risk instructions, operational intentions, and whether the frequency and sequence of operations are abnormal, as output by semantic analysis; Environmental risks include access time (whether it is outside of working hours), source network area (whether it comes from an untrusted network), and the severity of abnormal machine noises. Session context: The sensitivity level of the currently accessed resource (e.g., accessing the core database vs. the test server).

[0061] For example, the preset risk events include various risk events, such as detecting abnormal machine sounds in the background, or a user operating the core database. The adaptive weighting model applies weighted penalties to each factor based on preset risk events. For example, when the system detects abnormal machine noise in the background, it will automatically increase the weight of the "environmental risk" factor; when a user is operating the core database, it will increase the weight of the "behavioral compliance" factor.

[0062] In some implementations, S400: the step of obtaining an audio audit policy based on the trust score-audit policy library and trust score data includes: S410: Obtain trust level data based on trust score data; S420: Match the trust score-audit strategy library with the trust level data to obtain the audio audit strategy. The audio audit strategy includes segment capture mode, sampling audit mode and full audit mode. S430: If the audio auditing strategy matches the full audit mode, then allocate preset GPU resources to analyze and audit the initial audio data.

[0063] For example, risk levels are determined based on trust score data, and corresponding audio auditing strategies are selected based on the risk level. High trust only requires the segment capture mode, which triggers the recording and analysis of audio a few seconds before and after a suspected abnormal event is detected. Medium trust requires the sampling audit mode, which randomly records audio segments at a certain proportion and performs standard analysis. Low trust requires the full audit mode, which records all audio completely and initiates deep analysis modes of all modalities. In addition, when executing the full audit mode, preset GPU resources are allocated to analyze the initial audio data for auditing to ensure effective resource utilization.

[0064] For example, GPU (Graphics Processing Unit) resources refer to the computing resources of all the key hardware components within a graphics processor that can be used for computation and data processing.

[0065] For example, risk classification based on trust score data can be divided into three risk levels: High Trust: Segment Capture Mode: Recording and analysis of audio for several seconds before and after the event is triggered only when a suspected abnormal event is detected (such as a brief mismatch in voiceprint or the appearance of a single sensitive word). This greatly saves storage and computing power. Zhongxin Trust: Sampling Audit Model: Audio clips are randomly recorded at a certain percentage (e.g., 30%) and then analyzed using standard methods. This achieves a balance between security and efficiency. Low Trust → "Full Audit" Mode: Record all audio completely and initiate deep analysis modes for all modalities (such as using more complex models for semantic ambiguity analysis).

[0066] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating the process of privacy-processing audio feature data and encrypted audio audit results provided in an embodiment of this application.

[0067] In some implementations, prior to step S300: processing audio feature data according to a dynamic trust assessment algorithm to obtain trust score data, the method further includes: S301: Obtain privacy-preserving audio feature data by processing audio feature data based on differential privacy algorithm and preset noise data; S300: The steps for processing audio feature data according to the dynamic trust assessment algorithm to obtain trust score data include: S302: Process the privacy audio feature data according to the dynamic trust assessment algorithm to obtain trust score data.

[0068] For example, a differential privacy algorithm is used to add precisely calculated statistical noise to the feature vector after audio feature extraction and before the dynamic trust assessment algorithm processes the audio feature data. This protects the privacy of individual speakers while ensuring that the aggregated trust assessment results are basically accurate.

[0069] For example, differential privacy algorithms are a privacy protection framework with rigorous mathematical proofs. Differential privacy algorithms can minimize the risk of leakage to any individual record when extracting and publishing statistical information from a database containing sensitive information. In simple terms, differential privacy algorithms allow the secure sharing of overall patterns, trends, and insights of an entire dataset without exposing any specific personal data.

[0070] In some implementations, after step S500: auditing audio feature data according to an audio auditing strategy to obtain audio audit results, the method further includes: S303: The audio audit results are encrypted based on a semi-homomorphic encryption algorithm to obtain encrypted audio audit results. The encrypted audio audit results are subject to trust-adaptive access control based on the trust level of the accessing user, and can be retrieved and processed before decryption based on preset audit tags.

[0071] For example, a semi-homomorphic encryption algorithm is used to encrypt the audio audit results, allowing the system to retrieve and calculate certain audit tags in the audio audit results in ciphertext without decryption; at the same time, combined with trust-adaptive access control, administrators with low trust levels cannot decrypt and access highly sensitive raw audio data, but can only view the de-identified analysis report.

[0072] For example, Trust-Adaptive Access Control is a dynamic access control model based on real-time trust assessment. It dynamically adjusts access permissions by continuously monitoring multi-dimensional data on users, devices, environment, and behavior, achieving a balance between security and flexibility.

[0073] For example, semi-homomorphic encryption is an encryption technique that allows a specific single type of operation to be performed directly on encrypted data, supporting an unlimited number of addition or multiplication operations, but it cannot efficiently support any combination of the two.

[0074] In some implementations, S500: The step of auditing audio feature data according to an audio auditing strategy to obtain audio audit results includes: S510: Perform cross-modal correlation analysis and early warning on audio feature data to obtain audio audit results; After the step of auditing audio feature data according to the audio auditing strategy and obtaining audio audit results, the method further includes: S520: If the audio audit results include alarm information, execute the preset response strategy according to the alarm information.

[0075] For example, based on cross-modal correlation analysis of audio feature data, multi-dimensional feature correlation analysis and cross-validation of audio feature data can be realized, which can effectively eliminate false alarms of single-modal analysis; if there is an alarm message in the audio audit results, it indicates that there is an anomaly in the audio feature data, and a preset response strategy is executed according to the alarm message.

[0076] Optionally, multi-dimensional feature correlation analysis and cross-validation can be performed on the audio feature data, as shown in the following example: Rule 1: If the voiceprint matching degree is less than the voiceprint threshold and the semantic risk is greater than the semantic threshold, a high-risk "suspected identity impersonation" alarm will be triggered. Rule 2: If a DROP command is detected and the abnormal sound from the machine's disk exceeds the machine's threshold, a "suspected data corruption" high-risk alarm will be triggered. It should be noted that the above rules are only examples and not limitations. The rules for multi-dimensional feature correlation analysis and cross-validation of audio feature data are not limited to the above rules 1 and 2. Rules can be changed, added or deleted according to actual needs.

[0077] In some implementations, after step S520: executing a preset response strategy based on the alarm information, the method further includes: S530: Obtain the response and processing results of alarm information; S540: Mark the alarm feedback category of the alarm information according to the response processing result. The alarm feedback category includes correct alarm, false alarm, and missed alarm.

[0078] For example, alarm events and their corresponding processing results are recorded, and alarm information is marked as correct alarm, false alarm, or missed alarm. Thus, by analyzing the alarm information of false alarm or missed alarm, the audio auditing method can be dynamically adjusted to improve the accuracy of audio auditing.

[0079] In some implementations, after step S540: marking the alarm feedback category of the alarm information according to the processing result, the method further includes: S550: Obtain model adjustment data based on alarm feedback category; S560: Adjust the dynamic trust assessment algorithm or audio audit strategy based on the model adjustment data.

[0080] For example, the dynamic trust assessment algorithm or audio auditing strategy can be adjusted based on the model adjustment data, so that the dynamic trust assessment and multimodal fusion model can be continuously optimized to adapt to new operation and maintenance scenarios and attack methods. For example, for false alarm cases, the detection threshold or trust factor weight of the relevant modality can be fine-tuned; for missed alarms or newly discovered attack patterns, the relevant voiceprint, audio or semantic patterns can be added to the feature library of the audio auditing strategy to update the model.

[0081] In some implementations, S200: the step of extracting multimodal audio features based on initial audio data to obtain audio feature data includes: S210: Extract one or more of the following from the initial audio data: voiceprint biometric features, semantic features, machine health status features, and emotional stress features; S220: Obtain audio feature data based on one or more of the following features: voiceprint biometrics, semantic features, machine health status features, and emotional stress features.

[0082] For example, features from multiple dimensions such as voiceprint biometrics, semantic features, machine health status features, and emotional stress features are extracted in real time, providing rich and structured data for subsequent analysis.

[0083] In some implementations, combined Figures 1 to 3 The audio auditing method shown constructs a dynamic trust-driven, multimodal fusion bastion host audio auditing framework. It adaptively adjusts audio auditing strategies by calculating user trust levels in real time and integrates multi-dimensional information such as voiceprints, semantics, and machine audio for anomaly detection and risk assessment. Specific implementation details are as follows: Step 1: Parallel extraction and fusion of multimodal audio features; This step aims to extract features from the raw audio stream in parallel and in real time, providing rich, structured data for subsequent analysis. 1. Continuous extraction of voiceprint biometric features: Technical Implementation: An improved WaveNet network is combined with a temporal attention mechanism. WaveNet is responsible for extracting fine-grained temporal features from the original audio waveform, while the attention mechanism focuses on the most discriminative segments in the speech (such as vowel segments), thereby generating more robust speaker feature vectors. Output: In addition to outputting the current speaker's ID, the more crucial output is a real-time voiceprint matching score (0-1) for subsequent trust assessment. The system will continuously perform this task, not just once upon login. 2. In-depth analysis of semantic content and intent: Technical Implementation: A domain-adaptive end-to-end ASR (Automatic Speech Recognition) model (such as Conformer) is used. This model's vocabulary and language model are trained on massive amounts of operational logs and manuals, achieving extremely high recognition rates for IT terms (such as "Kubernetes" and "ROLLBACK"). The transcribed text is then fed into a lightweight NLP engine that integrates: Operations and maintenance knowledge graph: used to understand the contextual relationships of instructions (e.g., recognizing that "switching standby database" followed immediately by "deleting primary database" is an abnormal sequence). Risk keyword library: contains keywords with dynamic weights (e.g., "bypass" has a weight of 0.9, "test" has a weight of 0.2). Output: A structured semantic risk report, including transcribed text, identified risk instructions, and context-based classification of operational intent (e.g., routine maintenance, troubleshooting, malicious damage). 3. Abnormal machine sound detection and recognition: Technical Implementation: A dynamic attention mechanism and the lightweight MobileFaceNet network are employed. This network takes the audio log-Mel spectrum as input, and the attention mechanism automatically focuses on abnormal regions in the spectrum (such as sudden howling at specific frequencies). Anomalies are detected by comparing the acoustic fingerprint of a normally functioning device. Output: Machine health status vector, including anomaly type (e.g., disk bad sector noise, fan failure, server beeping alarm), severity, and timestamp of occurrence; 4. Auxiliary analysis of emotion and stress levels: Technical implementation: Extract paralinguistic features such as pitch jitter, speech rate, and energy changes from speech signals, and use a pre-trained classification model to determine the speaker's emotional state (calm, tense, angry) and stress level; Output: Stress index (0-1), serving as an auxiliary soft signal for trust assessment.

[0084] Step 2: Dynamic Trust Assessment and Score Calculation This step is the brain of the system. It integrates the multidimensional features extracted in step 1 and calculates the trust score that represents the real-time risk level of the current session. 1. Multi-dimensional trust factor collection: The system continuously collects factors across four dimensions: Identity credibility: The core is the voiceprint matching score, combined with multi-factor authentication status; Behavioral compliance: Based on the risk instructions, operational intentions, and whether the frequency and sequence of operations are abnormal, as output by semantic analysis; Environmental risks include access time (whether it is outside of working hours), source network area (whether it comes from an untrusted network), and the severity of abnormal machine noises. Session context: The sensitivity level of the currently accessed resource (e.g., accessing the core database vs. the test server); 2. Adaptive weight adjustment and trust integral calculation: Technical Implementation: Trust scores are not calculated using a fixed formula, but rather employ an adaptive weighting model. For example, when the system detects abnormal machine noise in the background, it automatically increases the weight of the "environmental risk" factor; when a user is operating the core database, it increases the weight of the "behavioral compliance" factor.

[0085] The following is an example of a calculation formula (conceptual): T_current = f( W_identity * S_voiceprint, W_behavior * R_semantic, W_context * R_environment, ...) Where T_current is the trust score data, W is the dynamic weight (W_identity is the dynamic weight of identity credibility, W_behavior is the dynamic weight of behavioral compliance, and W_context is the dynamic weight of context), S is the score (S_voiceprint is the voiceprint score), and R is the risk value (R_semantic is the semantic risk value, and R_environment is the environmental risk value); the function f is a non-linear fusion function that applies weighted penalties to sudden multiple risk events.

[0086] Output: A real-time dynamic trust score between 0 and 1, mapped to three trust levels: high (≥0.7), medium (0.4-0.7), and low (<0.4).

[0087] Step 3: Dynamic decision-making and execution of audit strategies: This step is the system's "hand," dynamically adjusting the intensity and method of the audit behavior according to the instructions of the brain (trust points); 1. Strategy Matching Engine: The system has a pre-built audit policy library that is tied to the trust level; High Trust → "Key Segment Capture" Mode: Recording and analyzing audio for several seconds before and after a suspected anomaly is detected only when a brief mismatch in voiceprint is detected (such as the appearance of a single sensitive word). This greatly saves storage and computing power. Zhongxin Trust → "Sampling Audit" Model: Audio clips are randomly recorded at a certain percentage (e.g., 30%) and then analyzed using standard methods. This strikes a balance between security and efficiency. Low Trust → "Full Audit" Mode: Record all audio completely and initiate deep analysis modes for all modalities (such as using more complex models for semantic ambiguity analysis). 2. Adaptive resource allocation: The system dynamically allocates computing resources based on policies. In low-trust scenarios, more GPU resources may be allocated to voiceprint and anomalous sound analysis to ensure analysis accuracy.

[0088] Step 4: Secure storage and processing under privacy protection: This step ensures that sensitive audio data is properly protected throughout the audit process; 1. Privacy protection during processing: Technical Implementation: Differential privacy technology is employed. After audio feature extraction and before being fed into the trust model, precisely calculated statistical noise is added to the feature vector. This protects the privacy of individual speakers while ensuring the overall accuracy of the aggregated trust assessment results. 2. Security and Encryption in Storage: Technical Implementation: Audit logs are encrypted using a semi-homomorphic encryption algorithm. This allows the system to retrieve and perform calculations on certain audit tags in their encrypted state without decryption. Simultaneously, combined with trust-adaptive access control, administrators with low trust levels cannot decrypt or access highly sensitive raw audio data; they can only view the anonymized analysis reports.

[0089] Step 5: Anomaly detection and real-time alarm for multimodal fusion: This step is the system's "early warning mechanism," which accurately analyzes the fused features and decides whether to issue an alert. 1. Cross-modal correlation analysis: Technical Implementation: The system is equipped with an association rule engine. For example: Rule 1: If the voiceprint matching degree is less than the threshold and the semantic risk is greater than the threshold, then a "suspected identity impersonation" high-risk alarm is triggered; Rule 2: If a DROP command is detected and abnormal disk noises exceed the threshold, a "Suspected Data Corruption" high-risk alarm will be triggered. This cross-validation effectively eliminates false alarms in single-modal analysis; 2. Real-time alarms and automatic response: Based on the risk level of the alarm, the system executes predefined response strategies, such as: popping up a red pop-up window in the management interface, sending an SMS to the security manager, or even automatically terminating the current operation and maintenance session.

[0090] Step 6: Feedback Loop and Continuous Optimization of the Trust Model: This step enables the system to learn, allowing it to improve the accuracy of anomaly detection with repeated use. 1. Feedback Data Collection: The system records all alarm events and their handling results. Security administrators need to confirm alarms, marking them as "true positive" (correct alarm), "false positive" (false alarm), or "missed alarm." 2. Incremental Model Update: Technical implementation: Employing online learning or periodic incremental training methods; For false positive cases, the system will fine-tune the detection threshold or trust factor weight of the relevant modality; For "missed" or newly discovered attack patterns, the system will add the relevant voiceprint, audio or semantic patterns to the feature library to update the model; This process enables the system's dynamic trust assessment and multimodal fusion model to continuously evolve, adapting to new operational scenarios and attack methods.

[0091] In some implementation scenarios, for example, operations engineer Zhang San (who has passed two-factor authentication and voiceprint registration) is performing a routine nighttime core database maintenance task; simultaneously, an attacker attempts to hijack Zhang San's session and perform malicious operations using a previously stolen session token. The audio auditing method provided in this application provides the following example of its specific workflow: Phase 1: Session Initiation — Establishing an Initial Trust Baseline; 1. Session initialization: Zhang San logs into the core database cluster via the bastion host. The system performs initial voiceprint verification, and the match is successful. Initial Trust Score: Based on its legitimate identity, normal working hours, and initial authentication, the system sets its initial trust score to 0.75 (which is a medium trust level). 2. Audit strategy activation: Based on the trust level, the system activates the [Sampling Audit] strategy: Audio recording: Record audio by randomly sampling 50% of the samples; Analysis: Perform continuous voiceprint verification and basic semantic analysis on the sampled audio; Phase Two: In Progress of the Session—Multimodal Fusion Detection and Dynamic Trust Decay; The attacker successfully injected commands, and anomalies began to occur in the session. The system's multimodal sensors immediately captured the abnormal signals, and the abnormal events and system responses are mapped in Table 1 below:

[0092] Table 1 - Mapping Table of Abnormal Events and System Responses Dynamic trust assessment results: Real-time trust score: 0.75 - 0.15 - 0.25 - 0.10 = 0.25; Trust level: Dropped sharply from Medium Trust to Low Trust.

[0093] Phase Three: Real-time Response and Risk Intervention; 1. Dynamic strategy upgrade: If the trust score falls below the threshold of 0.5, the system will immediately upgrade the audit strategy from [sampling audit] to [full audit + real-time alert]; Begin full recording of all audio and start deep scanning of all analysis modules; 2. Automatic safety response: Session interrupted: The system forcibly interrupted the maintenance session just before the DROP TABLE command was to be executed; Real-time alerts: The highest-level alert pops up on the Security Operations Center (SOC) dashboard, detailing the risk factors: [Urgent] Suspected session hijacking and malicious data corruption detected! Risky user: Zhang San (identity suspected of being impersonated); Related events: abnormal voiceprint + high-risk commands + unusual disk noises; Action performed: Session blocked; Two-factor authentication: The system sends a two-factor authentication request to Zhang San's registered mobile phone, requiring him to reconfirm his identity through biometrics; 3. Preservation of the chain of evidence: The system associates and tags the full recordings, voiceprint comparison results, semantic analysis logs, machine audio spectrograms, etc., and securely stores them using semi-homomorphic encryption to form a complete chain of evidence for judicial evidence collection.

[0094] Phase Four: Post-analysis and model optimization; 1. Event review: The security administrator confirmed that this was a successful insider threat containment. The attacker exploited an unknown middleware vulnerability to hijack sessions. The administrator marked this event as a "correct alert" in the system; 2. System self-learning: Feedback loop: The system receives confirmation feedback from the administrator; Model update: Trust Model: Fine-tuned to give a higher trust penalty weight when sensitive commands such as DROP are detected, especially if accompanied by abnormal voiceprints. Voiceprint model: Add unknown voiceprint fragments of attackers to the "blacklist voiceprint library" for faster identification in the future; Machine audio model: Enhanced the abnormal sound characteristics of "large-scale disk deletion operation", improving detection sensitivity.

[0095] By way of example, the audio auditing method provided in this application embodiment has at least the following beneficial effects: Improve anomaly detection accuracy: By using multimodal feature fusion and cross-validation, the limitations of single-dimensional analysis are overcome, and the false positive and false negative rates are significantly reduced; Optimize resource allocation: The adaptive auditing strategy based on dynamic trust assessment avoids indiscriminate full recording, significantly reducing storage space and computing resource consumption; Enhance real-time protection capabilities: Extend auditing from post-event analysis to in-event intervention, realize risk response based on real-time trust assessment, and achieve second-level detection and alarm for high-risk operations; Balancing security and privacy: By using privacy protection technologies, we minimize the exposure of raw privacy data without affecting the audit results, in compliance with regulations such as GDPR and the Personal Information Protection Act. Improve system adaptability: The dynamic trust assessment model can adapt to different operation and maintenance scenarios and new attack patterns, ensuring the long-term effectiveness of the system.

[0096] Please see Figure 4 , Figure 4 This is a structural block diagram of an audio auditing system provided in an embodiment of this application. The audio auditing system includes: The audio acquisition module 100 is used to acquire the initial audio data to be audited; The audio feature module 200 is used to extract multimodal audio features from the initial audio data to obtain audio feature data; The trust score module 300 is used to process audio feature data according to the dynamic trust assessment algorithm to obtain trust score data; Audit strategy module 400 is used to obtain audio audit strategies based on the trust score-audit strategy library and trust score data; The audit results module 500 is used to audit audio feature data according to the audio audit strategy and obtain audio audit results.

[0097] In some implementations, the trust score module 300 is also used to: obtain one or more data from identity credibility factor, behavioral compliance factor, environmental risk factor, and session context factor based on audio feature data; process one or more data from identity credibility factor, behavioral compliance factor, environmental risk factor, and session context factor based on an adaptive weight model to obtain trust score data, wherein the adaptive weight model applies weighted penalties to each factor based on preset risk events.

[0098] In some implementations, the audit strategy module 400 is also used to: obtain trust level data based on trust score data; match the trust score-audit strategy library based on the trust level data to obtain an audio audit strategy, the audio audit strategy including segment capture mode, sampling audit mode, and full audit mode; if the audio audit strategy matches the full audit mode, then allocate preset GPU resources to analyze and audit the initial audio data.

[0099] In some implementations, the audio auditing system further includes a privacy processing module for: processing audio feature data based on a differential privacy algorithm and preset noise data to obtain privacy-preserving audio feature data; The trust score module 300 is also used to: process privacy audio feature data according to the dynamic trust assessment algorithm to obtain trust score data.

[0100] In some implementations, the audio auditing system further includes an encryption module for: encrypting the audio audit results based on a semi-homomorphic encryption algorithm to obtain encrypted audio audit results, wherein the encrypted audio audit results perform trust-adaptive access control based on the trust level of the accessing user, and can be retrieved and processed without decryption based on preset audit tags.

[0101] In some implementations, the audit results module 400 is also used to: perform cross-modal correlation analysis and early warning on audio feature data, and obtain audio audit results; The audio auditing system also includes an alarm module, which is used to execute a preset response strategy based on the alarm information if the audio audit results include alarm information.

[0102] In some implementations, the alarm module is also used to: obtain the response processing result of the alarm information; and mark the alarm feedback category of the alarm information according to the response processing result, wherein the alarm feedback category includes correct alarm, false alarm, and missed alarm.

[0103] In some implementations, the audio auditing system also includes a model adjustment module for: obtaining model adjustment data based on alarm feedback categories; and adjusting the dynamic trust assessment algorithm or audio auditing strategy based on the model adjustment data.

[0104] In some implementations, the audio feature module 200 is also used to: extract one or more of voiceprint biometric features, semantic features, machine health status features, and emotional stress features from the initial audio data; and obtain audio feature data based on one or more of the voiceprint biometric features, semantic features, machine health status features, and emotional stress features.

[0105] It should be noted that the audio auditing system provided in this application embodiment is different from... Figures 1 to 3 The method embodiments shown correspond to each other, and will not be described again here to avoid repetition.

[0106] This application also provides an electronic device, please refer to [link to application]. Figure 5 , Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device may include a processor 510, a communication interface 520, a memory 530, and at least one communication bus 540. The communication bus 540 is used to enable direct communication between these components. In this embodiment, the communication interface 520 of the electronic device is used for signaling or data communication with other node devices. The processor 510 may be an integrated circuit chip with signal processing capabilities.

[0107] The processor 510 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor, or the processor 510 can be any conventional processor.

[0108] The memory 530 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory 530 stores computer-readable instructions. When these computer-readable instructions are executed by the processor 510, the electronic device can perform the aforementioned operations. Figures 1 to 3 The various steps involved in the method implementation examples.

[0109] Alternatively, the electronic device may also include a storage controller and an input / output unit.

[0110] The memory 530, storage controller, processor 510, peripheral interface, and input / output unit are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 540. The processor 510 is used to execute executable modules stored in the memory 530, such as software function modules or computer programs included in electronic devices.

[0111] The input / output unit is used to provide users with the ability to create tasks and to set optional start periods or preset execution times for those tasks, thereby enabling user-server interaction. The input / output unit may be, but is not limited to, a mouse and keyboard.

[0112] Understandable. Figure 5 The structure shown is for illustrative purposes only; the electronic device may also include components that are more advanced than those shown. Figure 5 The more or fewer components shown, or having the same Figure 5 The different configurations shown. Figure 5 The components shown can be implemented using hardware, software, or a combination thereof.

[0113] This application also provides a storage medium storing instructions. When the instructions are run on a computer, the computer program is executed by a processor to implement the method described in the method embodiment. To avoid repetition, the method will not be described again here.

[0114] This application also provides a computer program product that, when run on a computer, causes the computer to perform the method described in the method embodiment.

[0115] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0116] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0117] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0118] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0119] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0120] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. An audio auditing method, characterized in that, include: Obtain the initial audio data to be audited; Multimodal audio features are extracted from the initial audio data to obtain audio feature data; The audio feature data is processed according to the dynamic trust assessment algorithm to obtain trust score data; An audio audit strategy is obtained based on the trust score-audit strategy library and the trust score data; The audio feature data is audited according to the audio audit strategy to obtain audio audit results.

2. The audio auditing method according to claim 1, characterized in that, The steps of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data include: Based on the audio feature data, obtain one or more of the following data: identity credibility factor, behavioral compliance factor, environmental risk factor, and session context factor; The trust score data is obtained by processing one or more of the data from the identity credibility factor, the behavior compliance factor, the environmental risk factor, and the session context factor according to the adaptive weight model. The adaptive weight model applies weighted penalties to each factor based on preset risk events.

3. The audio auditing method according to claim 1 or 2, characterized in that, The steps for obtaining an audio audit strategy based on the trust score-audit strategy library and the trust score data include: Based on the trust score data, the trust level data is obtained; Based on the trust level data, the trust score-audit strategy library is matched to obtain an audio audit strategy, which includes a segment capture mode, a sampling audit mode, and a full audit mode. If the audio audit strategy matches the full audit mode, then preset GPU resources are allocated to analyze and audit the initial audio data.

4. The audio auditing method according to claim 1, characterized in that, Before the step of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data, the method further includes: The audio feature data is processed using a differential privacy algorithm and preset noise data to obtain privacy-preserving audio feature data. The steps of processing the audio feature data according to the dynamic trust assessment algorithm to obtain trust score data include: The privacy audio feature data is processed according to a dynamic trust assessment algorithm to obtain trust score data.

5. The audio auditing method according to claim 1 or 4, characterized in that, After the step of auditing the audio feature data according to the audio auditing strategy and obtaining the audio audit results, the method further includes: The audio audit results are encrypted using a semi-homomorphic encryption algorithm to obtain encrypted audio audit results. The encrypted audio audit results are subject to trust-adaptive access control based on the trust level of the accessing user, and can be retrieved and processed before decryption based on preset audit tags.

6. The audio auditing method according to claim 1, characterized in that, The steps of auditing the audio feature data according to the audio auditing strategy and obtaining the audio audit results include: Perform cross-modal correlation analysis and early warning on the audio feature data to obtain the audio audit results; After the step of auditing the audio feature data according to the audio audit strategy and obtaining the audio audit result, the method further includes: If the audio audit results include alarm information, a preset response strategy is executed based on the alarm information.

7. The audio auditing method according to claim 6, characterized in that, After executing a preset response strategy based on the alarm information, the method further includes: Obtain the response and processing results of the alarm information; The alarm feedback category of the alarm information is marked according to the response processing result. The alarm feedback category includes correct alarm, false alarm, and missed alarm.

8. The audio auditing method according to claim 7, characterized in that, After the step of labeling the alarm feedback category of the alarm information based on the processing result, the method further includes: Based on the alarm feedback category, obtain model adjustment data; The dynamic trust assessment algorithm or the audio audit strategy is adjusted based on the model adjustment data.

9. The audio auditing method according to claim 1, characterized in that, The steps of extracting multimodal audio features from the initial audio data to obtain audio feature data include: Extract one or more of the following from the initial audio data: voiceprint biometric features, semantic features, machine health status features, and emotional stress features. Audio feature data is obtained based on one or more of the following features: voiceprint biometrics, semantic features, machine health status features, and emotional stress features.

10. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the audio auditing method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the audio auditing method as described in any one of claims 1 to 9.

12. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the audio auditing method as described in any one of claims 1 to 9.