Audio content analysis method and device, electronic equipment, product and storage medium

By processing and mapping audio signals, interpretable concept activation sequences are generated, solving the problems of accuracy and interpretability in audio content analysis in driver-passenger interaction scenarios. This enables real-time risk monitoring and immediate alerts, and constructs a complete system for practical applications.

CN122050422APending Publication Date: 2026-05-15BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DIDI INFINITY TECH & DEV CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the analysis of audio content in driver-passenger interaction scenarios suffers from insufficient accuracy and interpretability, making it difficult to achieve real-time risk monitoring and early warning, and existing methods are insufficient to assist in further decision-making.

Method used

By processing the audio signal, a hidden state sequence is obtained and mapped to an interpretable concept activation sequence. The hidden state sequence is then mapped to the concept space using a concept bottleneck layer to generate an aligned concept activation sequence. Finally, a concept analysis report is output, including the concept analysis results of the audio signal and the concept analysis results of the time steps.

Benefits of technology

It improves the accuracy and stability of audio content analysis and detection, has good interpretability, can clearly give the reasons and justifications for the judgment results, realizes real-time streaming processing and instant alarms, and builds a complete system for practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050422A_ABST
    Figure CN122050422A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides an audio content analysis method and device, electronic equipment, a product and a storage medium. According to the embodiment of the invention, the high-dimensional hidden state sequence is extracted, and the hidden state sequence is mapped into the interpretable concept activation sequence, so that the continuous hidden vector which is originally difficult to interpret can be converted into the intermediate concept representation which can be understood and audited; the finally output concept analysis report not only can provide the concept analysis result of the audio signal and the concept analysis result of one or more time steps in the audio signal, but also can provide interpretable evidence of the concept analysis results. According to the embodiment of the invention, the stability and accuracy of audio content analysis and detection in a complex voice scene are improved, and the model judgment process has good interpretability, so that the technical problems of opaque decision basis and difficulty in tracing of an end-to-end audio content analysis and detection method in the prior art are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to methods, apparatus, electronic devices, products and storage media for audio content analysis. Background Technology

[0002] With the widespread adoption of ride-hailing and taxi services, voice interactions between drivers and passengers are becoming increasingly frequent. These interactions often involve various types of verbal behavior, such as friendly conversations, exchanges of social news, and expressions of emotions like joy, sadness, and happiness. Furthermore, the inherent concealment and difficulty in obtaining evidence in these interactions pose challenges to platform governance and passenger safety. Current technologies primarily rely on manual complaints and post-event reviews to address these verbal behaviors. However, these methods suffer from slow response times and high subjectivity, failing to meet the real-time risk monitoring and early warning needs of large-scale platforms. Therefore, how to automatically detect various verbal behaviors in driver-passenger interactions using in-vehicle voice data without disrupting the normal travel experience has become a crucial research direction in the fields of in-vehicle intelligent safety and voice interaction.

[0003] Currently, there are two main technical solutions for automatic analysis of speech audio. The first solution can be described as automatic audio detection and analysis based on a two-stage process: speech-to-text conversion and text analysis modeling. However, this solution is highly dependent on the accuracy of speech recognition and can lead to error accumulation. The second solution can be described as end-to-end automatic audio detection and analysis based on a large speech model. However, this solution lacks interpretability and is difficult to use to support further decision-making. Summary of the Invention

[0004] This application provides methods, apparatus, electronic devices, products, and storage media for audio content analysis, which can solve the technical problems of insufficient accuracy and interpretability of existing audio content analysis solutions, making it difficult to assist further decision-making.

[0005] In a first aspect, embodiments of this application provide a method for audio content analysis, including:

[0006] The audio signal is processed to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; The hidden state sequence is mapped to an interpretable concept activation sequence, the concept activation sequence including the concept activation vectors of the one or more time steps; Perform concept analysis on the concept activation sequence and output a concept analysis report, which includes at least one of the following: the concept analysis results of the audio signal, and the concept analysis results of one or more time steps in the audio signal.

[0007] In one possible implementation of the first aspect, processing the audio signal to obtain a hidden state sequence of the audio signal includes: The audio signal is preprocessed to obtain a preprocessed audio signal. The preprocessing includes at least one of the following: resampling, windowing, and frame splitting. The preprocessed audio signal is input into the audio big model to obtain the hidden state sequence output by the audio big model.

[0008] In one possible implementation of the first aspect, mapping the hidden state sequence to an interpretable concept activation sequence includes: The hidden state sequence is mapped to the concept space based on the concept bottleneck layer, and the mapped text concepts are aligned with the time steps of the audio signal to generate an aligned and interpretable concept activation sequence.

[0009] In one possible implementation of the first aspect, the conceptual bottleneck layer is a lightweight neural network, which includes any of the following: a fully connected layer, a multilayer perceptron, or a convolutional neural network.

[0010] In one possible implementation of the first aspect, the method further includes: During the training of the concept bottleneck layer, a correspondence is established between labeled data, time steps, and concepts; The loss function is set by comparing the concept activation value and the true labeled value of the concept bottleneck layer step by step.

[0011] In one possible implementation of the first aspect, the concept analysis of the concept activation sequence and the output of a concept analysis report include at least one of the following: The concept with the highest overall activation level or exceeding the first threshold is identified as the main concept of the audio signal; When the activation level of a concept exceeds a second threshold at multiple consecutive time steps, the start and end time steps of the concept are determined based on the multiple consecutive time steps.

[0012] Secondly, embodiments of this application provide an apparatus for audio content analysis, comprising: The first obtaining module is used to process the audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; The second mapping module is used to map the hidden state sequence into an interpretable concept activation sequence, wherein the concept activation sequence includes the concept activation vectors of the one or more time steps; The third output module is used to perform concept analysis on the concept activation sequence and output a concept analysis report. The concept analysis report includes at least one of the following: the concept analysis result of the audio signal, and the concept analysis result of one or more time steps in the audio signal.

[0013] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device performs the method as described in any one of the first aspects above.

[0014] Fourthly, embodiments of this application provide a computer program product, including a computer program, which, when run, causes the method as described in any one of the first aspects above to be performed.

[0015] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as described in any one of the first aspects above.

[0016] The beneficial effects of the first aspect of this application compared with the prior art are: This application embodiment processes an audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; maps the hidden state sequence to an interpretable concept activation sequence, the concept activation sequence including concept activation vectors at the one or more time steps; performs concept analysis on the concept activation sequence, and outputs a concept analysis report, the concept analysis report including at least one of the following: the concept analysis result of the audio signal, and the concept analysis result of one or more time steps in the audio signal. This application embodiment, by extracting a high-dimensional hidden state sequence and mapping the hidden state sequence to an interpretable concept activation sequence, can transform the originally difficult-to-interpret continuous latent vectors into understandable and auditable intermediate concept representations. The final output concept analysis report not only provides the concept analysis result of the audio signal, the concept analysis result of one or more time steps in the audio signal, but also provides interpretable evidence of the concept analysis result. The concept analysis report output by this application embodiment no longer relies solely on the direct classification result of the latent vectors, but is based on inference and judgment based on the combination and activation of concept sets in the concept space. This application not only improves the stability and accuracy of audio content analysis and detection in complex speech scenarios, but also makes the model judgment process highly interpretable, clearly indicating "whether it is a certain type of content" and "the reasons and justifications for determining it as a certain type of content." This effectively overcomes the technical problems of opaque and difficult-to-trace decision-making basis in existing end-to-end audio content analysis and detection methods. Furthermore, this application helps to build a complete system oriented towards practical applications. It not only improves the overall accuracy and stability of audio content analysis and detection, but more importantly, it possesses real-time streaming processing capabilities, enabling online analysis and immediate alerts for continuous speech streams. This marks a shift from "post-event analysis" to "real-time intervention," providing a directly deployable and comprehensive solution (detection, location, interpretation, and alerting) for content security application scenarios.

[0017] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an existing audio content analysis method. Figure 2 This is a flowchart illustrating another audio content analysis method in the existing technology; Figure 3 This is a flowchart illustrating an audio content analysis method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an audio content analysis device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0021] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0022] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0023] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0024] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0026] Figure 1 This is a flowchart illustrating an existing audio content analysis method, showing a two-stage process for automatic audio detection and analysis based on "speech-to-text" and "text analysis model".

[0027] In existing technologies, two-stage automatic audio detection and analysis is a relatively mature and widely adopted technical solution for speech detection in conversations between drivers and passengers inside vehicles. For example... Figure 1 As shown, this technical solution typically begins with audio acquisition by an in-vehicle audio acquisition module. Subsequently, an audio preprocessing module performs preprocessing operations such as noise reduction, VAD (Voice Activity Detection), and segmentation. Next, an Automatic Speech Recognition (ASR) module performs speech recognition processing on the preprocessed in-vehicle audio signal, converting continuous speech input into a corresponding text sequence. Then, a text processing module performs sentence segmentation and alignment on the text sequence. Finally, a text analysis module analyzes and detects the audio content of the dialogue to obtain detection results. In one application scenario, the text analysis module can detect offensive language such as insults and threats. In another application scenario, the text analysis module can classify audio content, such as social news content or lighthearted chat content.

[0028] In its implementation, this technical solution generally includes the following key steps. First, the voice signals of the driver and passengers are acquired through an in-vehicle microphone or in-vehicle audio acquisition device. These voice signals undergo preprocessing operations such as noise reduction, echo cancellation, and voice activity detection to reduce the impact of environmental noise and non-voice segments on subsequent processing. Then, the preprocessed voice signal is input into a speech recognition module (i.e., an ASR speech transcription-related model), which transcribes the voice content to generate corresponding text information. After obtaining the text information, the system typically further processes the text by segmenting, paragraphing, or time alignment, and then inputs the processed text into a text analysis module. The text analysis module used in this technical solution can be based on preset sensitive word rules, statistical features, or a text classification model designed based on machine learning, in existing technologies. With the development of large models in recent years, the text analysis module in this technical solution is usually based on a good open-source general-purpose large text model. Then, the model is fine-tuned based on various content corpora (such as corpora for conversation content classification, conversation emotion classification, etc.) to make it well adaptable to audio content analysis and detection scenarios.

[0029] The two-stage audio automatic detection and analysis framework can directly reuse existing mature speech recognition technology and text security detection model, thus it has the characteristics of relatively low implementation cost and fast engineering implementation.

[0030] However, two-stage automatic audio detection and analysis technologies are highly dependent on the accuracy of speech recognition, which can easily lead to error accumulation. Two-stage audio content analysis and detection technology based on speech-to-text conversion has a clear division between speech processing and text analysis, and its overall detection performance largely depends on the accuracy of the speech recognition results. When there are complex situations in the in-vehicle environment, such as noise interference, multiple speakers, or regional accents, speech-to-text errors are easily amplified in subsequent text analysis, leading to a decrease in the accuracy and stability of the audio content analysis and detection results.

[0031] Figure 2 This is a flowchart illustrating another audio content analysis method in the prior art, showing an end-to-end automatic audio detection and analysis process based on a large speech model.

[0032] With the development of deep learning and large speech model technology, audio content analysis and detection solutions based on end-to-end speech modeling have gradually attracted widespread attention. This type of solution no longer explicitly separates speech processing and semantic analysis into two independent stages. Instead, it directly inputs the collected in-vehicle speech signals into the speech model, and through overall modeling of the audio signal, achieves automatic analysis and judgment of speech behaviors for various content.

[0033] In practical implementation, this type of technology typically begins with audio acquisition by an in-vehicle audio acquisition module. Next, the acquired speech signal undergoes audio preprocessing, such as noise reduction, speech activity detection, or fixed-length slicing, to obtain audio segments suitable for model input. Subsequently, the preprocessed audio segments are directly input into a pre-trained or fine-tuned speech model, which simultaneously learns acoustic features, prosodic variations, emotional cues, and latent semantic information from the audio signal, forming an audio representation in a unified feature space. Based on this, the model, through its internal discriminative layer or task head, predicts whether the dialogue content corresponding to the input audio contains various conceptual elements. For example... Figure 2 As shown, in the Supervised Fine-Tuning (SFT) stage of the large-scale speech model, the speech encoder performs semantic encoding on the audio. Subsequently, the semantic vector understanding module performs semantic understanding and outputs a judgment result to achieve fine-grained annotation based on the application scenario. In one application scenario, the SFT stage can achieve fine-grained annotation for detecting aggressive behavior such as insults and threats. In another application scenario, the SFT stage can achieve fine-grained annotation for content classification of audio content, such as social news content or casual chat content.

[0034] Compared to two-stage speech-to-text technology, this end-to-end technical solution reduces intermediate processing steps in its structure, which can avoid the direct impact of speech transcription errors on the detection results to a certain extent. At the same time, it has the potential to jointly model tone, emotion and semantics from the original audio.

[0035] However, existing end-to-end audio content analysis and detection technologies based on large speech models typically employ highly integrated deep model structures, directly outputting audio content analysis and judgment results from the raw audio signal. Their internal decision-making processes are highly implicit, making it difficult to explicitly reflect the specific speech segments, semantic factors, or acoustic feature sources upon which the model bases its judgments. Due to the lack of interpretable descriptions of the criteria used for audio content analysis and judgment, such technical solutions struggle to clearly explain the specific time and location of the speech act corresponding to specific content, the corresponding speaker, or the content type. This results in low transparency of the detection results, hindering subsequent manual review, accountability determination, and system behavior auditing, thus leading to insufficient reliability and controllability in practical applications.

[0036] Figure 3 This is a flowchart illustrating an audio content analysis method provided in an embodiment of this application.

[0037] S31, process the audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states of one or more time steps.

[0038] Step S31 can be called the audio deep representation and temporal alignment stage. In this stage, the Audio Language Model (ALM) can be used as the core computing engine, typically employing a large-scale pre-trained architecture based on Transformers. The ALM receives pre-processed speech frame sequences and, through its deep self-attention mechanism and multi-layer perceptual network, performs hierarchical abstraction and understanding of the speech signals. Instead of directly outputting classification results, the ALM generates a high-dimensional hidden state vector at each time step (usually corresponding to tens of milliseconds of speech). This hidden state vector is a dense semantic representation, capturing not only linguistic information such as phonemes and vocabulary, but also encoding intonation variations, tone strength, emotional tendency, and even speaker-specific prosodic features. The final output hidden state sequence constitutes a feature tensor with high temporal resolution and rich semantic information, typically a multi-dimensional vector composed of batch size, time steps, and hidden dimensions. Here, "batch size" refers to the number of samples input into the model for parallel computation at once. The "time step" can be strictly correlated with the temporal sequence of the speech frames corresponding to the audio signal. The "hidden dimension" refers to the length of the model's hidden state vector, which can also be understood as the capacity of the model to "remember" or "understand" information at each step.

[0039] The inventiveness of this application lies in completely eliminating the reliance on Automatic Speech Recognition (ASR) for text transcription, thus avoiding the cascading effects of ASR recognition errors in noisy environments, accents, and proper nouns on subsequent analysis, and achieving end-to-end deep understanding of the original speech signal. Meanwhile, the strict time alignment characteristic is the foundation for the accurate positioning achieved by this technical solution; the end-to-end audio large model and the time alignment with preprocessing ensure real-time performance and stability.

[0040] S32, the hidden state sequence is mapped to an interpretable concept activation sequence, the concept activation sequence including the concept activation vectors of the one or more time steps.

[0041] Step S32 can be called the conceptual mapping and explicit supervised learning stage. This stage is a key bridge connecting deep features and human-understandable semantics, and is also the core design for achieving model interpretability and controllability.

[0042] For example, a low-dimensional, expert-defined "conceptual space for speech content subject classification" can be constructed first. Each concept in this conceptual space (e.g., "social news," "entertainment news," "casual chat") has a clear and unambiguous semantic boundary, and the conceptual space can cover, for example, 5 to 20 core categories. These concepts constitute the "semantic dictionary" for the system to reason and make decisions. Optionally, the conceptual space can also include concepts at multiple levels, using a hierarchical structure to expand the coverage of the conceptual space. For example, the main category of entertainment news can include subcategories such as film and television entertainment news, music entertainment news, and short video entertainment news. In the embodiments of this application, "speech content subject classification" is often used as an example for illustration. Those skilled in the art should understand that constructing the conceptual space with other concepts is also within the scope of protection of this application. For example, the conceptual space can also be constructed using concepts such as speech safety and insecurity, such as casual chat and unfriendly chat.

[0043] The Concept Bottleneck Layer (CBL) can be used as the execution module for this stage. Essentially, the CBL is a feature projection and transformer that receives a high-dimensional hidden state sequence from a large audio model and maps it to a predefined concept space using a lightweight neural network (such as a fully connected layer, multilayer perceptron, or convolutional neural network). For each time step's hidden state, the CBL outputs a concept activation vector, where each scalar value represents the association strength between the current speech segment and a specific concept. This process can be formalized as: Concept Activation Sequence = CBL{ALM(Hidden State Sequence)}. The hidden states of the time steps included in the hidden state sequence are mapped to the concept activation vectors of the time steps included in the concept activation sequence. The meaning of the CBL is that it forces the output of one hidden layer of the neural network to align with the physical meaning in the real world, thereby improving the interpretability of the neural network.

[0044] S33, perform concept analysis on the concept activation sequence and output a concept analysis report, the concept analysis report including at least one of the following: the concept analysis result of the audio signal, the concept analysis result of one or more time steps in the audio signal.

[0045] Step S33 can be termed the concept-based decision localization and output detection stage. In this stage, the interpretable intermediate representation (concept activation sequence) generated in S32 is transformed into structured decisions and actionable outputs oriented towards actual business operations. The decision logic is no longer the direct classification of a black-box neural network, but rather transparent reasoning based on concept activation. This transparent decision-making mechanism based on concept activation can support concept type identification and accurate time-based localization.

[0046] First, the activation intensity sequence of each speech content theme concept can be analyzed over time. By setting dynamic or static thresholds, it is possible to identify which time intervals show a sustained activation intensity exceeding the threshold. This process simultaneously accomplishes two key tasks: 1) Type identification: By determining which concept has the highest overall activation level or exceeds the threshold, the concept analysis result of the audio signal is determined, i.e., the main type of speech content theme corresponding to the overall audio signal. 2) Time localization: By identifying consecutive time periods exceeding the threshold, the start and end times of the speech content theme discussion can be precisely defined with frame-level accuracy (e.g., every 50 milliseconds). In other words, the concept analysis result of one or more time steps in the audio signal.

[0047] Based on the above analysis, a structured and information-rich results report can be generated. The concept analysis report includes at least one of the following: the concept analysis results of the audio signal, and the concept analysis results of one or more time steps in the audio signal.

[0048] Taking speech content subject matter as an example, the corresponding concept analysis report typically includes: the judgment result of speech content subject matter (yes / no), the specific type label of speech content subject matter, the precise start and end timestamps, and explanatory evidence centered on key concepts and their activation intensity.

[0049] Taking speech content security as an example, in real-time streaming scenarios, once an unsafe speech segment that meets the conditions is detected, a low-latency alarm signal can be triggered immediately to notify supervisors or trigger an automatic intervention mechanism (such as voice interruption, real-time muting, etc.).

[0050] The final concept analysis report output at this stage is not merely a simple classification label, but a complete "analysis report," allowing downstream applications to handle it flexibly. For example, taking speech content security as an example, the security platform can accurately intercept or label content based on its type and timestamp. Reviewers can directly jump to the offending segments for review, improving work efficiency. In the event of disputes regarding in-vehicle voice content, this application embodiment can provide concept activation evidence as a transparent record of the decision-making process. The concept analysis report provided by this application embodiment signifies that the concept analysis results of audio content have been upgraded from a single "detection tool" to a comprehensive "content security analysis platform."

[0051] This application embodiment processes an audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; maps the hidden state sequence to an interpretable concept activation sequence, the concept activation sequence including concept activation vectors at the one or more time steps; performs concept analysis on the concept activation sequence, and outputs a concept analysis report, the concept analysis report including at least one of the following: the concept analysis result of the audio signal, and the concept analysis result of one or more time steps in the audio signal. This application embodiment, by extracting a high-dimensional hidden state sequence and mapping the hidden state sequence to an interpretable concept activation sequence, can transform the originally difficult-to-interpret continuous latent vectors into understandable and auditable intermediate concept representations. The final output concept analysis report not only provides the concept analysis result of the audio signal and the concept analysis result of one or more time steps in the audio signal, but also provides interpretable evidence of the concept analysis result. The concept analysis report no longer relies solely on the direct classification result of the latent vectors, but is based on inference and judgment based on the combination and activation of concept sets in the concept space. This application not only improves the stability and accuracy of audio content analysis and detection in complex speech scenarios, but also makes the model judgment process highly interpretable, clearly indicating "whether it is a certain type of content" and "the reasons and justifications for determining it as a certain type of content." This effectively overcomes the technical problems of opaque and difficult-to-trace decision-making basis in existing end-to-end audio content analysis and detection methods. Furthermore, this application helps to build a complete system oriented towards practical applications. It not only improves the overall accuracy and stability of audio content analysis and detection, but more importantly, it possesses real-time streaming processing capabilities, enabling online analysis and immediate alerts for continuous speech streams. This marks a shift from "post-event analysis" to "real-time intervention," providing a directly deployable and comprehensive solution (detection, location, interpretation, and alerting) for content security application scenarios.

[0052] In some embodiments, S31 of the above embodiment, the processing of the audio signal to obtain the hidden state sequence of the audio signal, includes the following S311 and S312.

[0053] The core objective of step S31 is to establish a complete mapping pipeline from the raw audio signal to refined, semantic time-series features.

[0054] S311, preprocess the audio signal to obtain a preprocessed audio signal, wherein the preprocessing includes at least one of the following: resampling, windowing, and frame splitting.

[0055] The input continuous speech stream or audio signal file can first undergo standardization preprocessing to obtain a preprocessed audio signal. This preprocessing includes, but is not limited to, at least one of the following: resampling at a preset sampling rate, windowing, framing, etc. Preprocessing ensures the uniformity of the audio signal over time and establishes an accurate timestamp index for each audio frame. The preprocessing process provides a reliable time alignment reference for all subsequent modules.

[0056] S312, input the preprocessed audio signal into the audio large model to obtain the hidden state sequence output by the audio large model.

[0057] The audio big data model can receive pre-processed audio signals and, through its deep self-attention mechanism and multilayer perceptual network, perform hierarchical abstraction and understanding of the speech signals. Instead of directly outputting classification results, the audio big data model generates a high-dimensional hidden state vector at each time step (typically corresponding to tens of milliseconds of speech). This hidden state vector is a dense semantic representation that not only captures linguistic information such as phonemes and vocabulary, but also encodes intonation variations, strength of voice, emotional inclination, and even speaker-specific prosodic features.

[0058] In some embodiments, S32 in the above embodiments, which maps the hidden state sequence to an interpretable concept activation sequence, includes the following S321.

[0059] S321, Based on the concept bottleneck layer, the hidden state sequence is mapped to the concept space, and the mapped text concept is aligned with the time step of the audio signal to generate an aligned and interpretable concept activation sequence.

[0060] Here, by aligning the time steps of the text concepts with the time steps of the audio signals, the comparability between the time steps of the concept activation sequence and the time steps of the original audio signals can be ensured, thereby generating an aligned and interpretable concept activation sequence.

[0061] This application embodiment maps the hidden state sequence to the concept space based on the concept bottleneck layer, and maps audio features to a clear semantic concept space, thereby achieving model interpretability.

[0062] In some embodiments, the conceptual bottleneck layer is a lightweight neural network, which includes any of the following: a fully connected layer, a multilayer perceptron, or a convolutional neural network.

[0063] The concept bottleneck layer is essentially a feature projection and transformer. It receives a high-dimensional hidden state sequence from a large audio model and maps it to a predefined concept space using a lightweight neural network (such as a fully connected layer, multilayer perceptron, or convolutional neural network). For each time step of the hidden state, the concept bottleneck layer outputs a concept activation vector, where each scalar value represents the association strength between the current speech segment and a specific concept. Using convolutional neural networks to construct the concept bottleneck layer can better capture temporal dependencies.

[0064] In some embodiments, the audio content analysis method further includes the following steps S33 and S34.

[0065] S33, During the training of the concept bottleneck layer, establish the correspondence between labeled data, time steps, and concepts.

[0066] Here, a fine-grained, timestamped, concept-level strong supervision approach can be used to ensure the effectiveness of the concept bottleneck layer. The labeled data can clearly correspond to which time steps within which time period of the audio signal (e.g., from 3.2 seconds to 5.1 seconds), and which one or more concepts it corresponds to, thereby establishing the correspondence between labeled data, time steps, and concepts.

[0067] S34, set the loss function by comparing the concept activation value and the true labeled value of the concept bottleneck layer in a step-by-step manner.

[0068] Loss functions (such as masked binary cross-entropy loss) can compare the predicted concept activations with the true labeled values ​​at each time step, thereby forcing the concept bottleneck layer model to learn to activate the correct concepts at specific times.

[0069] The supervision method provided in this application offers two fundamental advantages. First, it constrains the representation learning of the concept bottleneck layer model, requiring it to focus its attention on acoustic-semantic features truly relevant to human-defined concepts. This suppresses the possibility of using false features such as background noise and speaker timbre for judgment, greatly improving the model's generalization ability and fairness. Second, it inherently provides interpretability, as any final decision of the concept bottleneck layer model can be traced back to the activation status of these intermediate concept layers. Furthermore, employing a time-stamped, concept-level explicit supervised training scheme can improve generalization and accuracy.

[0070] In some embodiments, S33 in the above embodiments, the concept analysis of the concept activation sequence and the output of the concept analysis report, includes at least one of the following S331 and S332.

[0071] S331, the concept with the highest overall activation level or exceeding the first threshold is determined as the main concept of the audio signal.

[0072] Here, the sum of the activation levels of each concept across all time steps corresponding to the audio signal can be used as the overall activation level. If the overall activation level of a particular concept is the highest or exceeds a first threshold, then that concept is identified as the dominant concept in the audio signal.

[0073] S332, when the activation level of a concept corresponding to multiple consecutive time steps exceeds a second threshold, the start and end time steps of the concept are determined based on the multiple consecutive time steps.

[0074] Here, if the activation of a concept corresponding to multiple consecutive time steps exceeds the second threshold, the concept can be identified as the main concept of this part of the time steps, and the start time step and end time step can be identified as the start and end time steps of the concept.

[0075] This application provides two specific implementation methods for outputting conceptual analysis reports. These embodiments contribute to building a complete system for practical applications. It not only improves the overall accuracy and stability of audio content analysis and detection, but more importantly, it possesses real-time streaming processing capabilities, enabling online analysis and immediate alerts for continuous audio streams. This marks a shift from "post-event analysis" to "real-time intervention," providing a directly deployable and comprehensive solution (detection, location, interpretation, and alerting) for content security application scenarios.

[0076] Figure 4 This is a schematic diagram of the structure of an audio content analysis device provided in an embodiment of this application.

[0077] like Figure 4 As shown, the audio content analysis device 4 includes: The first obtaining module 41 is used to process the audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states of one or more time steps; The second mapping module 42 is used to map the hidden state sequence into an interpretable concept activation sequence, the concept activation sequence including the concept activation vectors of the one or more time steps; The third output module 43 is used to perform concept analysis on the concept activation sequence and output a concept analysis report. The concept analysis report includes at least one of the following: the concept analysis result of the audio signal, and the concept analysis result of one or more time steps in the audio signal.

[0078] Another embodiment of the present invention discloses an audio content analysis device 4. This embodiment is based on the above... Figure 4 Based on the corresponding embodiment, the first obtaining module 41 is used for: The audio signal is preprocessed to obtain a preprocessed audio signal. The preprocessing includes at least one of the following: resampling, windowing, and frame splitting. The preprocessed audio signal is input into the audio big model to obtain the hidden state sequence output by the audio big model.

[0079] Another embodiment of the present invention discloses an audio content analysis device 4. This embodiment is based on the above... Figure 4 Based on the corresponding embodiment, the second mapping module 42 is used for: The hidden state sequence is mapped to the concept space based on the concept bottleneck layer, and the mapped text concepts are aligned with the time steps of the audio signal to generate an aligned and interpretable concept activation sequence.

[0080] Another embodiment of the present invention discloses an audio content analysis device 4. This embodiment is based on the above... Figure 4 Based on the corresponding embodiment, the conceptual bottleneck layer is a lightweight neural network, which includes any one of the following: a fully connected layer, a multilayer perceptron, or a convolutional neural network.

[0081] Another embodiment of the present invention discloses an audio content analysis device 4. This embodiment is based on the above... Figure 4 Based on the corresponding embodiment, the audio content analysis device 4 further includes: The fourth module is used to establish the correspondence between labeled data, time steps, and concepts during the training of the concept bottleneck layer. The fifth setting module is used to set the loss function by comparing the concept activation value and the true labeled value of the concept bottleneck layer in a step-by-step manner.

[0082] Another embodiment of the present invention discloses an audio content analysis device 4. This embodiment is based on the above... Figure 4 Based on the corresponding embodiment, the third output module 43 is used for at least one of the following: The concept with the highest overall activation level or exceeding the first threshold is identified as the main concept of the audio signal; When the activation level of a concept exceeds a second threshold at multiple consecutive time steps, the start and end time steps of the concept are determined based on the multiple consecutive time steps.

[0083] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0084] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0085] This application also provides an electronic device, such as... Figure 5 As shown, the electronic device 5 includes: at least one processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 executes the computer program 52 to implement the steps in any of the above-described method embodiments.

[0086] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0087] This application provides a computer program product, which includes a computer program that, when run, causes the steps described in the various method embodiments above to be performed.

[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0089] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0090] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0091] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0092] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0093] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for audio content analysis, characterized in that, include: The audio signal is processed to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; The hidden state sequence is mapped to an interpretable concept activation sequence, the concept activation sequence including the concept activation vectors of the one or more time steps; Perform concept analysis on the concept activation sequence and output a concept analysis report, which includes at least one of the following: the concept analysis results of the audio signal, and the concept analysis results of one or more time steps in the audio signal.

2. The method as described in claim 1, characterized in that, The process of processing the audio signal to obtain the hidden state sequence of the audio signal includes: The audio signal is preprocessed to obtain a preprocessed audio signal. The preprocessing includes at least one of the following: resampling, windowing, and frame splitting. The preprocessed audio signal is input into the audio big model to obtain the hidden state sequence output by the audio big model.

3. The method as described in claim 1, characterized in that, The step of mapping the hidden state sequence to an interpretable concept activation sequence includes: The hidden state sequence is mapped to the concept space based on the concept bottleneck layer, and the mapped text concepts are aligned with the time steps of the audio signal to generate an aligned and interpretable concept activation sequence.

4. The method as described in claim 3, characterized in that, The conceptual bottleneck layer is a lightweight neural network, which includes any of the following: a fully connected layer, a multilayer perceptron, or a convolutional neural network.

5. The method as described in claim 3 or 4, characterized in that, The method further includes: During the training of the concept bottleneck layer, a correspondence is established between labeled data, time steps, and concepts; The loss function is set by comparing the concept activation value and the true labeled value of the concept bottleneck layer step by step.

6. The method as described in claim 1, characterized in that, The concept analysis of the concept activation sequence, and the output of the concept analysis report, includes at least one of the following: The concept with the highest overall activation level or exceeding the first threshold is identified as the main concept of the audio signal; When the activation level of a concept exceeds a second threshold at multiple consecutive time steps, the start and end time steps of the concept are determined based on the multiple consecutive time steps.

7. An apparatus for audio content analysis, characterized in that, include: The first obtaining module is used to process the audio signal to obtain a hidden state sequence of the audio signal, the hidden state sequence including hidden states at one or more time steps; The second mapping module is used to map the hidden state sequence into an interpretable concept activation sequence, wherein the concept activation sequence includes the concept activation vectors of the one or more time steps; The third output module is used to perform concept analysis on the concept activation sequence and output a concept analysis report. The concept analysis report includes at least one of the following: the concept analysis result of the audio signal, and the concept analysis result of one or more time steps in the audio signal.

8. An electronic device, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device performs the method as described in any one of claims 1-6.

9. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1-6 to be performed.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.