Audio description generation method and device, electronic equipment, storage medium and product

Through expert network models and metadata processing, cross-domain audio descriptions are identified and generated, which solves the problem of feature adaptation in multi-domain audio scenarios and improves the accuracy and comprehensiveness of audio descriptions.

CN120636459APending Publication Date: 2025-09-12BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510796194.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies are unable to simultaneously adapt to the features of different audio types in multi-domain audio scenarios, resulting in loss or misjudgment of key information and reducing the accuracy of audio description.

Method used

An expert network model is used to identify audio types and generate audio descriptions based on the audio types. The expert network model is used to identify and predict features of different audio types, combined with confidence judgment and metadata to handle conflicts, and the thinking chain parsing capability is used to integrate audio features to generate descriptions.

Benefits of technology

It achieves cross-domain audio understanding, improves the accuracy and comprehensiveness of audio description, reduces misjudgment and information loss, and generates multi-angle and coherent audio description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636459A_ABST
    Figure CN120636459A_ABST
Patent Text Reader

Abstract

The invention relates to an audio description generation method and device, electronic equipment, a storage medium and a product. The audio description generation method comprises the following steps: acquiring an audio, and identifying an audio type to which the audio belongs so as to obtain at least one audio type; for the at least one audio type, determining audio characteristics of the audio according to the audio type so as to obtain audio characteristics corresponding to the at least one audio type; based on the audio characteristics corresponding to the at least one audio type, audio description of the audio is generated, and the audio description comprises description representing the audio characteristics. According to the invention, cross-domain understanding of the audio can be realized, and multi-domain audio description can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio types, and in particular to a method, device, electronic device, storage medium, and product for generating audio description. Background Art

[0002] In recent years, the field of audio, particularly in audio content understanding and description generation, has experienced a shift from small-scale, precise annotation to large-scale, weakly supervised annotation. Early research primarily relied on manually annotated datasets, which, while ensuring annotation quality, was limited in data size due to high costs. Current technological trends are achieving a breakthrough in data scale through automated annotation processes (such as speech recognition transcription) and the use of large language models (LLMs) to assist in annotation.

[0003] With the advancement of large language models and multimodal technologies, this field has expanded from traditional sound event detection (SED) to audio captioning tasks, requiring the model to output natural language descriptions to comprehensively represent the semantics, scenes, and emotions of the audio. However, in the scenario of multi-domain audio, there are significant differences between different fields. For example, in multi-domain audio scenarios corresponding to different audio types such as speech type and music type, related technologies are still unable to adapt to the characteristics of multi-domain audio at the same time, resulting in the loss or misjudgment of key information in multi-domain audio, which in turn leads to a decrease in the accuracy of the description. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a method, device, electronic device, storage medium and product for generating audio description.

[0005] According to a first aspect of some embodiments of the present disclosure, a method for generating an audio description is provided, comprising acquiring audio and identifying the audio type to which the audio belongs to obtain at least one audio type; for the at least one audio type, determining the audio characteristics of the audio according to the audio type to obtain audio characteristics corresponding to each of the at least one audio type; and generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type, wherein the audio description includes a description characterizing the audio characteristics.

[0006] In one implementation, one audio type corresponds to at least one audio characteristic.

[0007] In one embodiment, the at least one audio characteristic is predicted based on an expert network model corresponding to the audio type, and the expert network model includes at least one sub-model, wherein one audio type corresponds to one expert network model, and one sub-model is used to determine one audio characteristic.

[0008] In one embodiment, the method further includes: respectively determining the confidence of the audio characteristics corresponding to the at least one audio type; in response to conflicting audio characteristics, selecting the audio characteristics with the highest confidence among the conflicting audio characteristics as the audio characteristics used in the audio description.

[0009] In one embodiment, the method further includes: obtaining metadata of the audio, the metadata including information associated with an audio description of the audio; and generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type includes: generating an audio description of the audio based on the metadata and the audio characteristics corresponding to each of the at least one audio type.

[0010] In one embodiment, generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type includes: inputting the audio characteristics corresponding to each of the at least one audio type into a model, the model having a thought chain parsing capability, the thought chain being used to describe the display reasoning process of generating the audio description; determining the association relationship between the audio characteristics corresponding to each of the at least one audio type based on the thought chain parsing capability; displaying the association relationship, and determining a reliable association relationship, and using the description of the audio characteristics corresponding to the reliable association relationship as the audio description of the audio.

[0011] In one embodiment, the audio type includes at least one of the following: a voice type; a music type; an environment type; a combination type, wherein the combination type includes a combination of at least two of the voice type, the music type, and the environment type.

[0012] In one embodiment, in response to the audio type being the speech type, the audio characteristics include at least one of the following: language, speech recognition, emotion, age, gender, accent, and the paragraph corresponding to the speaker; in response to the audio type being the music type, the audio characteristics include at least one of the following: instrument, genre, and emotion; in response to the audio type being the ambient sound type, the audio characteristics include at least one of the following: quality, reverberation level, and sound intensity.

[0013] According to a second aspect of some embodiments of the present disclosure, a device for generating an audio description is provided, comprising a processing unit for acquiring audio and identifying the audio type to which the audio belongs, so as to obtain at least one audio type; a determination unit for determining the audio characteristics of the audio according to the audio type, so as to obtain the audio characteristics corresponding to each of the at least one audio type; and a generation unit for generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type, wherein the audio description includes a description characterizing the audio characteristics.

[0014] In one implementation, one audio type corresponds to at least one audio characteristic.

[0015] In one embodiment, the at least one audio characteristic is predicted based on an expert network model corresponding to the audio type, and the expert network model includes at least one sub-model, wherein one audio type corresponds to one expert network model, and one sub-model is used to determine one audio characteristic.

[0016] In one embodiment, the determination unit is further used to: respectively determine the confidence of the audio characteristics corresponding to each of the at least one audio type; in response to conflicting audio characteristics, select the audio characteristics with the highest confidence among the conflicting audio characteristics as the audio characteristics used in the audio description.

[0017] In one embodiment, the processing unit is further used to: obtain metadata of the audio, the metadata including information associated with the audio description of the audio; based on the audio characteristics corresponding to each of the at least one audio type, the generation unit generates the audio description of the audio in the following manner: based on the metadata and the audio characteristics corresponding to each of the at least one audio type, generate the audio description of the audio.

[0018] In one embodiment, based on the audio characteristics corresponding to each of the at least one audio type, the generation unit generates an audio description of the audio in the following manner: the audio characteristics corresponding to each of the at least one audio type are input into a model, the model having a thought chain parsing capability, and the thought chain is used to describe the display reasoning process of generating the audio description; based on the thought chain parsing capability, the association relationship between the audio characteristics corresponding to each of the at least one audio type is determined; the association relationship is displayed, and a reliable association relationship is determined, and the description of the audio characteristics corresponding to the reliable association relationship is used as the audio description of the audio.

[0019] In one embodiment, the audio type includes at least one of the following: a voice type; a music type; an ambient sound type; a combination type, wherein the combination type includes a combination of at least two of the voice type, the music type, and the ambient sound type.

[0020] In one embodiment, in response to the audio type being the speech type, the audio characteristics include at least one of the following: language, speech recognition, emotion, age, gender, accent, and the paragraph corresponding to the speaker; in response to the audio type being the music type, the audio characteristics include at least one of the following: instrument, genre, and emotion; in response to the audio type being the ambient sound type, the audio characteristics include at least one of the following: quality, reverberation level, and sound intensity.

[0021] According to a third aspect of some embodiments of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute the method for generating audio description described in the first aspect or any one of the implementations of the first aspect.

[0022] According to a fourth aspect of some embodiments of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor, the processor is enabled to execute the method for generating audio description described in the first aspect or any one of the embodiments of the first aspect.

[0023] According to a fifth aspect of some embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method for generating audio description as described in the first aspect or any one of the embodiments of the first aspect.

[0024] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: audio is acquired and the audio type to which the audio belongs is identified, thereby obtaining at least one audio type of the audio. For at least one audio type, the audio characteristics of the audio are determined according to the audio type, thereby further clarifying the audio characteristics of each audio type on the basis of clarifying the audio type of the audio, reducing the loss or misjudgment of the recognition of audio characteristics, and improving the ability to understand audio with different audio characteristics. Based on the audio characteristics corresponding to at least one audio type, an audio description is generated, and the audio description includes a description that characterizes the audio characteristics, thereby obtaining audio descriptions corresponding to multiple audio types, realizing the fusion of multi-domain audio descriptions, and realizing cross-domain audio understanding. It is also possible to obtain multi-angle audio descriptions and improve the accuracy of the audio description.

[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0027] Figure 1 The present invention is a flowchart of a method for generating audio description according to some embodiments of the present disclosure.

[0028] Figure 2 The figure is a flowchart showing a method for resolving conflicts according to some embodiments of the present disclosure.

[0029] Figure 3 The present invention is a flowchart of a method for generating audio description according to some embodiments of the present disclosure.

[0030] Figure 4 The present invention is a flowchart of a method for generating audio description according to some embodiments of the present disclosure.

[0031] Figure 5 This is a schematic diagram of a method for generating audio description according to some embodiments of the present disclosure.

[0032] Figure 6 FIG. 4 is a schematic diagram showing an audio description according to some embodiments of the present disclosure.

[0033] Figure 7 The present invention is a block diagram of a device for generating audio description according to some embodiments of the present disclosure.

[0034] Figure 8 It is a block diagram of an electronic device for generating audio description according to some embodiments of the present disclosure.

[0035] Figure 9 It is a block diagram of a device for generating audio description according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0036] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.

[0037] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0038] In the audio field, particularly in audio content understanding and description generation, there has been a shift from small-scale, precise annotation to large-scale, weakly supervised systems. To further analyze audio content understanding and description generation, the following comparison of three representative systems: a small-scale, precise annotation system, a medium-scale, LLM-assisted annotation system, and a large-scale, automated construction system, discusses their design differences and areas for improvement.

[0039] A small-scale, precise annotation system uses an encoder-decoder architecture based on collected audio data to encode audio into feature vectors. This system then generates audio descriptions using a long short-term memory (LSTM) decoder. This system then uses manual annotation to construct tens of thousands of high-quality audio-to-audio description pairs. While the annotation quality is high, its limited scale makes it difficult to handle complex, long audio scenarios. It lacks a deep understanding of audio content and focuses primarily on describing sounds and basic events, making it unable to handle complex audio content across multiple languages ​​and domains.

[0040] Medium-scale LLM-assisted annotation system: Through the question-answering model-assisted annotation strategy, the original audio descriptions from sources such as online databases (FreeSound) are converted into standardized audio descriptions, and about hundreds of thousands of audio-audio description pairs are constructed. A contrastive learning framework is used to map audio and text into a shared semantic space. Although it is large in scale, it has the following problems: each audio usually corresponds to a single description, and it lacks the ability to integrate multi-source information. Although weakly supervised annotations are processed by LLM, there is still noise and inconsistency; the auxiliary role of relevant metadata in audio understanding is not considered. When processing mixed audio (such as speech + music), there is a lack of specialized processing strategies.

[0041] Large-scale automated construction system: collects large-scale audio-text datasets, including about hundreds of thousands of audio-audio description pairs, covering a variety of types such as music, speech, and environmental sounds. The system uses automated network crawling and filtering strategies to build datasets, and uses audio-audio description similarity to screen relevant pairs. Although larger in scale than small-scale precision annotation systems, it has the following problems: the quality of automatically captured annotations varies, and there are noise and inaccurate descriptions. The system does not distinguish between different types of audio content and uses a unified encoding-decoding architecture, making it difficult to capture the specific features of different audio types such as speech and music. The system lacks a mechanism for fusing audio content with its related metadata, as well as the ability to handle possible conflicts among multi-source information, resulting in poor performance in complex, multi-source audio scenarios.

[0042] Related technologies typically use a single model structure to obtain audio descriptions corresponding to audio. However, this single model structure lacks the specialized processing capabilities to effectively integrate different audio types, resulting in insufficient unified understanding of cross-domain audio and an inability to process them.

[0043] In view of this, the present disclosure proposes a method for generating audio description, which is applied as an automatic annotation tool for large-scale audio datasets to generate audio description tags.

[0044] Among them, the generation of audio description is to create audio and audio description and form a training data set, which can be used as training data to train the audio base large model.

[0045] Figure 1 is a flowchart of a method for generating audio description according to some embodiments of the present disclosure. Figure 1 As shown, the following steps are included.

[0046] In step S11, audio is acquired, and the audio type to which the audio belongs is identified to obtain at least one audio type.

[0047] In the embodiments of the present disclosure, an audio type refers to a specific category or style of audio. For example, audio can be classified into different audio types based on different types of sounds in the audio. Different audio types correspond to different audio.

[0048] In the embodiment of the present disclosure, the audio type can be classified based on the audio classification model to determine at least one audio type of the audio. It is understandable that different audio types correspond to the same audio.

[0049] In some embodiments, the audio classification model may be an audio classification model based on consistent ensemble distillation (CED).

[0050] In some embodiments, the CED-based audio classification model can identify the audio type of the audio and obtain the audio type to which the audio belongs and the confidence level of the audio type.

[0051] In some embodiments, the audio type to which the audio belongs may be any audio type, or may be audio obtained by combining any two or more different audio types.

[0052] In step S12, for at least one audio type, audio characteristics of the audio are determined according to the audio type to obtain audio characteristics corresponding to the at least one audio type.

[0053] In the embodiments of the present disclosure, audio characteristics can be understood as characteristics extracted from and associated with the audio. Each audio type corresponds to its own audio feature, and corresponding characteristics can be extracted for audio of at least one audio type to obtain the corresponding audio characteristics.

[0054] In step S13, an audio description of the audio is generated based on the audio characteristics corresponding to the at least one audio type.

[0055] In the embodiment of the present disclosure, the audio description is usually in text form.

[0056] In the disclosed embodiments, the audio characteristics corresponding to at least one audio type are integrated to generate an audio description associated with the audio. It is understood that the audio description includes a description characterizing the audio characteristics. That is, if the audio contains multiple audio characteristics, the descriptions of the multiple audio characteristics need to be comprehensively considered to obtain the audio description of the audio.

[0057] In the embodiments of the present disclosure, considering that audio has a variety of different audio types, the present disclosure performs cross-domain recognition of audio of at least one audio type, enabling cross-domain understanding of audio. Furthermore, based on the audio characteristics corresponding to the at least one audio type, the characteristics of the audio in different domains and from different perspectives can be obtained, generating a multi-angle and cross-domain audio description.

[0058] In the embodiment of the present disclosure, the audio includes at least one audio type, one audio type corresponds to at least one audio characteristic, and each audio characteristic has a corresponding description.

[0059] In the disclosed embodiments, the audio type of the audio is identified, enabling different processing of the same audio of different types, thereby improving the accuracy of audio description generation. Furthermore, the audio characteristics corresponding to at least one audio type are predicted, enabling refined audio processing and ultimately generating a cross-domain and refined audio description.

[0060] In the disclosed embodiments, each audio type corresponds to an expert network model. Each expert network model includes at least one sub-model. A sub-model can be used to determine a single audio characteristic or multiple audio characteristics. In other words, at least one audio characteristic is derived based on at least one sub-model within the expert network model corresponding to the audio type, thereby enabling cross-domain understanding and processing of audio types.

[0061] In some embodiments, the audio type may include at least one of a speech type (Speech), a music type (Music), an ambient sound type (Sound), and a combination type. The combination type includes a combination of at least two of the speech type, the music type, and the ambient sound type.

[0062] If the audio is of the speech type, it can be understood as including the sound of living things speaking. For example, it can be the sound of human speech, such as a single person speaking, a group conversation, a speech, or a recitation. The music type can include the sound produced by musical instruments or human voices. The ambient sound type can be sounds produced in natural or man-made environments. This type of audio is generally not produced by human speech or music, such as the sound of wind, rain, or subway.

[0063] In some embodiments, the audio may include any one of a speech type, a music type, and an ambient sound type, or a combination of at least two of the two. For example, if a segment of audio includes the sound of rain, it can be determined through identification that the audio type to which the audio belongs is an ambient sound type. For example, if a segment of audio includes the sound of rain and the sound of human speech, it can be determined through identification that the audio type to which the audio belongs includes an ambient sound type and a speech type. For example, if a segment of audio includes the sound of rain, the sound of human speech, and a piece of music, it can be considered that the audio type to which the audio belongs includes an ambient sound type, a speech type, and a music type. The above embodiments are merely illustrative, and the present disclosure does not limit the audio type.

[0064] In an embodiment of the present disclosure, in response to the audio type being a speech type, the audio characteristics may include at least one of the following: language identification (Language Id), emotion (Emotion), speech recognition (Automatic Speech Recognition), age (Age), gender (Gender), accent (Accent), and the paragraph corresponding to the speaker (Speaker Diazation).

[0065] Language refers to the language used in the audio, such as English or Chinese. Emotion refers to the feelings expressed in the audio, and emotion recognition can be performed by analyzing the audio's pitch, speaking rate, and other factors. Speech recognition involves converting speech in the audio into text. Age refers to the age of the speaker in the audio, allowing for identification of whether the speaker is a child, young adult, or middle-aged. Gender refers to the gender of the speaker in the audio, such as male or female. Accent refers to the speaker's accent, such as a local dialect or accent. Speaker segmentation involves separating different speakers in the audio.

[0066] In some embodiments, at least one audio characteristic corresponding to the speech type is inferred through the audio of the speech type, thereby obtaining a prediction result for each audio characteristic, which can provide multi-angle and comprehensive speech information.

[0067] In the embodiment of the present disclosure, in response to the audio type being a music type, the audio characteristic includes at least one of the following: instrument, genre, and emotion.

[0068] Instrument refers to the type of instrument used in the audio, such as guitar or piano. This can be identified by analyzing the audio's spectral characteristics, time domain characteristics, or a combination of both. Genre categorizes audio by style, such as pop, rock, or classical. This can be determined by musical features such as rhythm and melody. Emotion refers to the emotion expressed in the audio, such as happiness or sadness. This can be determined by analyzing audio features such as pitch and rhythm.

[0069] In some embodiments, the audio type is music type audio, which can be pure music, that is, audio including a vocal part. It can also be audio including a vocal part and a non-vocal part. The music type audio is input into the expert network model corresponding to the music type, and the vocal part and the non-vocal part in the audio are separated based on the sub-model corresponding to Music Separation. In the expert network model corresponding to the music type, the audio characteristics of the non-vocal part are predicted, and the audio of the vocal part is input into the expert network model corresponding to the speech type, and the different audio characteristics of the vocal part are predicted based on multiple sub-models under the expert network model corresponding to the speech type.

[0070] In some embodiments, audio characteristics of music-type audio are predicted, thereby enabling professional analysis of music content.

[0071] In the embodiment of the present disclosure, in response to the audio type being the ambient sound type, the audio characteristic includes at least one of the following: quality, reverberation, and sound intensity.

[0072] Quality refers to the quality of the audio, such as noise and distortion. Reverberation refers to the reverberation characteristics of the audio, such as reverberation time or reverberation intensity, which can be determined through signal processing techniques. Sound intensity indicates the intensity characteristics of the audio, such as average intensity.

[0073] In some embodiments, audio characteristics of ambient sound type audio are predicted, audio quality features may be extracted and evaluated, and processed according to a preset time window to provide information related to sound quality.

[0074] In the embodiment of the present disclosure, different audio types correspond to different audio characteristics, and each audio characteristic can determine a prediction result of its audio characteristic through its own sub-model.

[0075] In an embodiment of the present disclosure, if the audio type includes a speech type, the audio is input into the expert network model corresponding to the speech type, and the prediction results of each audio characteristic are determined based on the sub-network under the expert network model. If the audio type includes a speech type and a music type, the audio is input into the expert network model corresponding to the speech type and the expert network model corresponding to the music type, and the audio characteristics are predicted based on the sub-models corresponding to the respective expert network models. When the audio of the present disclosure has other audio types, the audio is input into the expert network model of the corresponding audio type respectively, which will not be elaborated here. It can be understood that the audio characteristics corresponding to each type of audio are different. Based on the expert network type of each audio type, the detailed attributes of the audio such as speech, music, and sound quality are refined, which can improve the accuracy of the audio description and reduce the problem of excessive noise. It can also enrich the audio description and enhance the comprehensive understanding of audio in multiple fields.

[0076] In some embodiments, a pre-trained audio description model can be provided, and an overall audio description of audio types other than speech types can be provided based on the pre-trained audio description model. The overall audio description can be used as supplementary information to help enrich the audio description.

[0077] Because there are multiple audio features, conflicts may arise when obtaining prediction results for these features. This can be resolved by determining the confidence level corresponding to each audio feature and determining whether to retain the audio feature based on the relationship between the confidence levels.

[0078] Therefore, the embodiment of the present disclosure provides a method for resolving conflicts based on confidence. Figure 2 As shown, Figure 2 The flowchart of a conflict resolution method according to some embodiments of the present disclosure includes the following steps.

[0079] In step S21 , the confidence level of the audio characteristics corresponding to at least one audio type is determined respectively.

[0080] In the disclosed embodiment, each sub-model outputs a prediction result corresponding to the audio feature, along with a confidence score for the prediction result. Confidence is used to measure the degree of certainty in the prediction result. For example, if the audio is speech-type, the language of the audio is identified, and the result is: English (0.96). Here, English is the language corresponding to the audio, and 0.96 is the confidence score that the audio is English.

[0081] In step S22 , in response to the existence of conflicting audio features, an audio feature with the highest confidence level is selected as the audio feature used in the audio description among the conflicting audio features.

[0082] In the disclosed embodiments, if two or more of the at least one audio feature conflict, the prediction result of the retained audio feature may be determined based on their respective confidence levels. For example, the audio feature with the highest confidence level may be used as the audio feature for audio description, and the prediction result corresponding to that audio feature may be retained.

[0083] For example, an audio clip contains both music and speech. The music part is predicted as "popular style (0.9), relaxed (0.92)", and the speech part is predicted as "male (0.98), youth (0.97), angry emotion (0.9)". Then there is a conflict between relaxed and angry emotions. Since the confidence of relaxed emotion is higher than that of angry emotion, relaxed emotion can be retained and used as the result in audio description.

[0084] In the disclosed embodiment, the confidence level of each audio feature prediction result is outputted, so that a decision can be made when there is a conflict in audio features, thereby reducing the logical incoherence in the audio description caused by the conflict in audio features.

[0085] In the embodiment of the present disclosure, in the process of generating audio description for audio, in addition to the audio characteristics corresponding to each audio type in the audio, the audio metadata can also be obtained to jointly obtain the audio description, so that the system can understand a wide range of context information. And based on the metadata, the problem of audio characteristic conflicts can be further solved. Figure 3 As shown, Figure 3 The flowchart of a method for generating an audio description according to some embodiments of the present disclosure includes the following steps.

[0086] In step S31 , metadata of the audio is acquired.

[0087] In the embodiment of the present disclosure, metadata includes information associated with the audio description of the audio. The metadata can be information such as the audio title, description, etc.

[0088] In some embodiments, the description includes but is not limited to audio content, author, genre, etc. The metadata of the audio can be manually input, automatically extracted, or obtained through an interface call.

[0089] In step S32, an audio description of the audio is generated based on the metadata and the audio characteristics corresponding to the at least one audio type.

[0090] In the embodiment of the present disclosure, considering that at least one audio type corresponds to audio characteristics and metadata, an audio description including the audio characteristics and metadata can be obtained.

[0091] In the disclosed embodiments, metadata and audio characteristics are integrated to create a complete audio description. This can be used to make the audio description more relevant and accurate, even in the absence of clear audio cues, through the aid of metadata, and to provide a richer, more contextually informative audio description. Furthermore, when metadata conflicts with the audio characteristics corresponding to at least one audio type, the metadata can be used as a reference tag, retaining the audio characteristics associated with the metadata and discarding the others, thus resolving the conflict.

[0092] The following embodiment of the present disclosure further illustrates the method for generating audio description. Figure 4 is a flowchart of a method for generating audio description according to some embodiments of the present disclosure. Figure 4 As shown, the following steps are included.

[0093] In step S41, the audio characteristics corresponding to at least one audio type are input into the model, and the model has the ability to analyze thought chains.

[0094] In the disclosed embodiments, thought chains are used to describe the explicit reasoning process of audio description generation, analyzing the relationships and potential contradictions between various information sources. The information sources can be metadata or the predicted results of various audio characteristics.

[0095] In step S42, the association relationship between the audio characteristics corresponding to at least one audio type is determined based on the thought chain analysis capability.

[0096] In the disclosed embodiments, various information sources can be input into a large language model, and based on the chain-of-thought analysis capability, the relationships between the information sources can be determined by displaying the reasoning process. For example, this can be the correlation between the information sources, the contradictions between the information sources, or the dependencies between the information sources.

[0097] In step S43, the association relationship is displayed, and a reliable association relationship is determined, and the description of the audio characteristics corresponding to the reliable association relationship is used as the audio description of the audio.

[0098] In the disclosed embodiments, a reliable association relationship can be understood as information sources that are related and non-contradictory. By analyzing the association relationships between information sources and determining which ones to retain and which to discard, the remaining information sources are considered to have reliable associations. The descriptions corresponding to each information source are integrated and output, resulting in a coherent, natural, and accurate audio description.

[0099] In the disclosed embodiment, the system utilizes the ability to analyze thought chains to transparently reason out conflicts and relationships between different information sources. In the case of incomplete information or noise, reliable judgments can be made through an explicit reasoning process. For example, when the generated audio description is inconsistent with the speech recognition result, the system can determine which one is more reliable through reasoning. Moreover, the audio description not only covers the key events in the audio, but also describes them in a natural and coherent manner, thereby having better language fluency and narrative structure, avoiding the mechanical expression of simply listing events when generating the audio description.

[0100] In the embodiment of the present disclosure, Figure 5 The method for generating audio description is described. Figure 5 FIG. 1 is a schematic diagram illustrating a method for generating an audio description according to some embodiments of the present disclosure. Figure 5 As shown, the audio and audio-related metadata (title, description) are obtained. The metadata can be used as a supplement to domain knowledge to enhance the system's contextual understanding of the audio content.

[0101] In the embodiment of the present disclosure, in the multi-expert network model, Step 1 can be executed to classify the audio into three categories: speech, music, and ambient sound types. If the audio is determined to be of the speech type, it is input into the expert network model corresponding to Step 2, that is, the speech processing pipeline. If the audio is determined to be of the music type, it is input into the expert network model corresponding to Step 3, that is, the music processing pipeline. If the audio is determined to be of the ambient sound type, it is input into the expert network model corresponding to Step 4, that is, the audio quality analysis pipeline. It can be understood that the audio is input into the corresponding expert network model according to the audio type.

[0102] In the embodiment of the present disclosure, in the speech processing pipeline, the speech is classified in many aspects, including language, speech recognition, emotion, age, gender, accent and at least one of the paragraphs corresponding to the speaker, so as to provide multi-angle and comprehensive speech information. In the music processing pipeline, the audio is separated to distinguish between the vocal part and the non-vocal part. The non-vocal part is judged by the instrument, genre and emotion. The vocal part is transferred to the speech processing pipeline of Step 2 to achieve professional analysis of the music content. In the audio quality analysis pipeline, the audio quality features, reverberation level and sound intensity are extracted and evaluated, processed in a 2-second window, and sound quality related information is provided.

[0103] In some embodiments, based on the audio description pre-training model, an overall description of the audio in addition to the voice information can be provided. Although the overall description of the audio may not be accurate enough, it can help enrich the audio description as a supplementary information source.

[0104] In the embodiment of the present disclosure, based on the above-mentioned different audio characteristics, that is, the audio characteristics in each processing pipeline are predicted to obtain prediction results and confidence levels. That is, the multi-expert network model inputs audio and outputs prediction results and confidence levels. The title, description, prediction results, and confidence levels are used as inputs to the large language model, and the Chain-of-Thought parsing capability of the large language model is used to analyze the relationship between each information source and possible contradictions. The reliability of the information is judged through an explicit reasoning process to determine which information should be retained and which should be discarded. All reliable information is integrated to generate a coherent and accurate audio description to ensure the accuracy, contextual relevance, and fluency of the audio description.

[0105] For example, combined Figure 6 An example is given to illustrate the audio description obtained based on the multi-expert network model. Figure 6 FIG. 1 is a schematic diagram showing an audio description according to some embodiments of the present disclosure. Figure 6 As shown, audio is acquired and its audio type is identified as speech or music, meaning it contains both speech and music. Metadata related to the audio is acquired, which may include a title such as "Terminal Game Review: Can it run large-scale games smoothly?" and a description such as "Terminal: Latest Budget-Level Performance Model A." Expert network model analysis is then performed, using multiple sub-models within the corresponding expert network model to predict the audio characteristics of speech and music. Recognition by the respective expert network models can include speech recognition results such as "Discussing the performance comparison between Terminal A and Terminal B...", language identification: "English (confidence 0.96)", speaker analysis: "Speaker 00 (male, hoarse voice)", background sound analysis: "Electronic / experimental music (synthesizer, piano, 88 beats per minute (BPM), C minor) + speech", audio quality assessment: "Deep Noise Suppression Mean Opinion Score (DNSMOS): 2.77, Non-Intrusive Speech Quality Assessment Mean Opinion Score (NISQAMOS): 2.88 (medium)", and music sentiment analysis: "Interesting and cheerful (high pleasure, medium excitement)". Based on the above metadata, prediction results, and confidence, an audio description can be generated: "A male voice evaluating Terminal A, with a lively electronic rhythm as the background, combining technical commentary with an experimental synthesizer melody to create an urban atmosphere." It can be understood that the audio description includes the description of the above-mentioned corresponding audio characteristics, thereby obtaining a multi-domain and refined audio description.

[0106] In the disclosed embodiments, the multi-expert network model enables the system to simultaneously process and deeply understand speech content, music features, and sound. It excels in mixed audio scenes containing multiple sound sources, for example, being able to simultaneously identify speech content, background music genre, and ambient sound events. Compared to single-model architectures, it provides a more comprehensive and accurate understanding of complex audio scenes. By integrating metadata, the system can understand a wider range of contextual information, particularly in understanding audio content in specialized fields. In the absence of clear audio cues, metadata assists in making descriptions more relevant and accurate. For example, it can associate professional terms mentioned in video titles with the content heard in the audio, generating a more context-rich audio description. Leveraging the thought chain parsing capabilities of a large language model, the system can transparently reason about conflicts and relationships between different information sources. In situations where information is incomplete or noisy, more reliable judgments can be made through explicit reasoning. For example, when the content generated by the audio description model is inconsistent with the speech recognition results, the system can determine which is more reliable through reasoning. This allows the generated audio description to not only cover key events in the audio but also express them in a natural and coherent manner, avoiding the mechanical presentation of a simple listing of events. Compared with conventional datasets, it can better integrate multiple audio elements, such as speech content, music emotions, and environmental sound events.

[0107] Furthermore, the audio-to-audio description pairs generated in the disclosed embodiments can be used to train the Audio Base macro model, providing it with comprehensive and context-aware training data. Furthermore, the audio model trained based on the audio-to-audio description pairs can provide intelligent analysis and semantic search capabilities, enabling visually impaired users to understand the speech, music, and ambient sound details in media content. By generating accurate and detailed audio descriptions, it not only supports content classification, retrieval, and recommendation, but also improves the accessibility of media content, meeting the requirements for barrier-free use.

[0108] Based on the same concept, the embodiment of the present disclosure further provides an audio description generating device 100 .

[0109] It is understood that, in order to achieve the above functions, the audio description generation device 100 provided in the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function. In combination with the units and algorithm steps of the various examples disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware-driven hardware manner or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.

[0110] Figure 7 FIG1 is a block diagram of a device for generating audio description according to some embodiments of the present disclosure. Figure 7 The device includes a processing unit 101, a determining unit 102 and a generating unit 103.

[0111] The processing unit 101 is configured to acquire audio and identify the audio type to which the audio belongs, so as to obtain at least one audio type.

[0112] The determining unit 102 is configured to determine audio characteristics of at least one audio type according to the audio type, so as to obtain audio characteristics corresponding to the at least one audio type.

[0113] The generating unit 103 is configured to generate an audio description of the audio based on the audio characteristics corresponding to at least one audio type, where the audio description includes a description representing the audio characteristics.

[0114] In one implementation, one audio type corresponds to at least one audio characteristic.

[0115] In one embodiment, at least one audio characteristic is predicted based on an expert network model corresponding to the audio type, and the expert network model includes at least one sub-model, wherein one audio type corresponds to one expert network model, and one sub-model is used to determine one audio characteristic.

[0116] In one embodiment, the determination unit 102 is further configured to: determine the confidence level of the audio characteristics corresponding to at least one audio type; and in response to conflicting audio characteristics, select the audio characteristics with the highest confidence level among the conflicting audio characteristics as the audio characteristics used in the audio description.

[0117] In one embodiment, the processing unit 101 is further used to: obtain metadata of the audio, the metadata including information associated with the audio description of the audio; based on the audio characteristics corresponding to at least one audio type, the generation unit generates the audio description of the audio in the following manner: based on the metadata and the audio characteristics corresponding to at least one audio type, generate the audio description of the audio.

[0118] In one embodiment, based on the audio characteristics corresponding to at least one audio type, the generation unit 103 generates an audio description of the audio in the following manner: the audio characteristics corresponding to at least one audio type are input into a model, the model has a thought chain parsing capability, and the thought chain is used to describe the display reasoning process of audio description generation; based on the thought chain parsing capability, the association relationship between the audio characteristics corresponding to at least one audio type is determined; the association relationship is displayed, and a reliable association relationship is determined, and the description of the audio characteristics corresponding to the reliable association relationship is used as the audio description of the audio.

[0119] In one embodiment, the audio type includes at least one of the following: a voice type; a music type; an ambient sound type; a combination type, wherein the combination type includes a combination of at least two of the voice type, the music type, and the ambient sound type.

[0120] In one embodiment, in response to the audio type being a speech type, the audio characteristics include at least one of the following: language, speech recognition, emotion, age, gender, accent, and the paragraph corresponding to the speaker; in response to the audio type being a music type, the audio characteristics include at least one of the following: instrument, genre, and emotion; in response to the audio type being an ambient sound type, the audio characteristics include at least one of the following: quality, reverberation level, and sound intensity.

[0121] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0122] Figure 8 2 is a block diagram illustrating an electronic device 200 for generating audio description according to some embodiments of the present disclosure. For example, the electronic device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0123] Reference Figure 8 , electronic device 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0124] The processing component 202 generally controls the overall operation of the electronic device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0125] The memory 204 is configured to store various types of data to support operations on the electronic device 200. Examples of such data include instructions for any application or method operating on the electronic device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0126] The power component 206 provides power to the various components of the electronic device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 200.

[0127] The multimedia component 208 includes a screen that provides an output interface between the electronic device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the electronic device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0128] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.

[0129] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0130] The sensor assembly 214 includes one or more sensors for providing various aspects of status assessment for the electronic device 200. For example, the sensor assembly 214 can detect the open / closed state of the electronic device 200, the relative positioning of components, such as the display and keypad of the electronic device 200. The sensor assembly 214 can also detect changes in the position of the electronic device 200 or a component of the electronic device 200, the presence or absence of user contact with the electronic device 200, the orientation or acceleration / deceleration of the electronic device 200, and temperature changes of the electronic device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0131] The communication component 216 is configured to facilitate wired or wireless communication between the electronic device 200 and other devices. The electronic device 200 can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, 5G, other communication standards, or a combination thereof. In some embodiments of the present disclosure, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of the present disclosure, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0132] In some embodiments of the present disclosure, the electronic device 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0133] In some embodiments of the present disclosure, a storage medium including instructions is further provided, such as a memory 204 including instructions, and the instructions can be executed by the processor 220 of the electronic device 200 to perform the above method. For example, the storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0134] In some embodiments of the present disclosure, a storage medium is provided, which may be, for example, a non-transitory computer-readable storage medium.

[0135] In some embodiments of the present disclosure, when the instructions in the storage medium are executed by the processor of the electronic device 200, the electronic device 200 is enabled to execute the above-mentioned methods.

[0136] Figure 9 FIG. 3 is a block diagram of an apparatus 300 for generating audio description according to some embodiments of the present disclosure. For example, the apparatus 300 may be provided as a server. Figure 9 The apparatus 300 includes a processing component 322, which further includes one or more processors, and a memory resource represented by a memory 332 for storing instructions, such as an application, that can be executed by the processing component 322. The application stored in the memory 332 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 322 is configured to execute the instructions to perform the above-described method.

[0137] The device 300 may also include a power supply component 326 configured to perform power management of the device 300, a wired or wireless network interface 350 configured to connect the device 300 to a network, and an input / output (I / O) interface 358. The device 300 may operate based on an operating system stored in the memory 332, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0138] In some embodiments of the present disclosure, a storage medium including instructions is further provided, such as a memory 332 including instructions, and the instructions can be executed by the processing component 322 of the apparatus 300 to perform the above method. For example, the storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0139] In some embodiments of the present disclosure, a storage medium is provided, which may be, for example, a non-transitory computer-readable storage medium.

[0140] In some embodiments of the present disclosure, when the instructions in the storage medium are executed by the processor of the device 300, the device 300 is enabled to perform the above-mentioned methods.

[0141] In some embodiments of the present disclosure, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the method for generating audio description involved in any of the above embodiments is implemented.

[0142] In an exemplary embodiment, a processor for executing a computer program may be deployed in, for example, an electronic device.

[0143] Those skilled in the art will also appreciate that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present application.

[0144] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of these features.

[0145] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0146] In the description herein, the meaning of "a plurality" is at least two, and refers to two or more, such as two, three, etc., unless otherwise clearly defined. Other quantifiers are similar. The singular forms "a", "the" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise specified or clearly directed to a singular form from the context, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0147] It should be understood that, unless otherwise specifically noted, the features of some embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; "and / or" describes the association relationship of the associated objects, indicating that three relationships may exist, for example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. Similarly, "at least one of..." includes any one of the relevant listed items and any combination of any two or more thereof.

[0148] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0149] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then applying A or B is satisfied under any of the aforementioned instances.

[0150] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the rights. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise noted, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the components described, even if structurally not equivalent to the disclosed structure. In addition, although the specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial to any given or specific application. In addition, with respect to "including," "having," "having," "having," or variations thereof used in the present disclosure, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0151] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0152] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.

Claims

1. A method for generating audio description, characterized in that: include: Acquire audio, and identify the audio type of the audio to obtain at least one audio type; For the at least one audio type, determining audio characteristics of the audio according to the audio type to obtain audio characteristics corresponding to each of the at least one audio type; An audio description of the audio is generated based on the audio characteristics corresponding to each of the at least one audio type, where the audio description includes a description representing the audio characteristics.

2. The method according to claim 1, characterized in that An audio type corresponds to at least one audio characteristic.

3. The method according to claim 2, characterized in that The at least one audio characteristic is predicted based on an expert network model corresponding to the audio type, wherein the expert network model includes at least one sub-model, wherein one audio type corresponds to one expert network model, and one sub-model is used to determine one audio characteristic.

4. The method according to any one of claims 2 to 3, characterized in that The method further comprises: respectively determining a confidence level of an audio characteristic corresponding to each of the at least one audio type; In response to conflicting audio features, an audio feature with the highest confidence level is selected as the audio feature used in the audio description.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining metadata for the audio, the metadata comprising information associated with an audio description of the audio; Generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type includes: An audio description of the audio is generated based on the metadata and audio characteristics corresponding to the at least one audio type.

6. The method according to claim 1, characterized in that Generating an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type includes: Inputting audio characteristics corresponding to each of the at least one audio type into a model, wherein the model has a thought chain parsing capability, wherein the thought chain is used to describe a display reasoning process for generating the audio description; Determining, based on the thought chain parsing capability, an association relationship between audio characteristics corresponding to each of the at least one audio type; The association relationship is displayed, and a reliable association relationship is determined, and a description of the audio characteristics corresponding to the reliable association relationship is used as an audio description of the audio.

7. The method according to any one of claims 1 to 3, characterized in that The audio type includes at least one of the following: Voice type; Music genre; Ambient sound type; The combination type includes a combination of at least two of a voice type, a music type, and an ambient sound type.

8. The method according to claim 7, characterized in that In response to the audio type being the speech type, the audio characteristics include at least one of the following: language, speech recognition, emotion, age, gender, accent, and a paragraph corresponding to the speaker; In response to the audio type being the music type, the audio characteristics include at least one of the following: an instrument, a genre, and a mood; In response to the audio type being the ambient sound type, the audio characteristics include at least one of the following: quality, reverberation level, and sound intensity.

9. A device for generating audio description, characterized in that: include: a processing unit, configured to acquire audio and identify an audio type to which the audio belongs, so as to obtain at least one audio type; a determining unit, configured to determine, for the at least one audio type, audio characteristics of the audio according to the audio type, so as to obtain audio characteristics corresponding to each of the at least one audio type; The generating unit is configured to generate an audio description of the audio based on the audio characteristics corresponding to each of the at least one audio type, wherein the audio description includes a description representing the audio characteristics.

10. The device according to claim 9, characterized in that An audio type corresponds to at least one audio characteristic.

11. The device according to claim 10, characterized in that The at least one audio characteristic is predicted based on an expert network model corresponding to the audio type, wherein the expert network model includes at least one sub-model, wherein one audio type corresponds to one expert network model, and one sub-model is used to determine one audio characteristic.

12. The device according to any one of claims 10 to 11, characterized in that The determining unit is further configured to: respectively determining a confidence level of an audio characteristic corresponding to each of the at least one audio type; In response to conflicting audio features, an audio feature with the highest confidence level is selected as the audio feature used in the audio description.

13. The device according to claim 9, characterized in that The processing unit is further configured to: Obtaining metadata for the audio, the metadata comprising information associated with an audio description of the audio; Based on the audio characteristics corresponding to the at least one audio type, the generating unit generates the audio description of the audio in the following manner: An audio description of the audio is generated based on the metadata and audio characteristics corresponding to the at least one audio type.

14. The device according to claim 9, characterized in that Based on the audio characteristics corresponding to the at least one audio type, the generating unit generates the audio description of the audio in the following manner: Inputting audio characteristics corresponding to each of the at least one audio type into a model, wherein the model has a thought chain parsing capability, wherein the thought chain is used to describe a display reasoning process for generating the audio description; Determining, based on the thought chain parsing capability, an association relationship between audio characteristics corresponding to each of the at least one audio type; The association relationship is displayed, and a reliable association relationship is determined, and a description of the audio characteristics corresponding to the reliable association relationship is used as an audio description of the audio.

15. The device according to any one of claims 9 to 11, characterized in that The audio type includes at least one of the following: Voice type; Music genre; Ambient sound type; The combination type includes a combination of at least two of a voice type, a music type, and an ambient sound type.

16. The device according to claim 15, characterized in that In response to the audio type being the speech type, the audio characteristics include at least one of the following: language, speech recognition, emotion, age, gender, accent, and a paragraph corresponding to the speaker; In response to the audio type being the music type, the audio characteristics include at least one of the following: an instrument, a genre, and a mood; In response to the audio type being the ambient sound type, the audio characteristics include at least one of the following: quality, reverberation level, and sound intensity.

17. An electronic device, characterized in that: include: processor; memory for storing computer programs or instructions executable by a processor; The processor is configured to execute the computer program or instructions to implement the steps of the method for generating audio description according to any one of claims 1 to 8.

18. A storage medium, characterized in that The storage medium stores a computer program or instructions. When the computer program or instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method for generating audio description according to any one of claims 1 to 8.

19. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for generating an audio description according to any one of claims 1 to 8.