An audio question and answer method, device, equipment, storage medium and program product
Patent Information
- Application Number
- CN202510182481.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]对于音频类问答的应用场景,需要专门的音频问答技术来做支持,其中的输入数据主要为音频片段以及音频关联的问题描述文本,然而,由于音频应用场景下可询问的音频问题任务比较多样化,现有的音频问答技术并不能灵活地支持各不同音频问题任务地有效分析以及答复
[0021] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio question-and-answer method provided in any embodiment of this disclosure.
Smart Images

Figure CN122598623A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio processing technology, and in particular to an audio question-and-answer method, apparatus, device, storage medium, and program product. Background Technology
[0002] Question-and-answer matching is required in many application scenarios, where users ask questions and the appropriate answer is provided. Currently, the most common application scenarios are text-based or image-based question-and-answer, where the user inputs a text description of the question, or an image along with a related text description, and the question-and-answer server provides a suitable answer based on the input.
[0003] For audio question answering applications, specialized audio question answering technology is required. The input data mainly consists of audio segments and audio-related question description text. However, due to the diverse range of audio questions that can be asked in audio application scenarios, existing audio question answering technologies cannot flexibly support the effective analysis and response to various audio question tasks. Summary of the Invention
[0004] This disclosure provides an audio question-and-answer method, apparatus, device, storage medium, and program product, which realizes adaptive matching of response results in audio question-and-answer scenarios and improves the accuracy of response results.
[0005] In a first aspect, embodiments of this disclosure provide an audio question-and-answer method, the method comprising:
[0006] Obtain the audio sequence and the problem description text;
[0007] Determine the encoded feature sequence of the audio sequence, and determine the problem description features of the problem description text;
[0008] Based on the encoded feature sequence, the task-associated encoded features corresponding to at least one preset audio problem task are determined, and the problem description weights relative to each of the audio problem tasks are determined based on the problem description features.
[0009] Based on the task-related coding features and the corresponding problem description weights, the coding enhancement features associated with the problem description text are determined.
[0010] Based on the encoded enhancement features and the question description features, the question response results associated with the audio sequence are determined.
[0011] Secondly, embodiments of this disclosure also provide an audio question-and-answer device, the device comprising:
[0012] The acquisition module is used to acquire audio sequences and problem description text;
[0013] The first determining module is used to determine the encoded feature sequence of the audio sequence and the problem description features of the problem description text;
[0014] The second determining module is used to determine, based on the coding feature sequence, the task association coding features corresponding to at least one preset audio problem task, and to determine the problem description weight relative to each of the audio problem tasks based on the problem description features.
[0015] The third determining module is used to determine the coding enhancement features associated with the problem description text based on the coding features associated with each task and the corresponding problem description weights.
[0016] The response determination module is used to determine the question response result associated with the audio sequence based on the encoded enhancement features and the question description features.
[0017] Thirdly, embodiments of this disclosure also provide a computer device, the computer device comprising:
[0018] One or more processors;
[0019] Storage device for storing one or more programs.
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio question-and-answer method provided in any embodiment of this disclosure.
[0021] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio question-and-answer method provided in any embodiment of this disclosure.
[0022] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the audio question-and-answer method provided in any embodiment of this disclosure.
[0023] This disclosure provides an audio question-answering method, apparatus, device, storage medium, and program product. The method first acquires an audio sequence and question description text; determines the encoding feature sequence of the audio sequence and the question description features of the question description text; based on the encoding feature sequence, determines task-related encoding features corresponding to at least one preset audio question task, and determines the question description weight relative to each audio question task based on the question description features; determines the encoding enhancement features associated with the question description text based on each task-related encoding feature and the corresponding question description weight; and determines the question response result associated with the audio sequence based on the encoding enhancement features and the question description features. This embodiment's technical solution, based on the effective perception of audio features involved in question-answering through question description, achieves adaptive matching of different question descriptions in audio question-answering scenarios. This technical solution considers the correlation between the encoding features of audio data and different audio question tasks, while also taking into account the weight of the question description text relative to different audio question tasks. By highlighting the audio question task emphasized by the question description text through the weight results for different audio question tasks, it then combines the weights with the encoding features of the audio to strengthen the encoding features related to the emphasized audio question task. Finally, by combining the enhanced encoding features with the question description text, it is possible to analyze the audio response results associated with the audio sequence that are more suitable for the question description text. Compared with existing technical solutions, this solution achieves a comprehensive and effective understanding of the audio sequence, thereby perceiving the audio encoding features that better match the question description text. This enables adaptive matching of audio question answering and ensures the accuracy of audio question answering when faced with different question descriptions in audio question answering scenarios. Attached Figure Description
[0024] To more clearly illustrate the technical methods of the exemplary embodiments of this disclosure, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the embodiments to be described in this disclosure, and not all of them. Those skilled in the art can derive other drawings from these drawings without any creative effort.
[0025] Figure 1a A flowchart illustrating an audio question-and-answer method provided in an embodiment of this disclosure;
[0026] Figure 1b This is a model framework diagram of the audio question answering method provided in this embodiment of the present disclosure, which implements audio question answering through a determined audio question answering model;
[0027] Figure 2 This is a schematic diagram of the structure of an audio question-and-answer device provided in an embodiment of the present disclosure;
[0028] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0033] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0035] It is understood that before using the technical methods disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0036] For example, in response to receiving a user's active muting, a prompt message is sent to the user to explicitly inform them that the muting operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technology, based on the prompt message.
[0037] As an optional but non-limiting implementation, in response to receiving a user's active prompt, sending a notification message to the user could be done via a pop-up window, where the notification message could be presented in text format. Furthermore, the pop-up window could also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0038] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0039] It should be noted that one application scenario involved in this embodiment can be described as follows: in an audio question-answering scenario, questions can be posed based on the given audio segments. The dimensions of the questions are quite diverse, such as those related to speech recognition, audio theory, or environmental sound. In existing audio question-answering implementations, one type often excels only at a certain type of question and struggles to provide suitable answers for different types of questions. While some improved implementations can respond to different types of questions, they consider the ability to solve various types of questions equally during the audio question-answering analysis phase, making them more like a general-purpose question-answering implementation. Therefore, the answers they provide may not be the best or most complete. Thus, existing audio question-answering implementations cannot achieve adaptive question answering for different questions related to different audio clips.
[0040] Based on this, this embodiment provides an audio question-and-answer method that can adaptively respond to different questions involving different audio in an audio question-and-answer scenario.
[0041] Specifically, Figure 1 is a flowchart illustrating an audio question-and-answer method provided in an embodiment of this disclosure. This embodiment is applicable to question-and-answer scenarios in audio environments. The method can be executed by an audio question-and-answer device, which can be implemented by software and / or hardware and can be configured in a terminal and / or server to implement the audio question-and-answer method in this embodiment of the disclosure.
[0042] As shown in Figure 1, the audio question-and-answer method provided in this embodiment may include the following steps:
[0043] S101. Obtain the audio sequence and the problem description text.
[0044] In this embodiment, the audio sequence can be considered as the audio file used to ask questions in an audio question-and-answer scenario, which is often an audio format file containing audio data of a certain duration. Similarly, the textual question content entered in the audio question-and-answer scenario can be recorded as question description text, which can be entered between the question boxes involved in the stable audio scenario, or a question description text can be selected from a pre-set set of question description information.
[0045] S102. Determine the encoded feature sequence of the audio sequence and the problem description features of the problem description text.
[0046] It is understood that in audio question-and-answer scenarios, to obtain a suitable answer to a question, the audio needs to be understood and analyzed. To understand and analyze the audio, the characteristic information involved in the audio needs to be obtained; similarly, to know the specific content of the question, the characteristic information involved in the question also needs to be obtained.
[0047] In this embodiment, the encoded feature sequence of the audio sequence can be determined through this step. The audio sequence may contain audio data of different audio types. To obtain the audio features of different audio data, it is necessary to use encoders adapted to different audio types for audio encoding. Therefore, in one implementation of the encoded features, this embodiment can use multiple audio encoders to encode the audio sequence. Different audio encoders can focus on processing audio data of different audio types in the audio sequence, thereby obtaining the encoded features output by different audio encoders after encoding the audio data in the audio sequence. The output encoded features can be merged into the encoded feature sequence of the audio sequence.
[0048] It is known that audio encoders involved in audio sequence encoding can include language recognition or language translation encoders that focus on language data processing, classification encoders that focus on non-speech data and applicable audio event classification, and music design encoders that focus on music-related rhythm and melody analysis.
[0049] As described above, this embodiment can achieve the fusion of encoded features output by each encoder by increasing the channel dimension. Considering that the encoded features of the audio need to be combined with the question description features of the question description text to further determine the answer, the size of the encoded features that participate in subsequent processing along with the question description features needs to be limited. However, because the lengths of the processed audio sequences are different, the channel dimensions occupied by the encoded features of different audio sequences after fusion are different. This embodiment can use a window adjustment mechanism to limit the encoding features of different channel dimensions to be integrated at a pre-set length, generating a fixed-length encoded feature sequence that participates in subsequent execution logic.
[0050] In this embodiment, existing text feature processing algorithms can also be used to extract features from the problem description text, thereby obtaining the problem description features of the problem description text.
[0051] S103. Based on the encoded feature sequence, determine the task association encoded features corresponding to each of the preset at least one audio problem task, and based on the problem description features, determine the problem description weights relative to each of the audio problem tasks.
[0052] In this embodiment, one or more audio question tasks involved in the audio response scenario can be pre-defined. The audio question task can be understood as a question category involved in the audio scenario. An audio question task can correspond to a question category, representing the audio question involved in that question category.
[0053] In this embodiment, corresponding coding feature analysis logic is pre-defined for different audio question tasks. This step analyzes the coding feature sequence by executing different coding feature analysis logics, thereby obtaining relevant task-related coding features for different audio question tasks. The task-related coding features can be considered as the extraction of coding features used when answering the corresponding audio question task.
[0054] Analysis shows that the questions contained in the question description text can be associated with one or more audio question tasks. In this embodiment, the specific audio question task(s) involved in the question description text can be determined by determining the weight of the question description features relative to each audio question task.
[0055] In this embodiment, one way to determine the weight of the problem description can be described as follows: pre-determine matching weight matrices for different audio problem tasks, and then multiply the problem description features with each weight matrix. The result of the multiplication can be used to characterize the weight value of the problem description text relative to each audio problem text. This weight value can be recorded as the problem description weight.
[0056] It is known that if the problem description text contains questions related to one or more audio question tasks, then the weight value of the determined problem description weight will be relatively high.
[0057] S104. Based on the task-related coding features and the corresponding problem description weights, determine the coding enhancement features associated with the problem description text.
[0058] In this embodiment, relative to one or more pre-defined audio question tasks, the above steps can determine the associated task-related coding features and the question description weights relative to the question description features. This embodiment can combine the task-related coding features and question description weights associated with the same audio question task. For example, the task-related coding features can be weighted according to the question description weights, and the resulting feature information can be used as the coding enhancement features associated with the question description text.
[0059] Understandably, by using question description weights, we can identify which audio question tasks are involved in the question description text. Knowing the task-related coding features associated with these audio question tasks, we can determine the question description weights for those features, thereby increasing their importance in question-answering analysis. This embodiment obtains enhanced coding features relative to each audio question task from the task-related coding features through this step.
[0060] In this embodiment, the encoding enhancement feature can highlight the encoding features involved in the problem description text in the audio sequence, thereby achieving more effective perception that is more adapted to the encoding features.
[0061] S105. Determine the question response result associated with the audio sequence based on the encoded enhancement features and the question description features.
[0062] In this embodiment, by analyzing the encoded reinforcement features and the question description features, a question-and-answer result that better matches the audio sequence and the question description text can be adapted from the knowledge base. The question-and-answer result can be presented in text format. Specifically, a question-and-answer analysis model can be used to analyze the encoded reinforcement features and the question description features, and finally determine the answer from the knowledge base associated with the question-and-answer analysis model.
[0063] This embodiment provides an audio question-answering method that effectively perceives the audio features involved in the question and answer based on the question description, thereby achieving adaptive matching for different question descriptions in audio question-answering scenarios. This technical solution considers the correlation between the encoding features of audio data and different audio question tasks, while also considering the weight of the question description text relative to different audio question tasks. By highlighting the audio question task emphasized by the question description text through the weight results for different audio question tasks, it then considers combining the weights with the encoding features of the audio to strengthen the encoding features related to the emphasized audio question task. Finally, by combining the enhanced encoding features obtained after feature enhancement with the question description text, it is possible to analyze the audio response results associated with the audio sequence that are more suitable for the question description text. Compared with existing technical solutions, this technical solution achieves a comprehensive and effective understanding of the audio sequence, thereby perceiving the audio encoding features that better match the question description text, thus achieving adaptive matching of audio question-answering and ensuring the accuracy of audio question-answering when faced with different question descriptions in audio question-answering scenarios.
[0064] As a first optional embodiment of this example, based on the above embodiment, the steps for determining the encoded feature sequence of the audio sequence and the problem description features of the problem description text can be further optimized as follows:
[0065] a1) The audio sequence is encoded by the audio encoding layer in the determined audio question-answering model to obtain the encoding fusion features of the audio sequence. The audio encoding layer contains at least one audio encoder, and each audio encoder is used to encode different types of audio data in the audio sequence.
[0066] In this embodiment, an audio question-answering model is preferably used to achieve effective audio question-answering processing. The audio stabilization model set in this embodiment may include an audio coding layer, and one or more audio encoders may be integrated into the audio coding layer. In specific implementation, the integrated audio encoders can be used to encode the audio sequence, and the audio coding features output by each audio encoder can be obtained.
[0067] In this embodiment, various audio coding features can be fused. For example, the fusion of audio coding features can be achieved by concatenating features along the channel dimension, and finally obtaining the fused coding features output by the audio coding layer.
[0068] In this embodiment, the audio encoder included in the audio coding layer may include encoders that process different types of audio, such as encoders focused on speech tasks, which perform automatic speech recognition and language translation by extracting features such as phoneme structure and prosody. Another example is an encoder focused on non-speech audio, which models wide spectrum and event patterns through self-supervised learning to be suitable for tasks such as audio event classification. Yet another example is a music-focused encoder designed for music tasks, which can generally be used to capture harmony, rhythm, and melodic structure to achieve tasks such as generating music subtitles.
[0069] b1) Through the feature alignment layer in the audio question-answering model, the encoded fusion features are divided into a set number of encoded feature units, and the encoded feature units are aggregated to form the encoded feature sequence.
[0070] In this embodiment, the audio question-answering model also includes a feature alignment layer. This feature alignment layer can adopt a window-level audio and language alignment mechanism. Through this feature alignment layer, the number of alignments involved can be determined according to the set feature dimensions and the number of alignments. In this embodiment, the number of alignments is recorded as the set number. After determining the set number, the encoded feature sequence can be divided into the set number of encoded feature units. Each encoded feature unit can contain encoded features with a certain channel dimension.
[0071] It is known that the size of the audio sequence varies, and the size of the generated coding feature sequence also varies. After specifying the number of divisible units, the coding feature sequence can be adaptively divided to form a set number of coding feature units. Each coding feature unit can retain rich time information of the audio sequence. Finally, by summarizing each coding feature unit, a coding feature sequence can be formed.
[0072] c1) By using the text extraction layer in the audio question-answering model, text features are extracted from the question description text to obtain the question description features of the question description text.
[0073] In this embodiment, the audio question-answering model also includes a text extraction layer, which can be constructed using a text extraction network. Through this text feature extraction network, text features can be extracted from the question description text, and the question description features output by the text extraction layer relative to the question description text can be obtained.
[0074] This embodiment provides a specific implementation of the encoded feature sequence and question description features. This embodiment considers extracting multi-scale audio features from audio sequences by integrating multiple audio encoders. It also considers aligning the number of encoded features involved in different audio sequences. The aligned encoded feature sequence and question description features can provide basic data support for subsequent logic in audio question answering.
[0075] As a second optional embodiment of this example, based on the above embodiment, the steps of determining the task association coding features corresponding to at least one preset audio problem task according to the coding feature sequence, and determining the problem description weights relative to each of the audio problem tasks according to the problem description features, can be further optimized as follows:
[0076] a2) Input the encoded feature sequence and the question description features into the task awareness module in the constructed audio question answering model.
[0077] In this embodiment, when using an audio question-answering model for audio question answering, a task awareness module is added to the audio question-answering model. The encoded feature sequence and question description features determined through the above embodiments can be used as input data for the task awareness module.
[0078] The task awareness module can be understood as a network module that adaptively adapts the encoded feature sequence to the audio problem tasks involved in the problem description features. This task awareness module may include a feature processing submodule and a task routing submodule. The feature processing submodule can be used to determine task-related encoded features from the encoded feature sequence for multiple preset different audio problem tasks; the task routing submodule can be used to determine the weight information of the input problem description features relative to different audio problem tasks.
[0079] b2) The feature processing submodule in the task perception module processes the encoded feature sequence according to the included feature analysis networks to obtain the task-related encoded features relative to each set audio problem task.
[0080] In this embodiment, the feature processing submodule can be considered to include multiple feature analysis networks. Each feature analysis network can correspond to a preset audio problem task and can be used to process the encoded feature sequence. Finally, the task-related encoded features output by each feature analysis network can be obtained. The task-related encoded features can be considered as the encoded features associated with the corresponding audio problem task in the encoded feature sequence.
[0081] c2) The task routing submodule in the task perception module processes the problem description features according to the weight matrices of each task to obtain the problem description weights relative to each audio problem task.
[0082] In this embodiment, the task routing submodule can be considered to include multiple trained task weight matrices. Each task weight matrix also corresponds to a preset audio problem task. Each task weight matrix can be subjected to matrix operations with the problem description features. The result of the operation is a weight value. The weight value output by each task weight matrix can be considered as the problem description weight relative to each audio problem task.
[0083] This embodiment presents a specific implementation of determining task association encoding features for multiple different audio question tasks and assigning weights to question descriptions through the task awareness module in an audio question-answering model. This technical solution achieves adaptive perception of the audio question tasks involved in the question description text and assigns weights that characterize the importance of the associated audio question tasks. It also provides fundamental data support for the subsequent implementation of audio question-answering logic.
[0084] As a third optional embodiment of this embodiment, based on the above embodiment, the determination of the coding enhancement feature of the problem description text association according to the task association coding features and the corresponding problem description weights can be further optimized as follows: input the task association coding features and the corresponding problem description weights into the weighted operation layer in the constructed audio problem model to obtain the weighted coding features, and determine the weighted coding features as the coding enhancement feature of the problem description text association.
[0085] In this embodiment, the audio question answering model also includes a weighted operation layer after the task awareness module. Through this weighted operation layer, the task association coding features and question description weights output by the task awareness module relative to each audio question task can be input to the weighted operation layer. The weighted operation layer performs weighted processing on the task association coding features and question description weights, and a weighted coding feature is formed after the weighted processing. This coding feature can be recorded as the coding enhancement feature of the question description text association.
[0086] It is understandable that this encoding enhancement feature highlights the encoded features in the audio sequence that are related to the problem described in the text. Specifically, this encoding enhancement feature can be considered to have the same number of encoded feature units as the encoded feature sequence; therefore, this embodiment can be considered to have only implemented feature enhancement, without reducing or increasing the feature channel dimension.
[0087] The above-described technical solution in this embodiment provides a specific implementation for determining the encoding enhancement features. Unlike existing technologies that directly use the encoding features output by the audio encoder to participate in subsequent question-and-answer analysis, the above-described technical solution in this embodiment, through task awareness and feature enhancement extraction, is equivalent to a more comprehensive understanding of the audio encoding features and the question description text. It also ensures the compatibility between the encoding enhancement features and the question description text, providing more accurate basic data support for subsequent question-and-answer analysis.
[0088] As a fourth optional embodiment of this example, based on the above embodiments, the determination of the question-response result associated with the audio sequence according to the encoded enhancement features and the question description features can be further optimized as follows:
[0089] a3) By analyzing the question-answering analysis sub-model in the determined audio question-answering model, the encoding enhancement features and the question description features are analyzed to obtain the question-answering text features.
[0090] In this embodiment, the audio question-answering model also includes a question-answering analysis sub-model. This question-answering analysis sub-model can be built after the weighted operation layer mentioned above. The encoded enhancement features output by the weighted operation layer and the pre-determined question description features can be used as inputs to the question-answering analysis sub-model. By analyzing the encoded enhancement features and question description features, question response features that are more suitable for the question description text can be determined.
[0091] It can be seen that the question-answering analysis sub-model can be regarded as a network model built on a large language model. Through feature analysis and combination with the knowledge base, it can obtain the response features that are adapted to the question description text. The response features are set in text form and can be preferably recorded as question-answer text features.
[0092] b3) Decode the features of the question-and-response text, generate the question-and-response text, and identify the question-and-response text as the question-and-response result associated with the audio sequence.
[0093] In this embodiment, the determined question-and-response text features can be decoded through this step to restore the corresponding text information. This text information can be recorded as the generated question-and-response text. At the same time, the question-and-response text can also be used as the question-and-response result associated with the obtained audio sequence relative to the question description text.
[0094] The technical solution described in this embodiment provides an implementation for determining the question-and-answer result. The determined result is primarily obtained using pre-determined encoding enhancement features and question description features. The pre-determined encoding enhancement features incorporate a technical implementation that adaptively perceives the audio features associated with the question description text. Compared to existing technologies, the question-and-answer result determined by this solution achieves a comprehensive and effective understanding of the audio sequence, thereby perceiving audio encoding features that better match the question description text. This enables adaptive matching of audio question-and-answer and ensures the accuracy of audio question-and-answer when faced with different question descriptions in audio question-and-answer scenarios.
[0095] To facilitate a better understanding of the audio question-answering method provided in this embodiment, this embodiment introduces the use of an audio question-answering model to implement audio question-answering. Figure 1b This is a model framework diagram of the audio question answering method provided in this embodiment of the disclosure, which implements audio question answering through a determined audio question answering model, as shown below. Figure 1b The diagram illustrates the process by which an audio question-answering model processes audio sequences to achieve audio question answering. The audio sequence 11 may include speech, music, and other audio data, such as background noise or ambient sound. Figure 1b It also includes a question description text 12 as input data, which can describe the questions involved in the question-and-answer session.
[0096] As described above, the audio question-answering model 1 includes an audio encoding layer 13, a text extraction layer 14, a feature alignment layer 15, a task awareness module 16, a weighted operation layer 17, and a question-answering analysis sub-model 18. Through this audio question-answering model 1, the audio sequence 11 can be input into the audio encoding layer 13, and then processed by the feature alignment layer 15 to obtain an encoded feature sequence. The question description text 12 is then processed by the text extraction layer 14 to output question description features. The encoded feature sequence and question description features can continue to be input to the task awareness module 16. The output task-related encoded features and question description weights are processed by the weighted operation layer 17 to output encoded enhancement features with the same dimension as the encoded feature sequence. The encoded enhancement features and question description features are then analyzed again by the question-answering analysis sub-model 18 to output question-answer text features. Finally, the question-answer text involved in the audio question-answering can be obtained by decoding the question-answer text features.
[0097] Based on the optimizations of the above embodiments, this fifth optional embodiment can optimize the audio question answering model obtained by training the constructed initial audio question answering model.
[0098] It is understood that the audio question-answering model used in this embodiment needs to be trained in advance. Specifically, the training steps for optimizing the audio question-answering model include the following steps:
[0099] a4) Obtain a sample training set and a pre-built initial audio question answering model. The sample training set includes at least one sample triplet, which includes sample audio, sample question text, and real question answer results.
[0100] In this embodiment, this step can be used to obtain the sample training set required for model training, as well as the initial audio response model to participate in the training. Considering that the input data of the audio response model includes audio sequences and question description text, and considering that the loss function value needs to be determined by referring to the real results during model training, the sample data in the sample training set can be set as sample triples when generating the sample training set. The sample triples can include sample audio and sample question text as input data, as well as the real question response results that participate in the determination of the loss function.
[0101] In this embodiment, the pre-built initial audio question-answering model can be considered to have the same network structure as the audio question-answering model used in actual applications. The initial audio question-answering model includes an audio encoding layer, a text extraction layer, a feature alignment layer, a task awareness module, a weighted operation layer, and a question-answering analysis sub-model, and the network parameters of each included network structure can be considered to be initial set values.
[0102] b4) Select a sample triplet, input the sample audio and sample question text contained therein into the initial audio question answering model, and obtain the predictive coding enhancement features output by the task awareness module included in the initial audio question answering model.
[0103] In this embodiment, the training of the initial audio network model can begin with this step. First, a sample triplet can be selected from the sample training set, and the sample audio and sample question text contained therein are used as input data. The initial audio question answering model is used to process the input data. Through the network structure contained in the initial audio question answering model, the encoded enhancement features output by the task perception module after processing are extracted from the initial audio question answering model. In this embodiment, these encoded enhancement features can be denoted as predicted encoded enhancement features.
[0104] It should be noted that existing model training methods directly combine the model's output with the actual results to determine the loss function value. This training approach essentially considers only the cross-entropy loss value for aligning audio and text features. However, this alignment is only effective when the differences between the two objects being aligned are small. Since audio and text modalities differ significantly, audio sequences are often longer than text sequences, and audio easily contains repetitive or irrelevant information, complicating the alignment process. Therefore, training solely using this alignment method will negatively impact the model's accuracy.
[0105] Therefore, to enhance the cross-modal alignment between audio and text features, this embodiment also considers the impact of the predictive coding enhancement features output by the task awareness module on model training. Thus, this step first obtains the predictive coding enhancement features from the initial audio question-answering model to participate in the model training operation.
[0106] c4) Using the given semantic awareness alignment model, perform feature filtering on the predicted coding enhancement features to obtain coding enhancement filtering features that are semantically aligned with the sample question text.
[0107] Following the above description, this step can be used to optimize and train the output predicted coding enhancement features. Specifically, this embodiment additionally introduces a usable semantic awareness alignment model. This model performs feature filtering on the predicted coding enhancement features to remove redundant or repetitive features, thereby obtaining coding enhancement filtering features that can better semantically align with the sample question text.
[0108] In subsequent processing of the iterations, the predicted encoded enhanced features can be replaced by encoded enhanced filtering features as input to the question-answering analysis sub-model. This embodiment introduces semantic alignment training between audio encoded features and question text features by filtering encoded enhanced features during model training, which better ensures the accuracy and consistency of audio-text alignment.
[0109] In this embodiment, the process of obtaining the coding enhancement screening features through the semantic awareness alignment model can be described as follows: First, the similarity between the predicted coding enhancement features and the sample question description features of the sample question text is calculated to determine the similarity value of each feature in the predicted coding enhancement features. The predicted coding enhancement features can be clustered using the similarity value. The clustering results can constitute the coding enhancement screening features after filtering out redundant or duplicate features.
[0110] For example, as one implementation, the process of performing feature filtering on the predicted encoded enhancement features using a given semantic awareness alignment model to obtain encoded enhancement filtering features that are semantically aligned with the sample question text can be specified as the following steps:
[0111] c41) Input the predicted encoding enhancement features and the sample problem description features of the sample problem description text into the semantic awareness alignment model.
[0112] This step can be viewed as an input operation for the input data, which includes the predictive encoding enhancement features and the sample question description features. It is known that the sample question description features can be obtained by the text extraction layer in the initial audio question answering model through processing the sample question description text.
[0113] c42) Through the similarity determination network in the semantic awareness alignment model, determine the similarity score of each predicted feature unit included in the predicted encoding enhancement feature, and generate a feature aggregation decision matrix by combining each similarity score according to a set threshold.
[0114] In this embodiment, the semantic-aware alignment model can be considered to include a similarity determination network. This network structure enables the determination of similarity scores for each predicted feature unit included in the predicted encoded enhancement features. The similarity score can be understood as a numerical value representing the similarity between the predicted feature unit and the sample problem description features. Furthermore, this step can also filter the predicted feature units corresponding to each similarity score by setting a threshold, thereby removing predicted feature units with similarity scores below the set threshold, and generating a feature aggregation decision matrix using the retained predicted feature units.
[0115] c43) Through the feature aggregation network in the semantic awareness alignment model, each predicted feature unit is aggregated according to the feature aggregation decision matrix to obtain the aggregated feature aggregation unit, and each feature aggregation unit is used to form the encoded reinforcement screening feature.
[0116] In this embodiment, the feature aggregation decision matrix determined in the above steps can be used as input information for the feature aggregation network involved in this step. In this step, the feature aggregation network can perform clustering processing on the filtered predicted feature units according to the feature aggregation decision matrix, thereby obtaining the clustered feature aggregation units. This step can finally summarize the various feature aggregation units to form encoded enhanced filtering features that can be semantically aligned with the sample problem description features.
[0117] As described above, the semantic-aware alignment model used in this embodiment is considered to be a pre-trained model. In one optimized implementation, it is preferable that the semantic-aware alignment model updates network parameters according to a set regularization constraint loss function to achieve model training; and preferably, the loss function value of the regularization constraint loss function is the sum of a triplet loss function value and a regularization norm.
[0118] Specifically, the regularization constraint loss function can be considered as the loss function value used to determine the training of the semantic perception model. This loss function value actually includes two parts, specifically the sum of the two parts. One part is the triplet loss function value, and the other part is the regularization norm.
[0119] In an embodiment, the regularization norm can preferably be determined based on a set regularization parameter and the first-order norm of the similarity score. The similarity score can be considered as the similarity value determined by the semantic-aware alignment model relative to the predicted feature units in each predicted coding enhancement feature during the similarity determination stage. Furthermore, in this embodiment, the triplet loss function value can preferably be determined based on the sample problem description text and the coding enhancement screening features.
[0120] In determining the triplet loss function value, the main considerations are the impact of the sample problem text and its corresponding negative text on the encoded reinforcement filtering features. Based on the above optimization, as an implementation method, the step of determining the triplet loss function value according to the sample problem description text and the encoded reinforcement filtering features can be optimized as follows:
[0121] Obtain negative problem description text randomly sampled relative to the sample problem description text; determine the first cosine distance value between the sample problem description text and the encoded reinforcement filtering feature, and determine the second cosine distance value between the negative problem description text and the encoded reinforcement filtering feature; determine the distance difference between the first cosine distance value and the second cosine distance value, and obtain the sum of the distance difference value and the set hyperparameter; determine the maximum value between the sum and the set value as the triplet loss function value.
[0122] In this embodiment, the negative text corresponding to the sample problem description text can be obtained through random sampling. The obtained sample problem description text and the encoded reinforcement filtering features can determine the cosine distance value, and the negative problem description text and the encoded reinforcement filtering features can also determine the cosine distance value. In this embodiment, the difference between the two cosine distance values can be determined by calculating the difference between them. Then, the difference between the two cosine distance values can be added to a hyperparameter. Finally, the sum can be compared with a set value (such as 0), and the maximum value of the two can be selected as the triplet loss function value.
[0123] It can be seen that the triplet loss function value specifically involves three elements: encoded reinforcement screening features, sample problem description text, and problem description negative text. Therefore, the result value determined through the above steps is denoted as the triplet loss function value.
[0124] It is understood that the semantic awareness alignment model used in this embodiment is trained and then introduced into the training of the audio question answering model to further improve the training accuracy of the audio question answering model.
[0125] d4) Using the question-answering analysis sub-model in the initial audio question-answering model, analyze the encoded reinforcement filtering features and the sample question description features of the sample question text to obtain the predicted question response results associated with the sample audio.
[0126] In this embodiment, after determining the encoded enhancement screening features through step c4) above, the encoded enhancement screening features and the determined sample question description features can be used as input data for the question answering analysis sub-model in the initial audio question answering model. The question answering analysis sub-model is then used to perform question answering analysis, and finally, the predicted question answer results associated with the sample audio output by the initial audio question answering model under one iteration training can be obtained.
[0127] e4) Determine the target loss function value based on the predicted question response results and the actual question response results in the sample triplet.
[0128] It is understood that, in addition to introducing the above-mentioned operation of determining the encoding to strengthen the filtering features in the training of the audio question answering model, this embodiment also uses the loss function determination logic of real information and predicted information. Specifically, it can use the cross-entropy loss function to determine the target loss function value for global training, based on the predicted question answer result and the real question answer result included in the sample triplet.
[0129] f4) Update the network parameters in the initial audio question answering model according to the target loss function value, and return to re-execute step b4) until the training termination condition is met to obtain the trained audio question answering model.
[0130] In this embodiment, the network parameters in the initial audio question-answering model can be updated using the target loss function value. It is also known that training the initial audio question-answering model is an iterative process. After completing one iteration through the above steps, the process can return to step b4 to start the next iteration. Finally, when the training termination condition is met, the trained audio question-answering model can be obtained.
[0131] The technical solution described in this embodiment provides a training implementation for an audio question-answering model. During model training, it simultaneously considers the semantic alignment difficulties caused by the significant differences between audio and text modalities, thereby updating the encoded enhancement features of the input question-answering analysis sub-model. The training method provided in this embodiment effectively refines audio encoded features through semantic awareness, ensuring that the output encoded enhancement features of the trained audio question-answering model are accurately and consistently aligned with the question description features. This also better guarantees the processing accuracy of the audio question-answering model during audio question-answering.
[0132] Figure 2This is a schematic diagram of an audio question-and-answer device provided in an embodiment of the present disclosure. This embodiment is applicable to the situation of eliminating sibilance signals in audio signals. The device can be implemented by software and / or hardware, and can be configured in a terminal and / or server to implement the audio question-and-answer method in this embodiment of the present disclosure. Specifically, the device may include: an acquisition module 21, a first determination module 22, a second determination module 23, a third determination module 24, and a response determination module 25.
[0133] Among them, the acquisition module 21 is used to acquire the audio sequence and the problem description text;
[0134] The first determining module 22 is used to determine the encoded feature sequence of the audio sequence and the problem description features of the problem description text;
[0135] The second determining module 23 is used to determine, based on the coding feature sequence, the task association coding features corresponding to at least one preset audio problem task, and to determine the problem description weight relative to each of the audio problem tasks based on the problem description features.
[0136] The third determining module 24 is used to determine the coding enhancement features associated with the problem description text based on the coding features associated with each task and the corresponding problem description weights.
[0137] The response determination module 25 is used to determine the question response result associated with the audio sequence based on the encoded enhancement features and the question description features.
[0138] This embodiment provides an audio question-answering device that effectively perceives the audio features involved in the question and answer based on the question description, thereby achieving adaptive matching for different question descriptions in audio question-answering scenarios. This technical solution considers the correlation between the encoding features of audio data and different audio question tasks, while also considering the weight of the question description text relative to different audio question tasks. By highlighting the audio question task emphasized by the question description text through the weight results for different audio question tasks, it then considers combining the weight with the encoding features of the audio to strengthen the encoding features related to the emphasized audio question task. Finally, by combining the enhanced encoding features obtained after feature enhancement with the question description text, it is possible to analyze the audio response results associated with the audio sequence that are more suitable for the question description text. Compared with existing technical solutions, this technical solution achieves a comprehensive and effective understanding of the audio sequence, thereby perceiving the audio encoding features that better match the question description text, thus achieving adaptive matching of audio question and answer, and ensuring the accuracy of audio question and answer when faced with different question descriptions in audio question-answering scenarios.
[0139] Furthermore, the first determining module 22 can be used to:
[0140] The audio sequence is encoded using the audio coding layer in the determined audio question-answering model to obtain the coding fusion features of the audio sequence. The audio coding layer includes at least one audio encoder, and each audio encoder is used to encode different types of audio data in the audio sequence.
[0141] The feature alignment layer in the audio question-answering model divides the encoded fusion features into a set number of encoded feature units, and the encoded feature units are then aggregated to form the encoded feature sequence.
[0142] The text extraction layer in the audio question-answering model extracts text features from the question description text to obtain the question description features.
[0143] Furthermore, the second determining module 23 can specifically be used for:
[0144] The encoded feature sequence and the question description features are input into the task awareness module of the constructed audio question answering model;
[0145] The feature processing submodule in the task awareness module processes the encoded feature sequence according to the included feature analysis networks to obtain the task-related encoded features relative to each set audio problem task.
[0146] The task routing submodule in the task awareness module processes the problem description features based on the weight matrices of each task to obtain the problem description weights relative to each audio problem task.
[0147] Each of the feature analysis networks and each of the task weight matrices corresponds to an audio problem task.
[0148] Furthermore, the third determining module 24 can specifically be used for:
[0149] The task-associated encoding features and corresponding problem description weights are input into the weighted operation layer of the constructed audio problem model to obtain weighted encoding features, and the weighted encoding features are determined as the encoding enhancement features associated with the problem description text.
[0150] Furthermore, the response determination module 25 can specifically be used for:
[0151] By analyzing the question-answering analysis sub-model in the determined audio question-answering model, the encoded enhancement features and the question description features are analyzed to obtain the question-answer text features.
[0152] Decode the features of the question-and-response text, generate the question-and-response text, and identify the question-and-response text as the question-and-response result associated with the audio sequence.
[0153] Furthermore, the device may also include a model training module for obtaining an audio question-answering model by training the constructed initial audio question-answering model;
[0154] The model training module includes:
[0155] The acquisition unit is used to acquire a sample training set and a pre-built initial audio question answering model. The sample training set includes at least one sample triplet, which includes sample audio, sample question text, and real question answer results.
[0156] The feature acquisition unit is used to select a sample triplet, input the sample audio and sample question text contained therein into the initial audio question answering model, and obtain the predictive encoding enhancement features output by the task awareness module included in the initial audio question answering model.
[0157] The feature removal unit is used to perform feature removal processing on the predicted encoded enhancement features through a given semantic awareness alignment model to obtain encoded enhancement removal features that are semantically aligned with the sample question text.
[0158] The result prediction unit is used to analyze the encoded reinforcement screening features and the sample question description features of the sample question text using the question answering sub-model in the initial audio question answering model, and to obtain the predicted question answer result associated with the sample audio.
[0159] The loss determination unit is used to determine the target loss function value based on the predicted question response results and the actual question response results in the sample triplet.
[0160] The parameter update unit is used to update the network parameters in the initial audio question answering model according to the target loss function value, and return to the feature acquisition unit to re-execute until the training termination condition is met, so as to obtain the trained audio question answering model.
[0161] Furthermore, the feature removal unit can specifically be used for:
[0162] The predicted encoding enhancement features and the sample problem description features of the sample problem description text are input into the semantic awareness alignment model;
[0163] The similarity score of each predicted feature unit included in the predicted encoding enhancement feature is determined by the similarity determination network in the semantic awareness alignment model, and a feature aggregation decision matrix is generated by combining each similarity score according to a set threshold.
[0164] The feature aggregation network in the semantic awareness alignment model aggregates each predicted feature unit according to the feature aggregation decision matrix to obtain the aggregated feature aggregation unit, and uses each feature aggregation unit to form the encoded reinforcement screening feature.
[0165] Furthermore, the semantic-aware alignment model updates network parameters according to a set regularization constraint loss function;
[0166] The loss function value of the regularization constraint loss function is the sum of a triplet loss function value and a regularization norm;
[0167] The regularization norm is determined based on the set regularization parameters and the first-order norm of the similarity score;
[0168] The triplet loss function value is determined based on the sample problem description text and the encoded reinforcement screening features.
[0169] Furthermore, the feature removal unit also includes a semantic model training subunit, wherein the step of the semantic model training subunit determining the triplet loss function value based on the sample problem description text and the encoded reinforcement removal features includes:
[0170] Obtain negative text of the problem description that is randomly sampled relative to the sample problem description text;
[0171] Determine the first cosine distance value between the sample problem description text and the encoded reinforcement screening feature, and determine the second cosine distance value between the negative problem description text and the encoded reinforcement screening feature;
[0172] Determine the distance difference between the first cosine distance value and the second cosine distance value, and obtain the sum of the distance difference and the set hyperparameter;
[0173] The maximum value between the summed value and the set value is determined as the triplet loss function value.
[0174] The above-described apparatus can execute the methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the methods.
[0175] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0176] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Reference is made below. Figure 3It illustrates a computer device suitable for implementing embodiments of the present disclosure (e.g., Figure 3 The diagram below shows the structure of the terminal device or server 30. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0177] like Figure 3 As shown, the computer device 30 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 31, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 32 or a program loaded from a storage device 38 into a random access memory (RAM) 33. The RAM 33 also stores various programs and data required for the operation of the computer device 30. The processing unit 31, the ROM 32, and the RAM 33 are interconnected via a bus 35. An edit / output (I / O) interface 34 is also connected to the bus 35.
[0178] Typically, the following devices can be connected to I / O interface 34: input devices 36 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 37 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 38 including, for example, magnetic tapes, hard disks, etc.; and communication devices 39. Communication device 39 allows computer device 30 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computer device 30 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.
[0179] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 39, or installed from a storage device 38, or installed from a ROM 32. When the computer program is executed by the processing device 31, it performs the functions defined in the methods of embodiments of this disclosure.
[0180] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0181] The computer device provided in this embodiment and the audio question-and-answer method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0182] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the audio question-and-answer method provided in the above embodiments.
[0183] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0184] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0185] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0186] The aforementioned computer-readable medium may be included in the aforementioned computer device; or it may exist independently and not assembled into the computer device.
[0187] The aforementioned computer-readable medium carries one or more programs that, when executed by the computer device, cause the computer device to:
[0188] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0189] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0190] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0191] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0192] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0193] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to specific combinations of the above-described technical features, but should also cover other technical methods formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical methods formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0194] Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0195] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An audio question-and-answer method, characterized in that, include: Obtain the audio sequence and the problem description text; Determine the encoded feature sequence of the audio sequence, and determine the problem description features of the problem description text; Based on the encoded feature sequence, the task-associated encoded features corresponding to at least one preset audio problem task are determined, and the problem description weights relative to each of the audio problem tasks are determined based on the problem description features. Based on the task-related coding features and the corresponding problem description weights, the coding enhancement features associated with the problem description text are determined. Based on the encoded enhancement features and the question description features, the question response results associated with the audio sequence are determined.
2. The method according to claim 1, characterized in that, The process of determining the encoded feature sequence of the audio sequence and the problem description features of the problem description text includes: The audio sequence is encoded using the audio coding layer in the determined audio question-answering model to obtain the coding fusion features of the audio sequence. The audio coding layer includes at least one audio encoder, and each audio encoder is used to encode different types of audio data in the audio sequence. The feature alignment layer in the audio question-answering model divides the encoded fusion features into a set number of encoded feature units, and the encoded feature units are then aggregated to form the encoded feature sequence. The text extraction layer in the audio question-answering model extracts text features from the question description text to obtain the question description features.
3. The method according to claim 1, characterized in that, The step of determining the task-associated coding features corresponding to at least one preset audio problem task based on the coding feature sequence, and determining the problem description weights relative to each of the audio problem tasks based on the problem description features, includes: The encoded feature sequence and the question description features are input into the task awareness module of the constructed audio question answering model; The feature processing submodule in the task awareness module processes the encoded feature sequence according to the included feature analysis networks to obtain the task-related encoded features relative to each set audio problem task. The task routing submodule in the task awareness module processes the problem description features based on the weight matrices of each task to obtain the problem description weights relative to each audio problem task. Each of the feature analysis networks and each of the task weight matrices corresponds to an audio problem task.
4. The method according to claim 1, characterized in that, The step of determining the coding enhancement features associated with the problem description text based on the coding features associated with each task and the corresponding problem description weights includes: The task-associated encoding features and corresponding problem description weights are input into the weighted operation layer of the constructed audio problem model to obtain weighted encoding features, and the weighted encoding features are determined as the encoding enhancement features associated with the problem description text.
5. The method according to claim 1, characterized in that, The step of determining the question-response result associated with the audio sequence based on the encoded enhancement features and the question description features includes: By analyzing the question-answering analysis sub-model in the determined audio question-answering model, the encoded enhancement features and the question description features are analyzed to obtain the question-answer text features. Decode the features of the question-and-response text, generate the question-and-response text, and identify the question-and-response text as the question-and-response result associated with the audio sequence.
6. The method according to any one of claims 2-5, characterized in that, The audio question-answering model is obtained by training the initial audio question-answering model. The training steps of the audio question-answering model include: Obtain a sample training set and a pre-built initial audio question answering model. The sample training set includes at least one sample triplet, which includes sample audio, sample question text, and real question answer results. Select a sample triplet, input the sample audio and sample question text contained therein into the initial audio question answering model, and obtain the predictive encoding enhancement features output by the task awareness module included in the initial audio question answering model; Using a given semantic awareness alignment model, feature filtering is performed on the predicted coding enhancement features to obtain coding enhancement filtering features that are semantically aligned with the sample question text. Using the question-answering analysis sub-model in the initial audio question-answering model, the encoded reinforcement filtering features and the sample question description features of the sample question text are analyzed to obtain the predicted question-answer results associated with the sample audio. The target loss function value is determined based on the predicted question response results and the actual question response results in the sample triplet. The network parameters in the initial audio question answering model are updated according to the target loss function value, and the operation of inputting sample audio and sample question text into the initial audio question answering model is repeated until the training termination condition is met, and the trained audio question answering model is obtained.
7. The method according to claim 6, characterized in that, The step of performing feature filtering on the predicted encoded enhancement features using a given semantic awareness alignment model to obtain encoded enhancement filtering features that are semantically aligned with the sample question text includes: The predicted encoding enhancement features and the sample problem description features of the sample problem description text are input into the semantic awareness alignment model; The similarity score of each predicted feature unit included in the predicted encoding enhancement feature is determined by the similarity determination network in the semantic awareness alignment model, and a feature aggregation decision matrix is generated by combining each similarity score according to a set threshold. The feature aggregation network in the semantic awareness alignment model aggregates each predicted feature unit according to the feature aggregation decision matrix to obtain the aggregated feature aggregation unit, and uses each feature aggregation unit to form the encoded reinforcement screening feature.
8. The method according to claim 7, characterized in that, The semantic-aware alignment model updates network parameters according to the set regularization constraint loss function; The loss function value of the regularization constraint loss function is the sum of a triplet loss function value and a regularization norm; The regularization norm is determined based on the set regularization parameters and the first-order norm of the similarity score; The triplet loss function value is determined based on the sample problem description text and the encoded reinforcement screening features.
9. The method according to claim 8, characterized in that, The step of determining the triplet loss function value based on the sample problem description text and the encoded reinforcement screening features includes: Obtain negative text of the problem description that is randomly sampled relative to the sample problem description text; Determine the first cosine distance value between the sample problem description text and the encoded reinforcement screening feature, and determine the second cosine distance value between the negative problem description text and the encoded reinforcement screening feature; Determine the distance difference between the first cosine distance value and the second cosine distance value, and obtain the sum of the distance difference and the set hyperparameter; The maximum value between the summed value and the set value is determined as the triplet loss function value.
10. An audio question-and-answer device, characterized in that, include: The acquisition module is used to acquire audio sequences and problem description text; The first determining module is used to determine the encoded feature sequence of the audio sequence and the problem description features of the problem description text; The second determining module is used to determine, based on the coding feature sequence, the task association coding features corresponding to at least one preset audio problem task, and to determine the problem description weight relative to each of the audio problem tasks based on the problem description features. The third determining module is used to determine the coding enhancement features associated with the problem description text based on the coding features associated with each task and the corresponding problem description weights. The response determination module is used to determine the question response result associated with the audio sequence based on the encoded enhancement features and the question description features.
11. A computer device, characterized in that, The computer device includes: One or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the audio question-answering method as described in any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the audio question-and-answer method as described in any one of claims 1-9.
13. A computer program product comprising a computer program that, when executed by a processor, implements the audio question-and-answer method according to any one of claims 1-9.