Audio content recognition method, medium, apparatus, and computing device

CN118748007BActive Publication Date: 2026-09-11HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410733334.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2026-09-11
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

[0005]本公开提供一种音频内容识别方法、介质、装置和计算设备,以解决相关技术中对包含异常信息的语音识别准确性不足、容易出现误判的问题

Benefits of technology

[0024]According to the audio content recognition method, medium, apparatus, and computing device of this disclosure, semantic features and sub-semantic features corresponding to the same audio segment to be recognized are extracted, and the recognition result for the audio segment to be recognized is determined based on the semantic features and sub-semantic features. Therefore, by combining the semantic features and sub-semantic features of the audio segment to be recognized for content recognition, the problem of false positives and false negatives that easily occur when relying solely on sub-semantic features for recognition is effectively avoided. Compared with manual recognition, automated recognition filtering is achieved, thereby enabling the recognition of massive amounts of audio content, freeing up a large amount of manpower required for recognition, and saving labor costs. Compared with ordinary machine recognition, more accurate and more diverse recognition of target content is achieved, effectively applicable to various abnormal information recognition scenarios, and improving the efficiency of audio content recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118748007B_ABST
    Figure CN118748007B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio content recognition method, medium, device and computing device, relating to the technical field of neural networks. The audio content recognition method comprises: extracting semantic features and sub-semantic features corresponding to the same piece of to-be-recognized audio, and determining a recognition result for the to-be-recognized audio based on the semantic features and the sub-semantic features. The method of the present disclosure solves the problems of insufficient speech recognition accuracy and easy misjudgment in the related art containing abnormal information, effectively avoids the misjudgment and omission problems that are prone to occur when simply recognizing sub-semantic features, realizes automatic recognition filtering compared with manual recognition, so that the recognition of massive audio content can be realized, a large amount of manpower required for recognition is liberated, and manpower cost is saved. Compared with ordinary machine recognition, more accurate and more scene-specific accurate recognition of target content is realized, which is effectively applied to various abnormal information recognition scenarios, and the audio content recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of neural network technology, and more specifically, the embodiments of this disclosure relate to an audio content recognition method, medium, apparatus, and computing device. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be related art simply because it is included in this section.

[0003] In the field of audio detection technology, accurately identifying abnormal information in audio (such as unusual or sensitive words, or noise that is likely to cause discomfort) is an important application area. This type of speech content recognition containing sensitive information mainly includes two approaches: manual recognition and machine recognition. Manual recognition is less efficient and has higher labor costs; therefore, machine recognition is receiving increasing research and application.

[0004] In related technologies, machine recognition of speech containing abnormal information is mainly based on the judgment of parasemantic features that are unrelated to the text carried by the speech. However, this recognition method is prone to misjudging breathing sounds and background noises in normal speech as abnormal information, resulting in insufficient accuracy of the recognition results and difficulty in optimization. Summary of the Invention

[0005] This disclosure provides an audio content recognition method, medium, apparatus, and computing device to address the problems of insufficient accuracy and susceptibility to misjudgment in related technologies for speech recognition containing abnormal information.

[0006] In a first aspect of this disclosure, an audio content recognition method is provided, comprising:

[0007] Semantic features and para-semantic features are extracted, and the semantic features and para-semantic features correspond to the same audio segment to be identified;

[0008] Based on semantic features and para-semantic features, the recognition result for the audio to be recognized is determined; wherein, the recognition result indicates that the audio to be recognized contains the target recognition content; or, the recognition result indicates that the audio to be recognized does not contain the target recognition content.

[0009] In a second aspect of this disclosure, an audio content recognition model training method is provided, comprising:

[0010] The sample audio used for model training is input into the audio content recognition model, and the predicted sample content corresponding to the sample audio is output.

[0011] The audio content recognition model is trained based on the predicted sample content and the sample audio content corresponding to the sample audio.

[0012] In a third aspect of this disclosure, a computer-readable storage medium is provided, comprising:

[0013] The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the audio content recognition method of the first aspect of the present disclosure; and / or, when executed by a processor, are used to implement the audio content recognition model training method of the second aspect of the present disclosure.

[0014] In a fourth aspect of this disclosure, an audio content recognition device is provided, the device comprising:

[0015] The extraction module is used to extract semantic features and sub-semantic features, which correspond to the same audio segment to be identified.

[0016] The recognition module is used to determine the recognition result for the audio to be recognized based on semantic features and para-semantic features; wherein the recognition result indicates that the audio to be recognized contains the target recognition content; or, the recognition result indicates that the audio to be recognized does not contain the target recognition content.

[0017] In a fifth aspect of this disclosure, an audio content recognition model training apparatus is provided, the apparatus comprising:

[0018] The input module is used to input sample audio for model training into the audio content recognition model and output the predicted sample content corresponding to the sample audio.

[0019] The training module is used to train the audio content recognition model based on the predicted sample content and the sample audio content corresponding to the sample audio.

[0020] In a sixth aspect of this disclosure, a computing device is provided, comprising:

[0021] At least one processor;

[0022] and memory that is communicatively connected to at least one processor;

[0023] The memory stores instructions executable by at least one processor, which are executed by at least one processor to cause the computing device to perform the audio content recognition method as described in the first aspect of this disclosure; and / or to cause the computing device to perform the audio content recognition model training method as described in the second aspect of this disclosure.

[0024] According to the audio content recognition method, medium, apparatus, and computing device of this disclosure, semantic features and sub-semantic features corresponding to the same audio segment to be recognized are extracted, and the recognition result for the audio segment to be recognized is determined based on the semantic features and sub-semantic features. Therefore, by combining the semantic features and sub-semantic features of the audio segment to be recognized for content recognition, the problem of false positives and false negatives that easily occur when relying solely on sub-semantic features for recognition is effectively avoided. Compared with manual recognition, automated recognition filtering is achieved, thereby enabling the recognition of massive amounts of audio content, freeing up a large amount of manpower required for recognition, and saving labor costs. Compared with ordinary machine recognition, more accurate and more diverse recognition of target content is achieved, effectively applicable to various abnormal information recognition scenarios, and improving the efficiency of audio content recognition. Attached Figure Description

[0025] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:

[0026] Figure 1 An application scenario diagram illustrating an embodiment of the present disclosure is shown schematically;

[0027] Figure 2 A flowchart illustrating an audio content recognition method according to another embodiment of the present disclosure is shown schematically;

[0028] Figure 3a A flowchart illustrating an audio content recognition method according to another embodiment of the present disclosure is shown schematically;

[0029] Figure 3b schematically shown Figure 3a The diagram shows the structure of the audio content recognition model provided in the embodiment shown.

[0030] Figure 4 A flowchart illustrating an audio content recognition model training method according to another embodiment of the present disclosure is shown schematically;

[0031] Figure 5a A flowchart illustrating an audio content recognition model training method according to another embodiment of the present disclosure is shown schematically;

[0032] Figure 5b schematically shown Figure 5a A schematic diagram of the structure of the audio content recognition model to be trained provided in the embodiment shown;

[0033] Figure 5c schematically shown Figure 5aThe flowchart illustrates the processing method of sample semantic features and sample para-semantic features in the feature recognition part of the audio content recognition model provided in the embodiment.

[0034] Figure 5d schematically shown Figure 5a The flowchart of the mask processing method provided in the embodiment shown is shown.

[0035] Figure 5e schematically shown Figure 5a The flowchart of the mask processing method based on mask probability variation provided in the embodiment shown;

[0036] Figure 6 A flowchart illustrating an audio content recognition model training method according to another embodiment of the present disclosure is shown schematically.

[0037] Figure 7 A schematic diagram of the structure of a storage medium according to another embodiment of the present disclosure is shown;

[0038] Figure 8 A schematic diagram of the structure of an audio content recognition device according to another embodiment of the present disclosure is shown.

[0039] Figure 9 A schematic diagram of the structure of an audio content recognition model training apparatus according to another embodiment of the present disclosure is shown.

[0040] Figure 10 A schematic diagram of the structure of a computing device according to another embodiment of the present disclosure is shown.

[0041] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0042] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0043] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0044] According to embodiments of this disclosure, an audio content recognition method, medium, apparatus, and computing device are proposed.

[0045] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0046] The following is a description of the terminology used in this disclosure:

[0047] Semantic features: In this scheme, these are used to represent the corresponding features of the parts of the audio to be identified that can be converted into text.

[0048] Parasemantic features: In this scheme, these are used to represent the corresponding features of the parts of the audio to be identified that cannot be converted into text or whose meaning cannot be determined as text.

[0049] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0050] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0051] Furthermore, the number of any elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning. Invention Overview

[0053] The inventors have discovered that accurately identifying anomalous information (such as sensitive words or disturbing noise) in audio is an important application area in the field of audio detection technology, enabling the labeling and masking of audio containing such anomalous information. The recognition of speech containing anomalous information mainly involves two approaches: manual recognition and machine recognition. Manual recognition is less efficient and more costly, therefore machine recognition is receiving increasing research and application.

[0054] In related technologies, machine recognition of whether speech contains abnormal information is mainly based on the judgment of parasemantic features that are unrelated to the text carried by the speech. However, this recognition method is prone to misjudging breathing sounds and background noises in normal speech as abnormal information, resulting in insufficient accuracy of the recognition results and difficulty in optimization.

[0055] In this scheme, by combining semantic features and para-semantic features for recognition, the content in the speech to be recognized can be identified through multimodal correspondence, so as to more accurately determine whether it contains the target content and improve the accuracy and efficiency of the audio content recognition process.

[0056] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.

[0057] Application Scenarios Overview

[0058] First refer to Figure 1 As shown, during the audio content recognition process, the server 100 inputs the audio to be recognized 110 into the audio content recognition model 120. The audio content recognition model 120 outputs the corresponding recognition result 130 based on the semantic features and sub-semantic features of the audio to be recognized 110. Based on the recognition result 130, it determines whether the audio to be recognized 110 contains the target recognition content, thereby completing the audio content recognition process.

[0059] It should be noted that, Figure 1 In the scenario shown, only one of the publishing server, the audio to be identified, the audio content recognition model, and the recognition result is used for illustrative purposes. However, this disclosure is not limited to this. That is to say, the number of servers, the audio to be identified, the audio content recognition model, and the recognition result can be arbitrary.

[0060] Exemplary methods

[0061] The following is combined with Figure 1 Application scenarios, refer to Figures 2 to 3a To describe an audio content recognition method according to exemplary embodiments of the present disclosure, referencing Figures 4 to 6 This document describes an audio content recognition model training method according to exemplary embodiments of the present disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in any way. Rather, the embodiments of the present disclosure can be applied to any applicable scenario.

[0062] Figure 2 This is a flowchart illustrating an audio content recognition method provided in one embodiment of this disclosure. Figure 2 As shown, the audio content recognition method includes the following steps:

[0063] Step S201: Extract semantic features and sub-semantic features.

[0064] Among them, semantic features and sub-semantic features correspond to the same audio segment to be identified.

[0065] Specifically, this method is mainly applicable to scenarios where it is necessary to determine whether there is a specific type of abnormal information (i.e. target recognition content) in the audio, such as identifying whether the audio uploaded by the user may contain target recognition content. If the audio is found to contain the corresponding type of target recognition content, the audio needs to be filtered, blocked, and labeled.

[0066] In the above scenario, the audio that needs to be identified is the audio to be identified.

[0067] To identify whether the audio to be identified contains the target content, it is necessary to extract its corresponding semantic features and sub-semantic features.

[0068] Specifically, the method for extracting semantic features can employ a speech recognition algorithm that converts speech into text to obtain the text corresponding to the audio to be recognized, and then extract the semantic features corresponding to the text. The speech recognition algorithm can be an open-source speech recognition algorithm in the prior art, or it can be a non-open-source speech recognition algorithm; no limitation is made in this disclosure embodiment.

[0069] Specifically, the method for extracting secondary semantic features can employ acoustic feature recognition algorithms used for audio recognition to extract acoustic feature information, and then extract the required secondary semantic features from the acoustic feature information. The acoustic feature recognition algorithm can be an open-source recognition algorithm in the prior art, or it can be a non-open-source acoustic feature recognition algorithm; no limitation is made in this embodiment.

[0070] Step S202: Based on semantic features and sub-semantic features, determine the recognition result for the audio to be recognized.

[0071] The recognition result indicates that the audio to be recognized contains the target content; or, the recognition result indicates that the audio to be recognized does not contain the target content.

[0072] Specifically, by obtaining the semantic features and sub-semantic features of the audio to be identified, the semantic features and sub-semantic features are used to jointly determine whether the audio to be identified contains the target content.

[0073] For example, by using semantic features, it is possible to quickly determine whether the audio to be identified contains abnormal or sensitive keywords and expressions corresponding to abnormal information; by using para-semantic features, it is possible to quickly determine whether the audio to be identified contains pronunciation, tone and other para-semantic features that belong to abnormal information, and thus determine whether the audio to be identified contains the target content; by combining semantic features and para-semantic features for joint judgment, it is possible to judge some content that is easy to miss when judging solely from semantic features or para-semantic features (for example, a very normal voice in the corresponding text may contain some uncomfortable screams or noises. It is difficult to determine whether it belongs to abnormal information from the text alone, but by combining semantic features and para-semantic features for joint recognition, it can be judged to belong to abnormal information).

[0074] Therefore, by combining semantic and parasemantic features, the recognition result for the audio to be recognized can be automatically determined. Based on the recognition result, corresponding tags can be automatically added to the audio to be recognized, so that operations such as masking or backtracking can be performed, thereby significantly improving the recognition accuracy and efficiency.

[0075] According to the audio content recognition method of this disclosure, semantic features and secondary semantic features corresponding to the same audio segment to be recognized are extracted, and the recognition result for the audio segment to be recognized is determined based on the semantic features and secondary semantic features. Therefore, by combining the semantic features and secondary semantic features of the audio segment to be recognized for content recognition, the method effectively avoids the false positives and false negatives that are easily encountered when relying solely on secondary semantic features for recognition. Compared to manual recognition, it achieves automated recognition filtering, thereby enabling the recognition of massive amounts of audio content, freeing up a significant amount of manpower, and saving labor costs. Compared to ordinary machine recognition, it achieves more accurate and more diverse recognition of target content, effectively applicable to various abnormal information recognition scenarios, and improving the efficiency of audio content recognition.

[0076] Figure 3a This is a flowchart illustrating an audio content recognition method provided in one embodiment of this disclosure. Figure 3a As shown, it includes the following steps:

[0077] Step S301: Input the audio to be recognized into the audio content recognition model, and extract semantic features and sub-semantic features through the feature extraction part of the audio content recognition model.

[0078] Specifically, in this embodiment, the recognition result is mainly generated through an audio content recognition model. That is, the audio content recognition model can extract semantic features and para-semantic features, and obtain the corresponding recognition result based on the semantic features and para-semantic features.

[0079] The audio content recognition model provided in this embodiment includes a feature extraction part and a feature recognition part. The feature extraction part is used to extract semantic features and sub-semantic features, and the feature recognition part is used to obtain the corresponding recognition results based on the semantic features and sub-semantic features.

[0080] like Figure 3b The diagram shows the structure of an audio content recognition model. The audio content recognition model 300 includes a feature extraction section 310 and a feature recognition section 320. The feature extraction section 310 includes a first feature extraction module 311, a second feature extraction module 312, and a third feature extraction module 313. The third feature extraction module 313 includes a first feature extraction unit 314 (including a transformer encoder 315 and a first self-attention mechanism layer 316) that cooperates with the first feature extraction module 311, and a second feature extraction unit 317 (including a transformer encoder 318 and a second self-attention mechanism layer 319) that cooperates with the second feature extraction module 312. The feature recognition section 320 includes a fusion feature recognition module 321 (including a transformer decoder 322 and a self-attention pooling layer 323), a soft cue word module 324, and a classifier 325.

[0081] Furthermore, the arrows in the figure are used to indicate the direction of data transmission. The soft prompt word module 324 can also be included in the fusion feature recognition module 321, and the classifier 325 can also be a fully connected layer.

[0082] The working principle of the audio content recognition model described above will be further described below.

[0083] In one embodiment of this disclosure, the secondary semantic features include at least one of emotional features, timbre features, and tone features.

[0084] Specifically, firstly, the secondary semantic features referred to in this disclosure embodiment may correspond to the emotional features, timbre features, and tone features in the audio to be identified. These features may all correspond to the target content to be identified (for example, speaking normal text content with a special tone, timbre, and pitch can be identified as abnormal or sensitive information).

[0085] In one embodiment of this disclosure, the secondary semantic features also include phase information in the audio to be identified.

[0086] Specifically, the acoustic features included in the parasemantic features include common acoustic features such as FBank features (Fliter Banks) and MFCC features (Mel-Frequency Cepstral Coefficients), as well as phase information such as pitch, intensity, and duration. By combining phase information, it is possible to better identify the emotion, timbre, and tone features in the parasemantic features.

[0087] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a first feature extraction module, which is used to extract secondary semantic features based on the input audio to be recognized.

[0088] Specifically, the feature extraction part of the audio recognition module includes two feature extraction modules: a first feature extraction module for extracting secondary semantic features and a second feature extraction module for extracting semantic features. When the audio to be recognized is input into the audio content recognition model, the audio recognition module extracts semantic features and secondary semantic features through the first and second feature extraction modules, respectively.

[0089] In one embodiment of this disclosure, the first feature extraction module is a Hubert model obtained through self-supervised pre-training.

[0090] Specifically, the Hubert model is a self-supervised acoustic feature extraction model trained on speech. By selecting a self-supervised pre-trained model, it can be ensured that the model used for feature extraction has undergone pre-training on a large amount of unlabeled data, thereby guaranteeing the stability and reliability of the model.

[0091] Those skilled in the art may also choose other models capable of acoustic feature extraction (including convolutional neural networks, although the extraction results of different models will differ) instead of using the Hubert model as the first feature extraction module, without any restrictions here.

[0092] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a second feature extraction module, which is used to extract semantic features based on the input text content, where the text content corresponds to the audio to be recognized.

[0093] Specifically, the second feature extraction module is used to extract the semantic features of the text corresponding to the audio to be recognized. In this embodiment, the audio to be recognized can be directly input into the second feature extraction module and the corresponding semantic features can be obtained directly. In practical applications, the text corresponding to the audio to be recognized can be extracted first through the speech-text extraction model, and then the semantic features corresponding to the text can be extracted through the second feature extraction module (at this time, the speech-text extraction model is also included in the feature extraction part of the audio content recognition model).

[0094] In one embodiment of this disclosure, the second feature extraction module is a BERT model obtained through self-supervised pre-training.

[0095] Specifically, as mentioned above, those skilled in the art may also choose other models capable of semantic feature extraction instead of using the BERT model as the second feature extraction module (the semantic feature extraction effects of different models will vary), and no restrictions are imposed here.

[0096] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a third feature extraction module, which includes a first feature extraction unit and a second feature extraction unit. The first feature extraction unit is used to extract target sub-semantic features associated with the target recognition content based on the sub-semantic features output by the first feature extraction module. The second feature extraction unit is used to extract target semantic features associated with the target recognition content based on the semantic features output by the second feature extraction module.

[0097] Specifically, the first feature extraction module and the second feature extraction module obtain the sub-semantic features and semantic features of the entire audio to be identified. Since the target recognition content (corresponding semantic features and sub-semantic features) may only account for a part of the content in the audio to be identified (the audio to be identified may also contain meaningless blank spaces, white noise, etc.), it is necessary to further extract the parts related to the target recognition content from these sub-semantic features and semantic features, namely the target sub-semantic features and target semantic features.

[0098] In the audio content recognition model, the third feature extraction module processes the sub-semantic features and semantic features to obtain the target sub-semantic features and target semantic features.

[0099] In one embodiment of this disclosure, the third feature extraction module includes two sets of transformer encoders, which are used to encode the sub-semantic features and semantic features input to the third feature extraction module, respectively. The third feature extraction module also includes a self-attention mechanism based on acoustic feature modeling and a self-attention mechanism based on semantic feature modeling, so as to process the encoded sub-semantic features and semantic features, respectively, extract the partial features corresponding to the target recognition content, and thus obtain the target sub-semantic features and target semantic features.

[0100] Step S302: Input the semantic features and sub-semantic features into the feature recognition part of the audio content recognition model, and determine the recognition result based on the output of the audio content recognition model.

[0101] Specifically, the output results based on semantic features and para-semantic features are mainly obtained through the feature recognition part of the audio content recognition model, which will be explained in detail below.

[0102] In one embodiment of this disclosure, the target identification content is audio content containing abnormal information.

[0103] Specifically, abnormal information can be the target content of sensitive or abnormal words. This type of abnormal information can be directly identified through semantic features. Abnormal information can also be relatively ordinary information in the text, such as sensitive words that are jarring or have blurred or changing background sounds. Blurred sensitive words are easy to misidentify, while changed sensitive words are difficult to identify, and jarring background sounds are relatively more hidden. This type of abnormal information needs to be identified by combining semantic features and para-semantic features. Therefore, the following content in this step will mainly take the abnormal information that needs to be identified by combining semantic features and para-semantic features as an example for further explanation.

[0104] In one embodiment of this disclosure, the feature recognition part of the audio content recognition model includes a fusion feature recognition module, which is used to recognize target fusion features to determine that the audio to be recognized contains corresponding target recognition content; wherein, the target fusion feature is obtained by performing feature fusion processing on target sub-semantic features and target semantic features.

[0105] Specifically, after obtaining the target sub-semantic features and target semantic features, it is possible to determine whether the target recognition content is contained. The specific judgment action is performed by the fusion feature recognition module.

[0106] Before recognition is performed by the fusion feature recognition module, the target sub-semantic features and the target semantic features need to be fused (to obtain the target fusion features). The specific fusion processing method can be to directly concatenate the target sub-semantic features and the target semantic features, or to combine them based on cross-attention mechanisms, etc. There are no restrictions here.

[0107] In one embodiment of this disclosure, the fusion feature recognition module is a Q-Former model; the Q-Former model includes a Transformer decoder and a self-attention pooling layer; wherein, the Transformer decoder is used to perform feature fusion processing on the target sub-semantic features and the target semantic features, and the self-attention pooling layer is used to recognize the target fused features after feature fusion processing.

[0108] Specifically, the target fusion feature recognition process of the fusion feature recognition module can be completed by the Q-Former model. The Transformer decoder is used to receive the target sub-semantic features and the target semantic features and obtain the target fusion features through fusion processing. The self-attention pooling layer is used to output the corresponding recognition result based on the target fusion features (or, output the recognition result features and then input the recognition result features into a classifier or fully connected layer, and the classifier or fully connected layer outputs the final recognition result).

[0109] In one embodiment of this disclosure, the Q-Former model further includes a soft cue word module, which provides soft cue words related to the target recognition content so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

[0110] Specifically, to enhance recognition effectiveness and accuracy, when generating target fusion features, soft cue words are combined with target para-semantic features and target semantic features to jointly generate target fusion features. This results in the generated target fusion features containing more feature information related to the features corresponding to the target recognition content. When the content corresponding to the target recognition content in the audio to be recognized is not obvious, adding soft cue words can enhance recognition accuracy (because the more features related to the target recognition content are included, the less likely there will be feature omissions in the recognition judgment of such features, and the higher the accuracy of recognizing that it contains the target recognition content).

[0111] According to the audio content recognition method of this disclosure, the audio to be recognized is input into an audio content recognition model. Semantic features and para-semantic features are extracted by the feature extraction part of the audio content recognition model. These semantic and para-semantic features are then input into the feature recognition part of the audio content recognition model. Based on the output of the audio content recognition model, the recognition result is determined. Therefore, the audio content recognition model can determine whether the audio to be recognized contains the target content. Furthermore, the audio content recognition model can simultaneously combine the semantic and para-semantic features of the audio to be recognized for joint judgment, significantly improving the accuracy and efficiency of audio content recognition.

[0112] Figure 4 This is a flowchart illustrating an audio content recognition model training method provided in one embodiment of this disclosure. Figure 4 As shown, it includes the following steps:

[0113] Step S401: Input the sample audio used for model training into the audio content recognition model and output the predicted sample content corresponding to the sample audio.

[0114] Specifically, since the input content of the audio content recognition model is the audio to be recognized and the output content is the recognition result, the training of the audio content recognition model mainly involves inputting the sample audio into the audio content recognition model to be trained, obtaining the predicted sample content output by the audio content recognition model to be trained, and then feeding back the parameters within the audio content recognition model based on the predicted sample content.

[0115] In one embodiment of this disclosure, before inputting sample audio for model training into the audio content recognition model for training, the audio content recognition model training method further includes determining the structure of the audio content recognition model. The structure of the audio content recognition model can be found in [reference needed]. Figure 3a The structure described in the illustrated embodiment (i.e., the structure including the feature extraction part and the feature recognition part).

[0116] Step S402: Train the audio content recognition model based on the predicted sample content and the sample content corresponding to the sample audio.

[0117] Specifically, training the audio content recognition model with a large number of sample audio files can effectively ensure the accuracy and reliability of the trained audio content recognition model in identifying target content.

[0118] In one embodiment of this disclosure, the predicted sample content can be the judgment result of whether the sample audio contains the target recognition content. At this time, the sample content of the sample audio is the label of whether it contains the target recognition content. By comparing the judgment result with the label, the audio content recognition model is trained.

[0119] According to the audio content recognition model training method of this disclosure, sample audio for model training is input into the audio content recognition model, the predicted sample content corresponding to the sample audio is output, and the audio content recognition model is trained based on the predicted sample content and the sample content corresponding to the sample audio. Thus, an audio content recognition model that can determine whether the audio to be recognized contains the target recognition content can be obtained, thereby achieving automatic recognition of the audio to be recognized and improving the efficiency of audio content recognition.

[0120] Figure 5a This is a flowchart illustrating an audio content recognition model training method provided in one embodiment of this disclosure. Figure 5a As shown, it includes the following steps:

[0121] Step S501: Input the sample audio used for model training into the audio content recognition model, and extract the sample semantic features and sample secondary semantic features corresponding to the sample audio through the feature extraction part of the audio content recognition model.

[0122] Specifically, before inputting sample audio into the audio content recognition model, the structure of the audio content recognition model needs to be determined.

[0123] like Figure 5b As shown, this is a schematic diagram of the structure of the audio content recognition model to be trained. The audio content recognition model 500 includes a feature extraction part 510 and a feature recognition part 520, wherein the feature extraction part 510 and... Figure 3b The audio content recognition model 300 is the same as that in the audio content recognition model 300, including a first feature extraction module 511, a second feature extraction module 512 and a third feature extraction module 513. The third feature extraction module 513 includes a first feature extraction unit 514 (including a transformer encoder 515 and a first self-attention mechanism layer 516) that cooperates with the first feature extraction module 511 and a second feature extraction unit 517 (including a transformer encoder 518 and a second self-attention mechanism layer 519) that cooperates with the second feature extraction module 512. The feature recognition part 520 includes a fusion feature recognition module 521 (including a transformer decoder 522 and a self-attention pooling layer 525), a soft cue word module 524 and a classifier 525, and also includes a random modality masking unit 526.

[0124] Furthermore, the arrows in the figure are used to indicate the direction of data transmission. The soft prompt word module 524 and the random modality mask unit 526 can both be included in the fusion feature recognition module 521, and the classifier 525 can also be a fully connected layer.

[0125] The following section further describes the working principle of the training phase based on the audio content recognition model described above.

[0126] Once the structure of the audio content recognition model is determined, sample audio can be input into the audio content recognition model to train it.

[0127] The specific sample audio is processed in the audio content recognition model and Figure 3a The audio content recognition models in the embodiments have largely the same processing steps for the audio to be recognized. They all first extract the corresponding semantic features and sub-semantic features through the feature extraction part, that is, the sample semantic features and sample sub-semantic features corresponding to the sample audio. The main difference lies in the processing of the feature recognition part of the audio content recognition model.

[0128] Step S502: Input the sample semantic features and sample sub-semantic features into the feature recognition part of the audio content recognition model, and output the predicted sample content corresponding to the sample audio.

[0129] The feature recognition section includes a fusion feature recognition module.

[0130] Specifically, in the feature recognition part of the audio content recognition model, when processing the semantic features and para-semantic features of the samples, relative to the model application stage (i.e. Figure 3a The corresponding process in the illustrated embodiment has been enhanced with a random modal masking unit to randomly mask the sample semantic features and sample sub-semantic features, thereby ensuring the robustness of the trained audio content recognition model. This part will be described in detail below.

[0131] Furthermore, such as Figure 5c The diagram shown is a flowchart illustrating the processing method of sample semantic features and sample para-semantic features in the feature recognition part of the audio content recognition model. It includes the following:

[0132] Step S5021: Mask the sample semantic features and sample sub-semantic features using random modal masking units.

[0133] Specifically, the random modal masking unit, in conjunction with the Q-Former model, performs random masking on the sample semantic features and sample sub-semantic features before fusing them.

[0134] Furthermore, such as Figure 5d The diagram shown is a flowchart of the masking process, which includes the following steps:

[0135] Step A1: Based on the set training conditions, determine the mask probability of the output of the random modal mask unit.

[0136] Specifically, the random modality masking unit randomly selects one from the sample semantic features and sample sub-semantic features for masking. The masking probability here refers to the probability of whether or not to mask the selected feature; therefore, the masking probability needs to be determined first.

[0137] Among them, the training conditions can be that the audio content recognition model reaches a set number of training rounds, or that the recognition accuracy of the predicted sample content reaches a set accuracy threshold. The recognition accuracy is obtained by comparing the predicted sample content with the sample content.

[0138] Specifically, during actual training, the initial audio content recognition model may focus only on features of one modality. To enable the model to simultaneously focus on features of different modalities in the initial stage, we need to randomly mask one modality's features. As the training process iterates, the model stabilizes and maintains focus on both modalities, thus eliminating the need to mask features of one modality. This stable model maintains focus on both modalities based on the established training conditions.

[0139] The training conditions can be the number of training rounds, the recognition accuracy, or both (training stops when either condition is met, or when both conditions are met).

[0140] Step A2: Mask the sample semantic features and sample sub-semantic features based on the mask probability.

[0141] Specifically, after randomly selecting a feature from the sample semantic features and sample sub-semantic features for processing, the decision is made, based on the masking probability, whether to mask the selected feature. This masking process involves how the masking probability changes, which will be explained further below.

[0142] Furthermore, such as Figure 5e The diagram shows a flowchart of a masking process based on changes in mask probability, which includes the following steps:

[0143] Step B1: Randomly select one feature from the sample semantic features and sample sub-semantic features as the feature to be masked.

[0144] Specifically, random masking involves two layers of randomness: one layer is the random selection of a feature from the sample semantic features and the sample sub-semantic features, and the other layer is the randomness of whether or not to perform masking.

[0145] When randomly selecting a feature from the sample semantic feature and the sample sub-semantic feature, the probability of selecting the sample semantic feature and the sample sub-semantic feature can be 50%, that is, the probability of selecting the sample semantic feature or the sample sub-semantic feature each time is 50%, in order to avoid the bias of the audio content recognition model in the process of processing feature data.

[0146] The process of selecting features to be masked can be completed by a random modal masking unit. After inputting the sample semantic features and sample sub-semantic features into the random modal masking unit, the masked features will be obtained. Alternatively, the process of selecting features to be masked can be completed in the transformer decoder in the Q-Former model. In this case, the random modal masking unit will output the masking information (such as the information on whether to mask the features to be masked, which is also selected in the random modal masking unit, but the sample semantic features and sample sub-semantic features are not input into the random modal masking unit) and the sample semantic features and sample sub-semantic features together to the transformer decoder, where the masking and fusion processes are completed.

[0147] Step B2: Based on the mask probability, mask the features to be masked.

[0148] Specifically, after randomly selecting a feature as the feature to be masked each time, the masking probability is used to determine whether to mask the feature in that instance.

[0149] In one embodiment of this disclosure, the mask probability is negatively correlated with the number of training rounds of the audio content recognition model.

[0150] Specifically, as analyzed above, as the training process iterates, the model's focus on features of different modalities becomes more stable, and the need for masking decreases. Therefore, the masking probability decreases with increasing training iterations. The specific rate of decrease can be adjusted according to the configuration of the training conditions, which will not be elaborated here.

[0151] Step B3: If the masking probability is zero, do not mask the features to be masked.

[0152] Specifically, when the mask probability is zero, no further masking is performed. In actual training, the mask probability can also be non-zero, but rather kept at a low value (such as 1%, 5%, or 10%). In this case, the robustness of the trained model can be better guaranteed.

[0153] Whether to stop masking the features to be masked during the training phase can be chosen based on the actual situation. That is, step B3 technology will choose whether to execute it based on the actual situation.

[0154] Step S5022: Input the soft prompts output by the soft prompt module, the masked sample semantic features, and the masked sample secondary semantic features into the Transformer decoder to obtain the prediction fusion features output by the Transformer decoder.

[0155] Specifically, soft prompts can be input into the Transformer decoder along with the masked sample semantic features and sample sub-semantic features, or the masked information can be input along with soft prompts, sample semantic features, and sample sub-semantic features; there is no restriction here. Both methods achieve the goal of obtaining the predicted fusion features output by the Transformer decoder.

[0156] Step S5023: Input the predicted fusion features into the self-attention pooling layer and output the predicted sample content.

[0157] Specifically, after the predicted fusion features are input into the self-attention pooling layer, the predicted sample content or the corresponding features will be output. If the output is the corresponding features, it will be passed through a classifier (or a fully connected layer) and then output the predicted sample content.

[0158] Step S5024: Based on the predicted sample content and the sample content, train the soft cue word module, the Transformer decoder and the self-attention pooling layer.

[0159] Specifically, based on the predicted sample content and the sample content, the parameters in the audio content recognition model to be trained can be trained, including the parameters in the first feature extraction module, the second feature extraction module, the third feature extraction module of the feature extraction part, and the fusion feature recognition module (including the soft cue word module, the Transformer decoder and the self-attention pooling layer) of the feature recognition part, so as to obtain the trained audio content recognition model.

[0160] According to the audio content recognition model training method of this disclosure, sample audio for model training is input into the audio content recognition model. The feature extraction part of the audio content recognition model extracts the sample semantic features and sample para-semantic features corresponding to the sample audio. Then, the sample semantic features and sample para-semantic features are input into the feature recognition part of the audio content recognition model, and the predicted sample content corresponding to the sample audio is output. This ensures that the trained audio content recognition model can robustly combine the semantic features and para-semantic features in the audio to be recognized to jointly determine the corresponding recognition result, thereby ensuring the accuracy and reliability of its recognition of the target content, and thus ensuring the recognition efficiency of the target content.

[0161] Figure 6 This is a flowchart illustrating an audio content recognition model training method provided in one embodiment of this disclosure. Figure 6 As shown, it includes the following steps:

[0162] Step S601: Input the sample audio into the first feature extraction module to obtain the predicted secondary semantic features corresponding to the sample audio.

[0163] Specifically, before inputting sample audio into the audio content recognition model, it is necessary to determine the structure of the audio content recognition model, that is, to determine that the audio content recognition model includes the first feature extraction module, the second feature extraction module, the third feature extraction module of the feature extraction part, and the fusion feature recognition module (including the soft cue word module, the Transformer decoder and the self-attention pooling layer) of the feature recognition part, and then input the sample audio into the first feature extraction module for processing.

[0164] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a first feature extraction module, which is used to extract predicted secondary semantic features based on the input sample audio.

[0165] Specifically, during the training phase, the first feature extraction module extracts the corresponding predicted secondary semantic features based on the input sample audio, corresponding to the application phase (see...). Figure 3a The illustrated embodiment describes the process of obtaining secondary semantic features based on the audio to be identified.

[0166] In one embodiment of this disclosure, the first feature extraction module is a Hubert model that has been pre-trained under self-supervised supervision.

[0167] Specifically, during the training phase, it is necessary to first determine the first feature extraction module and the second feature extraction module to be used. Both of these modules can directly select self-supervised pre-trained models, such as the Hubert model. This can significantly reduce the training requirements for these two modules, and the output results of the corresponding models can be directly used as predicted secondary semantic features and predicted semantic features for subsequent training.

[0168] Step S602: Input the sample audio into the second feature extraction module to obtain the predicted semantic features corresponding to the sample audio.

[0169] Specifically, the process of obtaining the predicted semantic features corresponds to obtaining the predicted secondary semantic features, which will be described below.

[0170] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a second feature extraction module, which is used to extract predicted semantic features based on the sample text corresponding to the input sample audio.

[0171] Specifically, during the training phase, the second feature extraction module extracts the corresponding predicted semantic features based on the input sample audio, corresponding to the application phase (see...). Figure 3a The process of obtaining semantic features based on the audio to be identified (as shown in the embodiment) will not be described in detail here.

[0172] In one embodiment of this disclosure, the second feature extraction module is a BERT model obtained through self-supervised pre-training.

[0173] Specifically, both the BERT and Hubert models are open-source feature extraction models with high accuracy and reliability. Therefore, these two models can be used directly to perform the functions of the corresponding modules.

[0174] Step S603: Input the predicted secondary semantic features into the first feature extraction unit to obtain the sample secondary semantic features.

[0175] Specifically, after obtaining the predicted secondary semantic features, they are input into the subsequent first feature extraction unit to obtain the sample secondary semantic features (analogous to...). Figure 3a The target secondary semantic features are obtained in the illustrated embodiment.

[0176] In one embodiment of this disclosure, the feature extraction part of the audio content recognition model includes a third feature extraction module, which includes a first feature extraction unit and a second feature extraction unit. The first feature extraction unit is used to extract sample sub-semantic features associated with the predicted recognition content based on the predicted sub-semantic features output by the first feature extraction module. The second feature extraction unit is used to extract sample semantic features associated with the predicted recognition content based on the predicted semantic features output by the second feature extraction module.

[0177] Specifically, the specific structure of the third feature extraction module is as follows: Figure 3a The illustrated embodiment describes a process that further extracts features based on the features output by the first and second feature extraction modules. This process is consistent with the application stage (i.e.,...). Figure 3a The process of extracting the target sub-semantic features from the sub-semantic features and extracting the target semantic features from the semantic features in the illustrated embodiment is the same, and will not be described again here.

[0178] In one embodiment of this disclosure, the first feature extraction unit includes a Transformer encoder and a first self-attention mechanism layer based on predicted sub-semantic features; wherein, the Transformer encoder is used to encode the predicted sub-semantic features, and the first self-attention mechanism layer is used to extract sample sub-semantic features based on the encoded predicted sub-semantic features.

[0179] Specifically, in practical applications, the audio to be identified often contains segments that do not contain actual information, such as a portion of white noise or meaningless blank parts (such as the silent parts before and after the audio). Through the first self-attention mechanism layer, the sample sub-semantic features can be concentrated to reflect the features in the audio content after excluding these segments without actual information, thereby improving the accuracy of subsequent feature recognition.

[0180] Step S604: Input the predicted semantic features into the second feature extraction unit to obtain the sample semantic features.

[0181] Specifically, similarly, after obtaining the predicted semantic features, they are input into the subsequent second feature extraction unit to obtain the sample semantic features (analogous to...). Figure 3a The target semantic features obtained in the illustrated embodiment.

[0182] In one embodiment of this disclosure, the second feature extraction unit includes a Transformer encoder and a second self-attention mechanism layer based on predicted semantic features; wherein, the Transformer encoder is used to encode the predicted semantic features, and the second self-attention mechanism layer is used to extract sample semantic features based on the encoded predicted semantic features.

[0183] Specifically, similar to the first self-attention mechanism layer, the second self-attention mechanism layer also makes the semantic features of the obtained samples more focused on the parts of the audio to be identified that contain actual content, rather than the blank and meaningless parts of the audio to be identified.

[0184] Step S605: Based on the predicted sample content and the sample content, train the first feature extraction unit and the second feature extraction unit.

[0185] Specifically, Figure 5a The training in the embodiments mainly targets the feature recognition part of the audio content recognition model. In this embodiment, however, the training primarily focuses on the feature extraction part, specifically the first and second feature extraction units. In actual training, both the feature extraction and feature recognition parts are trained simultaneously, rather than separately.

[0186] According to the audio content recognition model training method of this disclosure, after determining the structure of the audio content recognition model, sample audio is input into a first feature extraction module to obtain the predicted secondary semantic features corresponding to the sample audio. The sample audio is then input into a second feature extraction module to obtain the predicted semantic features corresponding to the sample audio. The predicted secondary semantic features are then input into a first feature extraction unit to obtain sample secondary semantic features, and the predicted semantic features are input into a second feature extraction unit to obtain sample semantic features. Thus, based on the predicted sample content and the sample content, the first and second feature extraction units are trained. This ensures the accuracy and reliability of the trained audio content recognition model in feature extraction, thereby guaranteeing the accuracy and reliability of audio content recognition.

[0187] Exemplary media

[0188] After introducing the methods of exemplary embodiments of this disclosure, the following references are made. Figure 7 The storage medium of the exemplary embodiments of this disclosure will be described.

[0189] refer to Figure 7 As shown, a program product 70 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.

[0190] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0191] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.

[0192] Program code for performing the operations disclosed herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).

[0193] Exemplary device

[0194] Having introduced the medium of exemplary embodiments of this disclosure, the following references are made to... Figure 8 and Figure 9 An audio content recognition apparatus according to an exemplary embodiment of this disclosure will be described, wherein... Figure 8 The audio content recognition device shown is used to achieve Figures 2 to 3a The audio content recognition method in the illustrated embodiment, Figure 9 The audio content recognition device shown is used to achieve Figures 4 to 6 The audio content recognition method in the illustrated embodiment is similar in principle and technical effect to the aforementioned corresponding method embodiment, and will not be repeated here.

[0195] like Figure 8 As shown, the audio content recognition device 800 provided in this disclosure includes:

[0196] Extraction module 810 is used to extract semantic features and sub-semantic features, which correspond to the same audio segment to be identified;

[0197] The recognition module 820 is used to determine the recognition result for the audio to be recognized based on semantic features and para-semantic features; wherein the recognition result indicates that the audio to be recognized contains the target recognition content; or, the recognition result indicates that the audio to be recognized does not contain the target recognition content.

[0198] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: inputting the audio to be identified into an audio content recognition model, and extracting semantic features and sub-semantic features through the feature extraction part of the audio content recognition model; the recognition module 820 specifically includes: inputting the semantic features and sub-semantic features into the feature recognition part of the audio content recognition model, and determining the recognition result based on the output result of the audio content recognition model.

[0199] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: the feature extraction part of the audio content recognition model includes a first feature extraction module, which is used to extract sub-semantic features based on the input audio to be recognized.

[0200] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: a first feature extraction module is a Hubert model obtained through self-supervised pre-training.

[0201] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: the feature extraction part of the audio content recognition model includes a second feature extraction module, which is used to extract semantic features based on the input text content, where the text content corresponds to the audio to be recognized.

[0202] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: a second feature extraction module is a BERT model obtained through self-supervised pre-training.

[0203] In an exemplary embodiment of this disclosure, the extraction module 810 specifically includes: the feature extraction part of the audio content recognition model includes a third feature extraction module, the third feature extraction module includes a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract target sub-semantic features associated with the target recognition content based on the sub-semantic features output by the first feature extraction module; the second feature extraction unit is used to extract target semantic features associated with the target recognition content based on the semantic features output by the second feature extraction module.

[0204] In one exemplary embodiment of this disclosure, the recognition module 820 specifically includes: the feature recognition part of the audio content recognition model includes a fusion feature recognition module, which is used to recognize the target fusion features to determine that the audio to be recognized contains the corresponding target recognition content; wherein, the target fusion features are obtained by performing feature fusion processing on the target sub-semantic features and the target semantic features.

[0205] In one exemplary embodiment of this disclosure, the recognition module 820 specifically includes: a fusion feature recognition module is a Q-Former model; the Q-Former model includes a Transformer decoder and a self-attention pooling layer; wherein, the Transformer decoder is used to perform feature fusion processing on the target sub-semantic features and the target semantic features, and the self-attention pooling layer is used to recognize the target fused features after feature fusion processing.

[0206] In one exemplary embodiment of this disclosure, the recognition module 820 specifically includes: the Q-Former model further includes a soft cue word module, which is used to provide soft cue words related to the target recognition content so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

[0207] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: the subsemantic features include at least one of emotion features, timbre features, and tone features.

[0208] In one exemplary embodiment of this disclosure, the extraction module 810 specifically includes: the sub-semantic features also contain phase information in the audio to be identified.

[0209] In one exemplary embodiment of this disclosure, the identification module 820 specifically includes: the target identification content is audio content containing abnormal information.

[0210] like Figure 9 As shown, the audio content recognition model training device 900 provided in this disclosure includes:

[0211] The input module 910 is used to input sample audio for model training into the audio content recognition model and output the predicted sample content corresponding to the sample audio.

[0212] Training module 920 is used to train the audio content recognition model based on the predicted sample content and the sample content corresponding to the sample audio.

[0213] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: an audio content recognition model including a fusion feature recognition module, the fusion feature recognition module including a random modality masking unit, wherein the random modality masking unit is effective during the training phase of the audio content recognition model, but not effective during the usage phase of the audio content recognition model.

[0214] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: a fusion feature recognition module, which is a Q-Former model; the Q-Former model includes a soft cue word module, a Transformer decoder, and a self-attention pooling layer; wherein, the Transformer decoder is used to perform feature fusion processing on the predicted sub-semantic features and the predicted semantic features, the self-attention pooling layer is used to recognize the predicted fusion features processed by feature fusion, and the soft cue word module is used to provide soft cue words related to the target recognition content, so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

[0215] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: the feature fusion processing method includes: fusing the predicted sub-semantic features and the predicted semantic features based on a cross-attention mechanism; or, concatenating the predicted sub-semantic features and the predicted semantic features.

[0216] In one exemplary embodiment of this disclosure, the input module 910 is specifically configured to: input sample audio for model training into the audio content recognition model; extract sample semantic features and sample sub-semantic features corresponding to the sample audio through the feature extraction part of the audio content recognition model; input the sample semantic features and sample sub-semantic features into the feature recognition part of the audio content recognition model; and output the predicted sample content corresponding to the sample audio. The feature recognition part includes a fusion feature recognition module.

[0217] In an exemplary embodiment of this disclosure, the input module 910 is specifically configured to: perform masking processing on the sample semantic features and sample sub-semantic features through a random modality masking unit; input the soft prompts output by the soft prompt module, the masked sample semantic features, and the masked sample sub-semantic features into the Transformer decoder to obtain the prediction fusion features output by the Transformer decoder; input the prediction fusion features into a self-attention pooling layer to output the predicted sample content; and the training module 920 is specifically configured to: train the soft prompt module, the Transformer decoder, and the self-attention pooling layer based on the predicted sample content and the sample content.

[0218] In one exemplary embodiment of this disclosure, the input module 910 is specifically used to: determine the mask probability output by the random modality masking unit based on set training conditions; and perform masking processing on the sample semantic features and sample sub-semantic features based on the mask probability.

[0219] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: the training rounds of the audio content recognition model reaching a set training round, or the recognition accuracy of the predicted sample content reaching a set accuracy threshold, wherein the recognition accuracy is obtained based on the comparison between the predicted sample content and the sample content.

[0220] In one exemplary embodiment of this disclosure, the input module 910 is specifically configured to: randomly select a feature from the sample semantic features and the sample sub-semantic features as the feature to be masked; perform masking processing on the feature to be masked based on the masking probability; or, if the masking probability is zero, not perform masking processing on the feature to be masked.

[0221] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: the mask probability is negatively correlated with the training rounds of the audio content recognition model.

[0222] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: the feature extraction part of the audio content recognition model includes a first feature extraction module, which is used to extract predicted secondary semantic features based on the input sample audio.

[0223] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: the feature extraction part of the audio content recognition model includes a second feature extraction module, which is used to extract predicted semantic features based on the sample text corresponding to the input sample audio.

[0224] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: a second feature extraction module is a BERT model obtained through self-supervised pre-training.

[0225] In an exemplary embodiment of this disclosure, the input module 910 specifically includes: a feature extraction part of the audio content recognition model including a third feature extraction module, the third feature extraction module including a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract sample sub-semantic features associated with the predicted recognition content based on the predicted sub-semantic features output by the first feature extraction module; the second feature extraction unit is used to extract sample semantic features associated with the predicted recognition content based on the predicted semantic features output by the second feature extraction module.

[0226] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: a first feature extraction unit including a Transformer encoder and a first self-attention mechanism layer based on predicted sub-semantic features; wherein, the Transformer encoder is used to encode the predicted sub-semantic features, and the first self-attention mechanism layer is used to extract sample sub-semantic features based on the encoded predicted sub-semantic features.

[0227] In one exemplary embodiment of this disclosure, the input module 910 specifically includes: a second feature extraction unit including a Transformer encoder and a second self-attention mechanism layer based on predicted semantic features; wherein, the Transformer encoder is used to encode the predicted semantic features, and the second self-attention mechanism layer is used to extract sample semantic features based on the encoded predicted semantic features.

[0228] In an exemplary embodiment of this disclosure, the input module 910 is specifically configured to: input sample audio into a first feature extraction module to obtain predicted secondary semantic features corresponding to the sample audio; input sample audio into a second feature extraction module to obtain predicted semantic features corresponding to the sample audio; input predicted secondary semantic features into a first feature extraction unit to obtain sample secondary semantic features; and input predicted semantic features into a second feature extraction unit to obtain sample semantic features. The training module 920 is specifically configured to: train the first feature extraction unit and the second feature extraction unit based on the predicted sample content and the sample content.

[0229] Exemplary computing device

[0230] Having described the methods, media, and apparatus of exemplary embodiments of this disclosure, the following references... Figure 10 A computing device according to an exemplary embodiment of the present disclosure will be described.

[0231] Figure 10 The computing device 100 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0232] like Figure 10 As shown, the computing device 100 is presented in the form of a general-purpose computing device. The components of the computing device 100 may include, but are not limited to: at least one processing unit 1001, at least one storage unit 1002, and a bus 1003 connecting different system components (including the processing unit 1001 and the storage unit 1002).

[0233] Bus 1003 includes a data bus, a control bus, and an address bus.

[0234] Storage unit 1002 may include readable media in the form of volatile memory, such as random access memory (RAM) 10021 and / or cache memory 10022, and may further include readable media in the form of non-volatile memory, such as read-only memory (ROM) 10023.

[0235] Storage unit 1002 may also include a program / utility 10025 having a set (at least one) program module 10024, such program module 10024 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0236] The computing device 100 can also communicate with one or more external devices 1004 (e.g., keyboard, pointing device, etc.). This communication can be performed via the input / output (I / O) interface 1005. Furthermore, the computing device 100 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via a network adapter 10010. Figure 10 As shown, network adapter 10010 communicates with other modules of computing device 100 via bus 1003. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0237] It should be noted that although several units / modules or sub-units / modules of the supply chain strategy determination device and the object scoring model training device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0238] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0239] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. An audio content recognition method, characterized in that, include: The audio to be identified is input into the audio content recognition model, and semantic features and sub-semantic features are extracted through the feature extraction part of the audio content recognition model. The semantic features and the sub-semantic features correspond to the same audio segment to be identified. The semantic features and the secondary semantic features are input into the feature recognition part of the audio content recognition model, and the recognition result for the audio to be recognized is determined based on the output of the audio content recognition model; wherein, the recognition result indicates that the audio to be recognized contains the target recognition content; or, the recognition result indicates that the audio to be recognized does not contain the target recognition content. The feature extraction part of the audio content recognition model includes a first feature extraction module, a second feature extraction module, and a third feature extraction module. The third feature extraction module includes a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract target secondary semantic features associated with the target recognition content based on the secondary semantic features output by the first feature extraction module; the second feature extraction unit is used to extract target semantic features associated with the target recognition content based on the semantic features output by the second feature extraction module. The feature recognition part of the audio content recognition model includes a fusion feature recognition module, which is a Q-Former model. The Q-Former model includes a Transformer decoder and a self-attention pooling layer. The Transformer decoder is used to perform feature fusion processing on the target sub-semantic features and the target semantic features, and the self-attention pooling layer is used to recognize the target fused features after feature fusion processing.

2. The audio content recognition method according to claim 1, characterized in that, The first feature extraction module is used to extract the secondary semantic features based on the input audio to be identified.

3. The audio content recognition method according to claim 2, characterized in that, The first feature extraction module is a Hubert model obtained through self-supervised pre-training.

4. The audio content recognition method according to claim 2, characterized in that, The second feature extraction module is used to extract the semantic features based on the input text content, where the text content corresponds to the audio to be identified.

5. The audio content recognition method according to claim 4, characterized in that, The second feature extraction module is a BERT model obtained through self-supervised pre-training.

6. The audio content recognition method according to claim 5, characterized in that, The fusion feature recognition module is used to recognize the target fusion features to determine that the audio to be recognized contains the corresponding target recognition content; wherein, the target fusion feature is obtained by performing feature fusion processing on the target sub-semantic feature and the target semantic feature.

7. The audio content recognition method according to claim 1, characterized in that, The Q-Former model also includes a soft cue word module, which provides soft cue words related to the target recognition content so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

8. The audio content recognition method according to any one of claims 1 to 7, characterized in that, The secondary semantic features include at least one of emotional features, timbre features, and tone features.

9. The audio content recognition method according to claim 8, characterized in that, The secondary semantic features also include phase information in the audio to be identified.

10. The audio content recognition method according to any one of claims 1 to 7, characterized in that, The target identification content is audio content containing abnormal information.

11. A method for training an audio content recognition model, characterized in that, include: The sample audio used for model training is input into the audio content recognition model, and the sample semantic features and sample secondary semantic features corresponding to the sample audio are extracted through the feature extraction part of the audio content recognition model; The sample semantic features and the sample secondary semantic features are input into the feature recognition part of the audio content recognition model, and the predicted sample content corresponding to the sample audio is output. The feature recognition part includes a fusion feature recognition module. The feature extraction part of the audio content recognition model includes a first feature extraction module, a second feature extraction module, and a third feature extraction module. The first feature extraction module is used to extract predicted secondary semantic features based on the input sample audio. The second feature extraction module is used to extract predicted semantic features based on the sample text corresponding to the input sample audio. The third feature extraction module includes a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract sample secondary semantic features associated with the predicted recognition content based on the predicted secondary semantic features output by the first feature extraction module; the second feature extraction unit is used to extract sample semantic features associated with the predicted recognition content based on the predicted semantic features output by the second feature extraction module. The audio content recognition model is trained based on the predicted sample content and the sample content corresponding to the sample audio. The fusion feature recognition module is a Q-Former model; the Q-Former model includes a Transformer decoder and a self-attention pooling layer; wherein, the Transformer decoder is used to perform feature fusion processing on the predicted sub-semantic features and the predicted semantic features, and the self-attention pooling layer is used to recognize the predicted fusion features processed by feature fusion.

12. The audio content recognition model training method according to claim 11, characterized in that, The fusion feature recognition module includes a random modality masking unit, which is effective during the training phase of the audio content recognition model but not during the usage phase of the audio content recognition model.

13. The audio content recognition model training method according to claim 12, characterized in that, The Q-Former model also includes a soft cue word module; the soft cue word module is used to provide soft cue words related to the target recognition content, so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

14. The audio content recognition model training method according to claim 11, characterized in that, The feature fusion processing methods include: The predicted secondary semantic features and the predicted semantic features are fused based on a cross-attention mechanism; Alternatively, the predicted secondary semantic features and the predicted semantic features can be concatenated.

15. The audio content recognition model training method according to claim 13, characterized in that, The step of inputting the sample semantic features and the sample secondary semantic features into the feature recognition part of the audio content recognition model, and outputting the predicted sample content corresponding to the sample audio, includes: The random modal masking unit is used to mask the sample semantic features and the sample sub-semantic features; The soft prompts output by the soft prompt module, together with the masked sample semantic features and the masked sample secondary semantic features, are input into the Transformer decoder to obtain the prediction fusion features output by the Transformer decoder. The predicted fusion features are input into the self-attention pooling layer, and the predicted sample content is output. The step of training the audio content recognition model based on the predicted sample content and the sample content corresponding to the sample audio includes: Based on the predicted sample content and the sample content, the soft cue word module, Transformer decoder, and self-attention pooling layer are trained.

16. The audio content recognition model training method according to claim 15, characterized in that, The step of masking the sample semantic features and the sample sub-semantic features using the random modality masking unit includes: Based on the set training conditions, the mask probability output by the random modal masking unit is determined; Based on the mask probability, the sample semantic features and the sample sub-semantic features are masked.

17. The audio content recognition model training method according to claim 16, characterized in that, The training conditions include: The audio content recognition model has reached the set number of training rounds. Alternatively, the recognition accuracy of the predicted sample content reaches a set accuracy threshold, and the recognition accuracy is obtained by comparing the predicted sample content with the sample content.

18. The audio content recognition model training method according to claim 17, characterized in that, The masking process for the sample semantic features and the sample sub-semantic features based on the mask probability includes: Randomly select one feature from the sample semantic features and the sample sub-semantic features as the feature to be masked; Based on the masking probability, the feature to be masked is masked. Alternatively, if the masking probability is zero, the feature to be masked is not masked.

19. The audio content recognition model training method according to claim 18, characterized in that, The mask probability is negatively correlated with the number of training rounds of the audio content recognition model.

20. The audio content recognition model training method according to claim 11, characterized in that, The first feature extraction module is a Hubert model that has been pre-trained under self-supervised supervision.

21. The audio content recognition model training method according to claim 20, characterized in that, The second feature extraction module is a BERT model obtained through self-supervised pre-training.

22. The audio content recognition model training method according to claim 11, characterized in that, The first feature extraction unit includes a Transformer encoder and a first self-attention mechanism layer based on predicted secondary semantic features; wherein, the Transformer encoder is used to encode the predicted secondary semantic features, and the first self-attention mechanism layer is used to extract the sample secondary semantic features based on the encoded predicted secondary semantic features.

23. The audio content recognition model training method according to claim 11, characterized in that, The second feature extraction unit includes a Transformer encoder and a second self-attention mechanism layer based on predicted semantic features; wherein, the Transformer encoder is used to encode the predicted semantic features, and the second self-attention mechanism layer is used to extract the sample semantic features based on the encoded predicted semantic features.

24. The audio content recognition model training method according to claim 23, characterized in that, The audio content recognition model involves inputting sample audio for model training into the audio content recognition model, and extracting sample semantic features and sample secondary semantic features corresponding to the sample audio through the feature extraction part of the audio content recognition model. The audio content recognition model includes: The sample audio is input into the first feature extraction module to obtain the predicted secondary semantic features corresponding to the sample audio; The sample audio is input into the second feature extraction module to obtain the predicted semantic features corresponding to the sample audio; The predicted secondary semantic features are input into the first feature extraction unit to obtain the sample secondary semantic features; The predicted semantic features are input into the second feature extraction unit to obtain the sample semantic features; The step of training the audio content recognition model based on the predicted sample content and the sample content corresponding to the sample audio includes: Based on the predicted sample content and the sample content, the first feature extraction unit and the second feature extraction unit are trained.

25. A computer-readable storage medium comprising: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the audio content recognition method as described in any one of claims 1 to 10; and / or, when executed by a processor, the computer-executable instructions are used to implement the audio content recognition model training method as described in any one of claims 11 to 24.

26. An audio content recognition device, characterized in that, The device includes: The extraction module is used to input the audio to be identified into the audio content recognition model, and extract semantic features and sub-semantic features through the feature extraction part of the audio content recognition model. The semantic features and the sub-semantic features correspond to the same audio segment to be identified. The recognition module is used to input the semantic features and the sub-semantic features into the feature recognition part of the audio content recognition model, and determine the recognition result for the audio to be recognized based on the output result of the audio content recognition model; wherein, the recognition result indicates that the audio to be recognized contains the target recognition content; or, the recognition result indicates that the audio to be recognized does not contain the target recognition content. The extraction module specifically includes: the feature extraction part of the audio content recognition model includes a first feature extraction module, a second feature extraction module, and a third feature extraction module. The third feature extraction module includes a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract the target secondary semantic features associated with the target recognition content based on the secondary semantic features output by the first feature extraction module; the second feature extraction unit is used to extract the target semantic features associated with the target recognition content based on the semantic features output by the second feature extraction module. The recognition module specifically includes: a fusion feature recognition module based on a Q-Former model; the Q-Former model includes a Transformer decoder and a self-attention pooling layer; wherein, the Transformer decoder is used to perform feature fusion processing on the target sub-semantic features and the target semantic features, and the self-attention pooling layer is used to recognize the target fused features after feature fusion processing.

27. The audio content recognition device according to claim 26, characterized in that, The extraction module specifically includes: the feature extraction part of the audio content recognition model includes a first feature extraction module, which is used to extract secondary semantic features based on the input audio to be recognized.

28. The audio content recognition device according to claim 27, characterized in that, The first feature extraction module is a Hubert model obtained through self-supervised pre-training.

29. The audio content recognition device according to claim 27, characterized in that, The extraction module specifically includes: the feature extraction part of the audio content recognition model includes a second feature extraction module, which is used to extract semantic features based on the input text content, where the text content corresponds to the audio to be recognized.

30. The audio content recognition device according to claim 29, characterized in that, The extraction module specifically includes: the second feature extraction module is a BERT model obtained through self-supervised pre-training.

31. The audio content recognition device according to claim 26, characterized in that, The recognition module specifically includes: the feature recognition part of the audio content recognition model includes a fusion feature recognition module, which is used to recognize the target fusion features to determine whether the audio to be recognized contains the corresponding target recognition content; wherein, the target fusion features are obtained by performing feature fusion processing on the target sub-semantic features and the target semantic features.

32. The audio content recognition device according to claim 26, characterized in that, The recognition module specifically includes: The Q-Former model also includes a soft cue word module, which is used to provide soft cue words related to the target recognition content so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

33. The audio content recognition device according to any one of claims 26 to 32, characterized in that, The extraction module specifically includes: parasemantic features, including at least one of sentiment features, timbre features, and tone features.

34. The audio content recognition device according to claim 32, characterized in that, The extraction module specifically includes: the secondary semantic features also contain phase information in the audio to be identified.

35. The audio content recognition device according to any one of claims 26 to 32, characterized in that, The recognition module specifically includes: the target recognition content is audio content containing abnormal information.

36. An audio content recognition model training device, characterized in that, The device includes: The input module is used to input sample audio for model training into the audio content recognition model, and extract and output the sample semantic features and sample secondary semantic features corresponding to the sample audio through the feature extraction part of the audio content recognition model; The sample semantic features and sample para-semantic features are input into the feature recognition part of the audio content recognition model, and the predicted sample content corresponding to the sample audio is output. The feature recognition part includes a fusion feature recognition module. The training module is used to train the audio content recognition model based on the predicted sample content and the sample content corresponding to the sample audio. The feature extraction part of the audio content recognition model includes a first feature extraction module, a second feature extraction module, and a third feature extraction module. The first feature extraction module is used to extract predicted secondary semantic features based on the input sample audio. The second feature extraction module is used to extract predicted semantic features based on the sample text corresponding to the input sample audio. The third feature extraction module includes a first feature extraction unit and a second feature extraction unit, wherein: the first feature extraction unit is used to extract sample secondary semantic features associated with the predicted recognition content based on the predicted secondary semantic features output by the first feature extraction module; the second feature extraction unit is used to extract sample semantic features associated with the predicted recognition content based on the predicted semantic features output by the second feature extraction module. The fusion feature recognition module is a Q-Former model, which includes a Transformer decoder and a self-attention pooling layer. The Transformer decoder is used to perform feature fusion processing on the predicted sub-semantic features and the predicted semantic features, and the self-attention pooling layer is used to recognize the predicted fusion features processed by feature fusion.

37. The audio content recognition model training device according to claim 36, characterized in that, The fusion feature recognition module includes a random modality masking unit, which is effective during the training phase of the audio content recognition model but not during the usage phase.

38. The audio content recognition model training device according to claim 37, characterized in that, The Q-Former model also includes a soft cue word module; the soft cue word module is used to provide soft cue words related to the target recognition content, so that the Transformer decoder can perform feature fusion processing on the target sub-semantic features and the target semantic features.

39. The audio content recognition model training device according to claim 37, characterized in that, The input module specifically includes the following: the feature fusion processing methods include: fusing the predicted sub-semantic features and the predicted semantic features based on the cross-attention mechanism; or concatenating the predicted sub-semantic features and the predicted semantic features.

40. The audio content recognition model training device according to claim 36, characterized in that, The input module is specifically used to: mask the sample semantic features and sample sub-semantic features through a random modality masking unit; input the soft prompts output by the soft prompt module, along with the masked sample semantic features and the masked sample sub-semantic features, into the Transformer decoder to obtain the predicted fusion features output by the Transformer decoder; input the predicted fusion features into a self-attention pooling layer to output the predicted sample content; the training module is specifically used to: train the soft prompt module, the Transformer decoder, and the self-attention pooling layer based on the predicted sample content and the sample content.

41. The audio content recognition model training device according to claim 40, characterized in that, The input module is specifically used to: determine the mask probability of the output of the random modality masking unit based on the set training conditions; and perform masking processing on the sample semantic features and sample sub-semantic features based on the mask probability.

42. The audio content recognition model training device according to claim 41, characterized in that, The input module specifically includes: the audio content recognition model reaching a set number of training rounds, or the recognition accuracy of the predicted sample content reaching a set accuracy threshold, with the recognition accuracy obtained by comparing the predicted sample content with the sample content.

43. The audio content recognition model training device according to claim 42, characterized in that, Specifically, the input module is used to: randomly select one feature from the sample semantic features and sample sub-semantic features as the feature to be masked; Based on the masking probability, the features to be masked are masked; or, if the masking probability is zero, the features to be masked are not masked.

44. The audio content recognition model training device according to claim 42, characterized in that, The input module specifically includes a mask probability that is negatively correlated with the number of training rounds of the audio content recognition model.

45. The audio content recognition model training device according to claim 36, characterized in that, The input module specifically includes: the second feature extraction module is a BERT model obtained through self-supervised pre-training.

46. ​​The audio content recognition model training device according to claim 36, characterized in that, The input module specifically includes: a first feature extraction unit comprising a Transformer encoder and a first self-attention mechanism layer based on predicted sub-semantic features; wherein, the Transformer encoder is used to encode the predicted sub-semantic features, and the first self-attention mechanism layer is used to extract sample sub-semantic features based on the encoded predicted sub-semantic features.

47. The audio content recognition model training device according to claim 46, characterized in that, The input module specifically includes: a second feature extraction unit comprising a Transformer encoder and a second self-attention mechanism layer based on predicted semantic features; wherein, the Transformer encoder is used to encode the predicted semantic features, and the second self-attention mechanism layer is used to extract sample semantic features based on the encoded predicted semantic features.

48. The audio content recognition model training device according to claim 47, characterized in that, The input module is specifically used for: inputting sample audio into the first feature extraction module to obtain the predicted secondary semantic features corresponding to the sample audio; inputting sample audio into the second feature extraction module to obtain the predicted semantic features corresponding to the sample audio; inputting the predicted secondary semantic features into the first feature extraction unit to obtain the sample secondary semantic features; and inputting the predicted semantic features into the second feature extraction unit to obtain the sample semantic features. The training module is specifically used to train the first feature extraction unit and the second feature extraction unit based on the predicted sample content and the sample content.

49. A computing device, comprising: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions executable by at least one processor, which are executed by at least one processor to cause the computing device to perform the audio content recognition method as described in any one of claims 1 to 10; and / or to cause the computing device to perform the audio content recognition model training method as described in any one of claims 11 to 24.

Citation Information

Patent Citations

  • Abnormal language detection method, electronic equipment and storage medium

    CN117493495A