Data processing method and device, electronic equipment and storage medium

By constructing and enhancing the sound effects dataset, the problem of insufficient sample data quality in the sound effects generation model was solved, achieving higher quality sound effects generation and improving the user experience.

CN120769114BActive Publication Date: 2026-01-23BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511259472.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-01-23
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In existing technologies, the quality of sample data used to train sound effect generation models is insufficient, resulting in low-quality generated sound effects that fail to meet users' demands for realistic sound.

Method used

The first sound effect dataset is constructed and an initial sound effect description text is generated. The initial sound effect description text is then processed using data augmentation techniques to generate a target sound effect description text. A second sound effect dataset is then constructed for training the sound effect generation model, including data cleaning, data preprocessing, and data augmentation steps.

Benefits of technology

The quality of the sound effect generation model has been improved, making the generated sound effects more accurate, meeting user needs, and enhancing the human-computer interaction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769114B_ABST
    Figure CN120769114B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, electronic equipment and storage medium, belonging to the technical field of artificial intelligence. For the sound effect generation scene, the present disclosure can construct high-quality sound effect data for sound effect generation model training. In detail, considering that the user may have multiple input forms in the request sound effect generation stage, the present scheme introduces a data enhancement step. In the data enhancement stage, the present scheme can perform data enhancement on the sound effect description text generated for each sound effect data based on the possible input form of the user in the request sound effect generation stage. Since the sound effect description text with more accurate content can be obtained through data enhancement, in the request sound effect generation stage, no matter which form of input the user performs, the sound effect generation model can output the sound effect adapted to the user's demand, ensuring the generation quality of the sound effect generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Sound effect generation is a technology that uses technical means to create, synthesize, or simulate sound effects in a specific scene. It is widely used in film, games, virtual reality / augmented reality, or interactive media, with the aim of enhancing the immersion and expressiveness of content through sound and meeting users' needs for realism in the sound dimension.

[0003] With the development of artificial intelligence technology, most current methods employ sound effect generation models to generate sound effects. Since the quality of the sample data used to train these models directly determines the quality of the generated sound effects, how to construct high-quality sample data to train these models has become a pressing problem for those skilled in the art. Summary of the Invention

[0004] This disclosure provides a data processing method, apparatus, electronic device, and storage medium. The technical solution of this disclosure is shown below.

[0005] According to a first aspect of the present disclosure, a data processing method is provided, comprising:

[0006] Construct a first sound effect dataset and generate initial sound effect description text for the sound effect data included in the first sound effect dataset;

[0007] Based on the possible input formats of the user during the request to generate sound effects, the initial sound effect description text is augmented to obtain the target sound effect description text.

[0008] Based on the sound effect data included in the first sound effect dataset and the target sound effect description text, a second sound effect dataset is constructed for training the sound effect generation model.

[0009] In some embodiments, the step of performing data augmentation on the initial sound effect description text based on possible user input forms during the request-to-generate sound effect stage to obtain the target sound effect description text includes:

[0010] Under given constraints, based on the possible input forms of the user during the request to generate sound effects, the initial sound effect description text is augmented with data to obtain the target sound effect description text.

[0011] The constraints are used to control the elements included in the target sound effect description text.

[0012] In other embodiments, the constraints include one or more of the following:

[0013] The output audio effect description text should include an audio quality description for each audio effect data segment;

[0014] The output sound effect description text should include an audio category description for each sound effect data segment;

[0015] The output sound effect description text should include a detailed description of the sound effect for each sound effect data segment;

[0016] The output sound effect description text should include a detailed description of the background music for each sound effect data segment.

[0017] In other embodiments, the input format includes keyword input;

[0018] The step of performing data augmentation on the initial sound effect description text based on the possible input forms of the user during the request-to-generate sound effect stage to obtain the target sound effect description text includes:

[0019] Add keyword descriptions to the initial sound effect description text.

[0020] In other embodiments, the input format further includes timing input;

[0021] The step of performing data augmentation on the initial sound effect description text based on the possible input forms of the user during the request-to-generate sound effect stage to obtain the target sound effect description text includes:

[0022] In the first sound effect dataset, filter first type of sound effect data and second type of sound effect data; wherein, the first type of sound effect data is used to describe a first type of event with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe a second type of event with a sound effect duration less than a second threshold; the first threshold is greater than the second threshold;

[0023] Each time, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data respectively;

[0024] After combining the two selected sound effect data segments, a sound effect description text including temporal relationships is generated for the combined sound effect data based on the initial sound effect description text of the two sound effect data segments.

[0025] In other embodiments, the two audio effect data segments include first audio effect data and second audio effect data; the first audio effect data comes from the first type of audio effect data, and the second audio effect data comes from the second type of audio effect data; the method further includes:

[0026] The first sound effect data is divided into multiple time periods, and the second sound effect data is inserted into different time periods of the first sound effect data to obtain multiple combined sound effect data.

[0027] Add the combined audio data of the multiple segments to the first audio data set.

[0028] In other embodiments, the input format further includes no text input; the step of performing data augmentation on the initial sound effect description text based on the possible input formats of the user during the request-to-generate sound effect phase to obtain the target sound effect description text includes:

[0029] Under the condition of meeting the preset duration limit, at least two audio effect data segments in the first audio effect dataset are spliced ​​together to obtain spliced ​​audio effect data.

[0030] Based on the initial sound effect description text of the at least two sound effect data segments, a sound effect description text including temporal relationships is generated for the concatenated sound effect data.

[0031] In other embodiments, constructing the first sound effects dataset includes:

[0032] The first sound effect dataset is constructed based on open-source shared sound effect data and unlabeled sound effect data;

[0033] The types of sound effect data in the first sound effect dataset include:

[0034] Audio data including sound effects;

[0035] Video data including sound effects.

[0036] In other embodiments, constructing the first sound effect dataset based on open-source shared sound effect data and unlabeled sound effect data includes:

[0037] If the shared sound effect data is website data, candidate videos are filtered from the website data based on preset tag text to obtain a candidate video set;

[0038] If the shared sound effect data comes from multiple open-source datasets, establish a label mapping rule, and convert the original labels of the sound effect data included in the multiple open-source datasets into labels under the label mapping rule to obtain the label-converted sound effect data;

[0039] The unlabeled sound effect data is subjected to sound effect recognition. Based on the obtained sound effect recognition results, the data in the unlabeled sound effect data that does not contain sound effects is filtered out to obtain the target data.

[0040] Based on the extracted audio data, the sound effect data converted from the tags, and the target data, the first sound effect dataset is constructed.

[0041] In other embodiments, constructing the first sound effect dataset based on the extracted audio data, the tagged sound effect data, and the target data includes:

[0042] The extracted audio data, the labeled sound effect data, and the target data are preprocessed to obtain a third sound effect dataset;

[0043] The sound effect data included in the third dataset is cleaned to obtain the first sound effect dataset;

[0044] The data preprocessing includes one or more of the following:

[0045] Perform quality filtering on video data, including audio effects;

[0046] Transcode audio data, including sound effects;

[0047] The video data and the audio data are sliced.

[0048] In other embodiments, the data cleaning of the sound effect data included in the third dataset to obtain the first sound effect dataset includes one or more of the following:

[0049] The audio data is subjected to quality filtering; wherein the audio data subjected to quality filtering includes the original audio data in the third sound effect dataset and the audio data extracted from the video data included in the third sound effect dataset;

[0050] Filter out data in the third sound effect dataset that does not include sound effects;

[0051] Filter out the audio effect data with incorrect labels in the third audio effect dataset.

[0052] In other embodiments, generating initial sound effect description text for the sound effect data included in the first sound effect dataset includes:

[0053] The sound effect data included in the first sound effect dataset is classified to obtain the classification label of each sound effect data in the first sound effect dataset;

[0054] For any segment of sound effect data in the first sound effect dataset, a model matching the sound effect data is determined based on at least one of the classification label, label quality, or video correlation of the sound effect data; based on the model matching the sound effect data, the initial sound effect description text is generated for the sound effect data.

[0055] The video correlation is used to indicate whether the audio data is video data that includes audio effects.

[0056] According to a second aspect of the present disclosure, a sound effect generation method is provided, the method comprising:

[0057] During the model training phase, a second sound effect dataset is obtained; wherein, the second sound effect dataset is a dataset obtained based on any of the above data processing methods; the model is trained based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task;

[0058] During the sound effect activation phase, user input data is acquired, and the sound effect generation model is invoked to generate sound effects based on the user input data, thereby obtaining sound effects that match the user input data.

[0059] According to a third aspect of the present disclosure, a data processing apparatus is provided, comprising:

[0060] The first building module is configured to build the first sound effects dataset;

[0061] The text generation module is configured to generate initial sound effect description text for the sound effect data included in the first sound effect dataset;

[0062] The data augmentation module is configured to augment the initial sound effect description text based on the possible input forms of the user during the request generation sound effect stage, so as to obtain the target sound effect description text.

[0063] The second building module is configured to build a second sound effect dataset for training a sound effect generation model based on the sound effect data included in the first sound effect dataset and the target sound effect description text.

[0064] In some embodiments, the data enhancement module is configured to:

[0065] Under given constraints, based on the possible input forms of the user during the request to generate sound effects, the initial sound effect description text is augmented with data to obtain the target sound effect description text.

[0066] The constraints are used to control the elements included in the target sound effect description text.

[0067] In other embodiments, the constraints include one or more of the following:

[0068] The output audio effect description text should include an audio quality description for each audio effect data segment;

[0069] The output sound effect description text should include an audio category description for each sound effect data segment;

[0070] The output sound effect description text should include a detailed description of the sound effect for each sound effect data segment;

[0071] The output sound effect description text should include a detailed description of the background music for each sound effect data segment.

[0072] In other embodiments, the input format includes keyword input;

[0073] The data enhancement module is configured to add keyword descriptions to the initial sound effect description text.

[0074] In other embodiments, the input format further includes timing input; the data enhancement module is configured to:

[0075] In the first sound effect dataset, filter first type of sound effect data and second type of sound effect data; wherein, the first type of sound effect data is used to describe a first type of event with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe a second type of event with a sound effect duration less than a second threshold; the first threshold is greater than the second threshold;

[0076] Each time, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data respectively;

[0077] After combining the two selected sound effect data segments, a sound effect description text including temporal relationships is generated for the combined sound effect data based on the initial sound effect description text of the two sound effect data segments.

[0078] In other embodiments, the two audio effect data segments include first audio effect data and second audio effect data; the first audio effect data comes from the first type of audio effect data, and the second audio effect data comes from the second type of audio effect data;

[0079] The data enhancement module is also configured to:

[0080] The first sound effect data is divided into multiple time periods, and the second sound effect data is inserted into different time periods of the first sound effect data to obtain multiple combined sound effect data.

[0081] Add the combined audio data of the multiple segments to the first audio data set.

[0082] In other embodiments, the input format further includes no text input; the data enhancement module is configured to:

[0083] Under the condition of meeting the preset duration limit, at least two audio effect data segments in the first audio effect dataset are spliced ​​together to obtain spliced ​​audio effect data.

[0084] Based on the initial sound effect description text of the at least two sound effect data segments, a sound effect description text including temporal relationships is generated for the concatenated sound effect data.

[0085] In other embodiments, the first building module is configured as follows:

[0086] The first sound effect dataset is constructed based on open-source shared sound effect data and unlabeled sound effect data;

[0087] The types of sound effect data in the first sound effect dataset include:

[0088] Audio data including sound effects;

[0089] Video data including sound effects.

[0090] In other embodiments, the first building module is configured as follows:

[0091] If the shared sound effect data is website data, candidate videos are filtered from the website data based on preset tag text to obtain a candidate video set;

[0092] If the shared sound effect data comes from multiple open-source datasets, establish a label mapping rule, and convert the original labels of the sound effect data included in the multiple open-source datasets into labels under the label mapping rule to obtain the label-converted sound effect data;

[0093] The unlabeled sound effect data is subjected to sound effect recognition. Based on the obtained sound effect recognition results, the data in the unlabeled sound effect data that does not contain sound effects is filtered out to obtain the target data.

[0094] Based on the extracted audio data, the sound effect data converted from the tags, and the target data, the first sound effect dataset is constructed.

[0095] In other embodiments, the first building module is configured as follows:

[0096] The extracted audio data, the labeled sound effect data, and the target data are preprocessed to obtain a third sound effect dataset;

[0097] The sound effect data included in the third dataset is cleaned to obtain the first sound effect dataset;

[0098] The data preprocessing includes one or more of the following:

[0099] Perform quality filtering on video data, including audio effects;

[0100] Transcode audio data, including sound effects;

[0101] The video data and the audio data are sliced.

[0102] In other embodiments, the first building module is configured to perform one or more of the following:

[0103] The audio data is subjected to quality filtering; wherein the audio data subjected to quality filtering includes the original audio data in the third sound effect dataset and the audio data extracted from the video data included in the third sound effect dataset;

[0104] Filter out data in the third sound effect dataset that does not include sound effects;

[0105] Filter out the audio effect data with incorrect labels in the third audio effect dataset.

[0106] In other embodiments, the generation module is configured to:

[0107] The sound effect data included in the first sound effect dataset is classified to obtain the classification label of each sound effect data in the first sound effect dataset;

[0108] For any segment of sound effect data in the first sound effect dataset, a model matching the sound effect data is determined based on at least one of the classification label, label quality, or video correlation of the sound effect data; based on the model matching the sound effect data, the initial sound effect description text is generated for the sound effect data.

[0109] The video correlation is used to indicate whether the audio data is video data that includes audio effects.

[0110] According to a fourth aspect of the present disclosure, a sound effect generation apparatus is provided, the apparatus comprising:

[0111] The training module is configured to acquire a second sound effect dataset during the model training phase; wherein the second sound effect dataset is a dataset obtained based on any of the above-mentioned data processing methods; and to train the model based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task.

[0112] The sound effect generation module is configured to, during the sound effect activation phase, acquire user input data, call the sound effect generation model to generate sound effects based on the user input data, and obtain sound effects that match the user input data.

[0113] According to a fifth aspect of the present disclosure, an electronic device is provided, the electronic device comprising:

[0114] One or more processors;

[0115] Memory used to store the executable program code of the processor;

[0116] The processor is configured to execute the program code to implement the aforementioned data processing method or the aforementioned sound effect generation method.

[0117] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein program code in the computer-readable storage medium, when executed by a processor of an electronic device, enables the electronic device to perform the data processing method or the sound effect generation method described above.

[0118] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor of an electronic device, implements the data processing method or the sound effect generation method described above.

[0119] For sound effect generation scenarios, the data processing scheme provided in this disclosure can construct high-quality sound effect data for training a sound effect generation model. Specifically, considering the multiple possible input formats from users during the sound effect generation request stage, this scheme introduces a data augmentation step to enable the model to learn the corresponding knowledge during training. In the data augmentation stage, this scheme can augment the sound effect description text previously generated for each sound effect data segment based on the possible input formats of the user during the sound effect generation request stage. After data augmentation, training data for training the sound effect generation model is obtained. Since data augmentation yields more accurate sound effect description text, regardless of the user's input format during the sound effect generation request stage, the sound effect generation model trained based on the above training data can output sound effects adapted to the user's needs, ensuring the generation quality of the sound effect generation model.

[0120] In summary, this solution makes the sound effects generated in the sound effect generation task more accurate and better meet user needs, thereby improving the human-computer interaction effect.

[0121] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0122] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0123] Figure 1 This is a schematic diagram illustrating the implementation environment of a data processing method according to an exemplary embodiment.

[0124] Figure 2 This is a flowchart illustrating a data processing method according to an exemplary embodiment.

[0125] Figure 3 This is a schematic diagram illustrating a data mining, preprocessing, and cleaning process according to an exemplary embodiment.

[0126] Figure 4 This is a schematic diagram illustrating a process for generating captions and data augmentation according to an exemplary embodiment.

[0127] Figure 5 This is a flowchart illustrating another data processing method according to an exemplary embodiment.

[0128] Figure 6 This is a schematic diagram illustrating a structured description according to an exemplary embodiment.

[0129] Figure 7 This is a flowchart illustrating a sound effect generation method according to an exemplary embodiment.

[0130] Figure 8 This is a block diagram illustrating a data processing apparatus according to an exemplary embodiment.

[0131] Figure 9 This is a block diagram illustrating a sound effect generation apparatus according to an exemplary embodiment.

[0132] Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0133] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0134] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0135] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by the parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0136] Figure 1 This is a schematic diagram illustrating the implementation environment of a data processing method according to an exemplary embodiment.

[0137] See Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 is an electronic device used by the user. In this embodiment, a target application with sound effect generation functionality is installed on the terminal 101. Exemplarily, the target application is an instant messaging application, such as a short video application; this disclosure does not limit this application.

[0138] In some embodiments, the terminal 101 is a device such as a smartphone, desktop computer, or laptop. Figure 1 The example of terminal 101 being a smartphone is provided for illustration only. Furthermore, those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only a few terminals, or dozens or hundreds of devices, or even more. This disclosure does not limit the number or type of terminals.

[0139] In other embodiments, server 102 provides background services for the target application. Additionally, server 102 is connected to terminal 101 via a wireless or wired network. Furthermore, server 102 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server; this disclosure does not limit this. Furthermore, the servers involved in the embodiments of this disclosure may also include other functional servers to provide more comprehensive and diversified services.

[0140] Based on the aforementioned implementation environment, this disclosure provides a data processing scheme that not only constructs a database containing large-scale, high-quality audio effect data, but also generates refined and diverse audio effect description text, i.e., captions, for the audio effect data. Specifically, this disclosure, based on a cross-modal data processing system, can generate data combining audio-text, audio-video, and audio-video-text modalities. This data can be used for individual or combined training of TTA (Text to Audio), VTA (Video to Audio), or TVTA (Text and Video to Audio) tasks. In detail, this scheme achieves the following aspects.

[0141] (1) This solution constructs a database of high-quality sound effect data and establishes a pipeline for data mining, preprocessing, and cleaning. It ensures both the quantity and quality of the sound effect data. In other words, this solution can construct a massive and high-quality sound effect dataset from massive amounts of internet data and open-source datasets for training sound effect generation models.

[0142] (2) In order to match the natural language input of users in the application stage and the demand for fine control of sound effects, this solution builds a caption generation and enhancement system that integrates multimodal large models and supports structured detailed caption generation.

[0143] For example, this solution integrates the semantic parsing capabilities of VLM (Visual Language Model) for video scenes, the description and generation capabilities of LLM (Large Language Model) for complex actions, and the understanding capabilities of ALM (Audio Language Model) for acoustic features.

[0144] (3) This solution supports the generation of TTA data (audio-text), VTA data (audio-video) and TVTA data (audio-video-text).

[0145] The data processing scheme provided in this disclosure will be described in detail below through the following embodiments.

[0146] Figure 2 This is a flowchart illustrating a data processing method according to an exemplary embodiment, such as... Figure 2 As shown, this data processing method is applied in electronic devices, such as... Figure 1 In the server 102 shown, the data processing method includes the following steps.

[0147] In 201, the electronic device constructs a first sound effect dataset and generates initial sound effect description text for the sound effect data included in the first sound effect dataset.

[0148] In this embodiment of the disclosure, the first sound effect dataset refers to Figure 3 The clean data obtained after data cleaning.

[0149] It should be noted that the short duration, high background noise, multi-source heterogeneity, and strong scene dependence of sound effect data make data cleaning and annotation difficult. Furthermore, the significant long-tail distribution (rare sound effect samples are scarce), weak spatiotemporal correlation, and the need to balance copyright compliance and privacy protection all exacerbate the difficulty of constructing large-scale, high-quality sound effect datasets.

[0150] To construct a large-scale, high-quality audio effects dataset, this solution establishes a data processing pipeline for data mining, data preprocessing, and data cleaning. Among these, Figure 3 This is a schematic diagram illustrating a data mining, preprocessing, and cleaning process according to an exemplary embodiment. Figure 3 The data mining, data preprocessing, and data cleaning shown can build a database of high-quality sound effect samples. The sound effect samples in this database are also referred to as the first sound effect dataset in the text.

[0151] After constructing the first sound effects dataset, in order to match possible user input formats during the application phase, this solution will understand the semantics of the natural language input and generate sound effects that meet user needs. This solution will also perform processes such as... Figure 4 The processing shown. Figure 4 This is a schematic diagram illustrating a process for generating captions and data augmentation according to an exemplary embodiment. See also Figure 4 During the Caption generation stage, a natural language Caption, i.e., the initial sound effect description text, is generated for each sound effect data.

[0152] It should be noted that this solution can process not only audio data but also video data, achieving multimodal data processing. Therefore, the types of audio effect data in the first audio effect dataset include, but are not limited to, the following two types.

[0153] Audio data including sound effects. For example, audio clips that include sound effects.

[0154] Video data including sound effects. For example, video clips that include sound effects.

[0155] Accordingly, by generating the Caption stage, one can obtain either a sound effect description text for an audio clip that includes sound effects or a sound effect description text for a video clip that contains sound effects.

[0156] In step 202, the electronic device performs data augmentation on the initial sound effect description text based on the possible input forms of the user during the request to generate sound effects, thereby obtaining the target sound effect description text.

[0157] In this embodiment of the disclosure, the request to generate sound effects stage refers to the application stage, that is, the stage of generating sound effects based on the trained sound effect generation model.

[0158] Considering that users may input different types of prompts during the application phase, and in order to achieve fine-grained control over the generated sound effects, it is also necessary to construct corresponding scenarios from data so that the model can learn the corresponding knowledge during training. Therefore, after generating a version of the Caption, this solution also adds a data augmentation step.

[0159] For example, see Figure 4 During the sound effect generation stage, possible user input formats include: no text input, keyword input, and sequential input (i.e., multiple events occurring in chronological order), which this disclosure does not limit. In addition, this solution introduces structured descriptions to control the elements included in the final output caption. The structured description is essentially a set of preset element templates, whose function is to constrain the information that must be included in the final output sound effect description text. For details on the data augmentation process, please refer to the description below.

[0160] In 203, the electronic device constructs a second sound effect dataset for training the sound effect generation model based on the sound effect data included in the first sound effect dataset and the target sound effect description text.

[0161] In this embodiment of the disclosure, the second sound effect dataset refers to Figure 4 The training data obtained after data augmentation. This training data consists of combinations of three modalities: audio-text, audio-video, and audio-video-text, and can be used for training TTA, VTA, and TVTA tasks individually or in combination.

[0162] For sound effect generation scenarios, the data processing scheme provided in this disclosure can construct high-quality sound effect data for training a sound effect generation model. Specifically, considering the multiple possible input formats from users during the sound effect generation request stage, this scheme introduces a data augmentation step to enable the model to learn the corresponding knowledge during training. In the data augmentation stage, this scheme can augment the sound effect description text previously generated for each sound effect data segment based on the possible input formats of the user during the sound effect generation request stage. After data augmentation, training data for training the sound effect generation model is obtained. Since data augmentation yields more accurate sound effect description text, regardless of the user's input format during the sound effect generation request stage, the sound effect generation model trained based on the above training data can output sound effects adapted to the user's needs, ensuring the generation quality of the sound effect generation model.

[0163] In conclusion, this solution makes the sound effects generated in the sound effect generation task more accurate and better meet user needs, thereby improving the human-computer interaction effect.

[0164] The above Figure 2 The diagram shown is merely the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on a specific implementation method. Figure 5 This is a flowchart illustrating another data processing method according to an exemplary embodiment, such as... Figure 3 As shown, this data processing method is applied in electronic devices, such as... Figure 1 In the server 102 shown, the data processing method includes the following steps.

[0165] In 501, the electronic device constructs the first sound effects dataset.

[0166] As an example, see Figure 3 The construction of the first sound effects dataset involves three steps: data mining, data preprocessing, and data cleaning. These three steps will be described below.

[0167] Data mining

[0168] like Figure 3 As shown, this solution constructs the first sound effect dataset based on open-source shared sound effect data and unlabeled sound effect data. The open-source shared sound effect data includes, but is not limited to, data from open-source websites and sound effect data from open-source datasets.

[0169] A. If the shared audio data in an open-source state is website data, this solution will filter candidate videos from these website data based on preset tag text to obtain a candidate video set; and extract audio data from this candidate video set.

[0170] This set of videos is also known as the candidate video set. After obtaining the candidate video set and completing audio extraction, further steps can be taken, such as... Figure 3 The audio separation and signal-to-noise ratio filtering shown are not limited in this disclosure. The following describes... Figure 3 The process of data mining in open-source website data.

[0171] For data from open-source websites, see [link / reference]. Figure 3 First, a tag-based text retrieval mechanism is employed to filter candidate video sets through keyword matching. This tag-based text retrieval mechanism uses pre-labeled tag text as a bridge to find, match, and filter target data (especially non-text data, such as audio or video). The core is to transform the features of non-text data into structured or semi-structured tag text, and then use text matching algorithms to associate user queries with target data. Essentially, it uses text retrieval logic to solve the problem of finding non-text data.

[0172] Next, audio data is extracted from the candidate videos. Then, an audio separation model is used to extract the target audio data from this audio data. The audio separation is used to automatically separate the target sound source (such as a single human voice, piano sound, or car engine sound) from mixed audio (such as audio containing human voices, musical instruments, or background noise) using a deep learning model.

[0173] Next, evaluation metrics such as signal-to-noise ratio are introduced for quality filtering to ensure the acquisition of high-quality audio clips. Because this data mining process automates the extraction of in-site audio data, it significantly improves data processing efficiency.

[0174] It should be noted that, Figure 3 The audio separation step and signal-to-noise ratio filtering step shown may also be omitted, and this disclosure does not limit this.

[0175] B. If the shared sound effect data in the open source state comes from multiple open source datasets, establish a label mapping rule, and convert the original labels of the sound effect data included in the multiple open source datasets into labels under the label mapping rule to obtain the label-converted sound effect data.

[0176] For audio effect data from multiple open-source datasets, the data processing system first integrates these high-quality datasets and then establishes a unified tag mapping system. This enables rapid location of target audio effect data through automated tag matching and filtering mechanisms. Specifically, the tag mapping system transforms scattered and heterogeneous open-source tags into an interoperable and reusable standardized tag system using unified rules and logic, thereby achieving cross-dataset integration, retrieval, and analysis. The overall logic is: integrating scattered open-source data, unifying tags from different sources, and rapidly matching and filtering using automated tools to efficiently acquire target audio effect data.

[0177] C. Perform sound effect recognition on the unlabeled sound effect data. Based on the obtained sound effect recognition results, filter out the data in the unlabeled sound effect data that does not contain sound effects to obtain the target data.

[0178] Unlabeled data refers to audio effect data that has not been manually labeled. This type of data usually comes from complex sources and may contain a large amount of non-target data, which can cause interference if used directly for model training.

[0179] For example, for unlabeled data, this solution uses AED (Audio Event Detection) technology for sound effect recognition, and this disclosure does not limit this.

[0180] For this type of data processing, by constructing a sound effect feature library and establishing a multi-dimensional filtering mechanism, non-target data such as speech or singing in unlabeled sound effect data can be effectively removed, which significantly improves the accuracy and targeting of data mining and can provide high-quality data for subsequent sound effect generation tasks.

[0181] Data preprocessing

[0182] See Figure 3 After the data mining process described above (AC), the data preprocessing step can be initiated. For example, before the mined sound effect data enters the data preprocessing process, it can be determined whether the sound effect data actually contains sound effects. Only data containing sound effects can enter the data preprocessing process; this disclosure does not limit this step.

[0183] In this embodiment of the disclosure, by preprocessing the mined sound effect data—namely, the audio data obtained in step A, the tag-converted sound effect data obtained in step B, and the target data obtained in step C—a preprocessed sound effect dataset can be obtained. This sound effect dataset is also referred to herein as the third sound effect dataset. As an example, such as... Figure 3 As shown, data preprocessing includes one or more of the following.

[0184] a. Perform quality filtering on video data, including audio effects.

[0185] Quality filtering of video data, including audio effects, is performed to ensure that the video encoder correctly understands the video, thus serving the subsequent caption generation and data augmentation processes.

[0186] For example, based on the assumption that low-quality video implies low-quality audio, this solution filters video resolution, such as filtering out video data with a resolution less than 720P. Additionally, this solution performs OCR (Optical Character Recognition) subtitle recognition on the video data and filters out video data with subtitles that are too large in proportion to the video content.

[0187] b. Transcode the audio data, including sound effects.

[0188] For example, this solution will transcode the audio data, including sound effects, into 44k sampling rate, 16-bit, dual-channel WAV format data to facilitate subsequent processing.

[0189] c. Slice the video data including sound effects and the audio data including sound effects.

[0190] This solution will perform unified segmentation processing on video data including sound effects and audio data including sound effects. For example, the segment duration is 10 seconds, but this disclosure does not limit it.

[0191] It should be noted that other types of preprocessing can be performed on the mined sound effect data during the data preprocessing stage, and this disclosure does not limit this.

[0192] Data cleaning

[0193] Because the preprocessed audio data still contains meaningless silent segments, messy and poor audio quality, and inaccurate text labels, data cleaning is still required.

[0194] See Figure 3 After cleaning the sound effect data included in the third dataset, the first sound effect dataset is obtained. As an example, data cleaning includes one or more of the following.

[0195] 1. Perform quality filtering on the audio data. The audio data to be quality filtered includes the original audio data in the third sound effects dataset and the audio data extracted from the video data included in the third sound effects dataset.

[0196] For example, quality filtering includes filtering out audio data with excessively long silent segments. The specific process involves determining the percentage of silence in each audio effect data segment in the third audio effect dataset and deleting audio effect data with a silence percentage greater than a preset threshold. For instance, this solution uses a VAD detection tool to detect silent segments and valid sound segments in each audio effect data segment and deletes audio effect data with excessively long silent segments. For example, audio effect data with a valid sound percentage less than 0.8 is deleted.

[0197] In addition, audio data can be quality filtered based on signal-to-noise ratio, MOS (Mean Opinion Score), clipping ratio, audio bandwidth, etc., to obtain high-quality sound data. This disclosure does not limit this.

[0198] 2. AED model filtering, which filters out data in the third sound effect dataset that does not include sound effects.

[0199] As an example, the AED model is a four-class classification model (speech, music, singing, sound effects), and this disclosure does not limit it. It should be noted that, in addition to training an AED model for filtering, open-source models (multi-class classification) such as BEATs and PANNs can also be used, and this disclosure also does not limit it.

[0200] 3. CLAP (Contrastive Language-Audio Pretraining) model filtering, which filters out audio effect data with incorrect labels in the third audio effect dataset.

[0201] This approach uses a pre-trained CLAP model to filter mislabeled audio effect data. Furthermore, because audio effect data comes from a wide range of sources, many data points are outside the CLAP model's domain; therefore, a relatively low score threshold (score > 0) is typically used for data filtering. Here, "outside the CLAP model's domain" refers to data whose distribution differs from the model's pre-training data.

[0202] See Figure 3 After three steps—data mining, data preprocessing, and data cleaning—a relatively clean audio effects dataset was obtained. This dataset includes both audio and audio-video modalities. Audio refers to audio clips containing sound effects, while audio-video refers to video clips containing sound effects.

[0203] It should be noted that, in Figure 3 In addition to filtering candidate video sets through audio tag text, search engines can also map relevant videos of user accounts in specific vertical categories (such as games) to the corresponding vertical category.

[0204] In addition, the audio quality filtering step in the data cleaning process can also be done without using MOS scoring and VAD detection tools, but by directly filtering out audio data with a high proportion of silence through energy detection. This disclosure does not limit this.

[0205] Additionally, see Figure 4 The data cleaning process may also include audio-video synchronization filtering. This filtering step can be implemented based on a CVAP (Contrastive Audio-Visual Pretraining) model, which is not limited in this disclosure. Of course, this step can also be skipped, retaining some samples with audio-visual asynchrony, thereby preserving data diversity.

[0206] The following is combined Figure 4 We will begin by introducing the process of generating captions and data augmentation.

[0207] In step 502, the electronic device generates initial sound effect description text for the sound effect data included in the first sound effect dataset.

[0208] In some embodiments, initial sound effect description text is generated for the sound effect data included in the first sound effect dataset, including but not limited to the following steps.

[0209] 5021. Classify the sound effect data included in the first sound effect dataset to obtain the classification label for each sound effect data segment in the first sound effect dataset.

[0210] like Figure 4 As shown, this scheme first uses an expert model in the audio domain to classify the audio data in the first sound effect dataset, and then uses an expert model in the visual domain to classify the video data in the first sound effect dataset.

[0211] For example, when classifying audio data, common classification methods such as Beats and CLAP models can be used, or other classification models can be used for more detailed classification; this disclosure does not limit this approach. When classifying video data, the CLIP model can be used to obtain classification labels by extracting frames from the video data.

[0212] 5022. For any segment of sound effect data in the first sound effect dataset, determine a model that matches the segment of sound effect data based on at least one of the classification label, label quality, or video correlation of the segment of sound effect data; generate initial sound effect description text for the segment of sound effect data based on the model that matches the segment of sound effect data; wherein, video correlation is used to indicate whether the segment of sound effect data is video data that includes sound effects.

[0213] In this embodiment of the disclosure, large model technology is used to generate the caption, i.e., the audio description text.

[0214] As an example, for audio data with missing or inaccurate labels, or audio data where label accuracy drops significantly after slicing, this solution uses ALM (such as Qwen-Audio) or a large music model (such as lp-music-caps) to generate captions.

[0215] As another example, for audio data with more accurate labels, this approach calls LLM to generate captions.

[0216] As another example, for audio data in video format, this solution uses the VLM model to extract visual information from the video, and then combines it with the classification labels given by the expert model, and finally generates the caption by LLM.

[0217] In summary, during the Caption generation stage, the audio description text for each audio effect data segment can be obtained.

[0218] It should be noted that other models can be added as expert models during the caption generation stage to enrich the classification information of the audio and video data; this disclosure does not limit this. Furthermore, during the caption generation stage, a sound detection model can be introduced to determine the start and end points of the audio, and a video detection model can be introduced to determine the location, size, and other information of objects in the video; this disclosure does not limit this.

[0219] In 503, under given constraints, the electronic device performs data augmentation on the initial sound effect description text based on the possible input forms of the user during the request to generate sound effects, thereby obtaining the target sound effect description text.

[0220] In this embodiment of the disclosure, the constraints, also referred to as structured control input or structured description, are used to control the elements included in the target sound effect description text. Regardless of the user input format during the application phase, the structured description, when constructing training data for training the sound effect generation model, forces the data processing system to ultimately output a caption containing preset elements.

[0221] As an example, the constraints above may include one or more of the following.

[0222] The output audio effect description text should include an audio quality description for each audio effect data segment;

[0223] The output sound effect description text should include an audio category description for each sound effect data segment;

[0224] The output sound effect description text should include a detailed description of the sound effect for each sound effect data segment;

[0225] The output sound effect description text should include a detailed description of the BGM (Background Music) for each sound effect data segment.

[0226] Figure 6 This is a schematic diagram illustrating a structured description according to an exemplary embodiment. The following is in conjunction with... Figure 6 This section introduces descriptions of audio quality, audio category, sound effects details, and background music (BGM) details.

[0227] The audio quality description is used to describe the quality of this sound effect data, scoring it based on audio quality. This element can improve the audio quality generated during the application phase. The audio category description is used to describe whether this sound effect data belongs to speech, singing, music, or sound effects, distinguishing the sound effect data into broad categories.

[0228] The detailed description of sound effects includes, but is not limited to, the type, scene, distance, material, action, and timing of the sound effects contained in this sound effect data, and this disclosure does not limit this. The detailed description of BGM includes, but is not limited to, the style, genre, mood, instruments, scene, and music theory of the BGM, and this disclosure does not limit this.

[0229] No text input by the user

[0230] During the application phase, users may only upload videos without providing additional text input. Furthermore, due to the concise nature and limited information in short audio or video clips, captions can easily become overly simplistic or inaccurate. To ensure accurate sound effect generation even without user text input during the application phase, and to fully utilize short audio and video clips, this solution uses video stitching to combine two short-duration sound effect data segments into a longer-duration one. This significantly increases the amount of usable data and allows the model to learn multi-camera data generation through data stitching, enabling smooth audio conversion during shot transitions and compensating for the gaps in multi-camera data.

[0231] In summary, the possible input forms for users during the sound effect generation stage include no text input. Accordingly, based on the possible input forms for users during the sound effect generation stage, data augmentation is performed on the initial sound effect description text to obtain the target sound effect description text. This includes: concatenating at least two sound effect data segments from the first sound effect dataset under a preset duration limit to obtain concatenated sound effect data; and generating a sound effect description text including temporal relationships for the concatenated sound effect data based on the initial sound effect description text of the at least two sound effect data segments.

[0232] It should be noted that when splicing audio data, videos are spliced ​​together, and audio files are spliced ​​together. Additionally, the spliced ​​audio data meets the preset duration limit.

[0233] In addition, under given constraints, this solution generates a sound effect description text including temporal relationships for the spliced ​​sound effect data based on the initial sound effect description text of at least two sound effect data segments.

[0234] User input keywords

[0235] During the application phase, user input text may be highly irregular. For example, descriptions may be simplified, using phrases or even single words instead of complete sentences. User input may also be fragmented and grammatically incomplete, such as inputting "rain dripping on the ground." This irregular text input makes it difficult for the model to accurately understand user needs, thus affecting the accuracy of the generated sound effects. To increase the model's robustness to different user inputs, this solution introduces LLM to enhance the initial sound effect description text by adding keyword descriptions to address issues of simplified descriptions and incomplete sentences. This data processing method reduces the user's learning cost during the application phase and also reduces the learning difficulty for the model after mapping user input to the latent space.

[0236] In summary, possible user input formats during the sound effect generation stage also include keyword input. Accordingly, based on these possible user input formats, data augmentation is performed on the initial sound effect description text to obtain the target sound effect description text. This includes optimizing the generated caption using LLM to supplement keyword descriptions. In other words, this solution adds keyword descriptions to the initial sound effect description text using LLM.

[0237] For example, adding keyword descriptions using the text understanding and generation capabilities of LLM can be: expanding the simplified description of user input in detail under given constraints, or integrating fragmented input that does not form sentences into a coherent description while retaining the complete description of the core keywords. This disclosure does not limit this.

[0238] User inputs timing sequence

[0239] To more accurately control the order of events, this solution also proposes temporal data augmentation. By combining at least two audio data segments, mixed audio data under different temporal conditions can be obtained, which greatly increases the diversity of data and improves the model's semantic compliance with multi-event description requests.

[0240] In summary, the possible input forms for users during the sound effect generation stage also include temporal input. Accordingly, based on the possible input forms for users during the sound effect generation stage, data augmentation is performed on the initial sound effect description text to obtain the target sound effect description text, including the following steps.

[0241] First, filter the first type of sound effect data and the second type of sound effect data in the first sound effect dataset; wherein, the first type of sound effect data is used to describe the first type of events (long events) with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe the second type of events (short events) with a sound effect duration less than a second threshold.

[0242] The first threshold is greater than the second threshold. For example, the first type of sound effect data is used to describe events with a sound effect duration of more than 5 seconds; the second type of sound effect data is used to describe events with a sound effect duration of less than 3 seconds, and this disclosure does not impose any restrictions on this.

[0243] Next, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data each time. After combining the two selected segments of sound effect data, under given constraints, based on the initial sound effect description text of the two segments of sound effect data, LLM is used to generate a sound effect description text including temporal relationships for the combined sound effect data.

[0244] Taking the two audio effect data segments, including the first audio effect data and the second audio effect data, where the first audio effect data comes from the first type of audio effect data and the second audio effect data comes from the second type of audio effect data, as an example, the method provided in this embodiment further includes: dividing the first audio effect data into multiple time periods and inserting the second audio effect data into different time periods of the first audio effect data to obtain multi-segment combined audio effect data; adding the multi-segment combined audio effect data to the first audio effect dataset.

[0245] For example, the first sound effect data can be divided into three time periods, that is, the second sound effect data can be inserted at different positions before, during, and after the first sound effect data, respectively. This disclosure does not limit this. That is, the longer sound effect data is used as an event that runs through the entire process, and the shorter sound effect data is used as events that occur at different positions before, during, and after the process. As an example, if the longer sound effect data is used to describe the sound of running water from a faucet for 5 seconds and the shorter sound effect data is used to describe the sound of a plate colliding for 2 seconds, then the combination may include, but is not limited to, the following situations.

[0246] Short events appearing earlier in the sequence: the sound of running water from a faucet (0-5 seconds), followed by the sound of a plate being inserted and colliding (0-2 seconds).

[0247] Short events occur in the middle: the sound of running water from a faucet (0-5 seconds), followed by the sound of a plate being inserted and colliding (2-4 seconds).

[0248] Short events occur in the following positions: the sound of running water from a faucet (0-5 seconds), followed by the sound of a plate being inserted and collapsing (3-5 seconds).

[0249] As another example, the data-augmented caption after using a structured description is as follows: This audio has a quality of 0.9. This audio does not contain speech. This audio does not contain vocals. This audio is described as follows: In a mid-range sound field within a modern home kitchen, clear vocals and continuous metallic running water create a stepped soundscape. Metallic clanging sounds occur when the speaker adjusts a valve. The running water sound persists in the mid-background and alternates with the vocals. This audio has a 90% probability of containing music. Its musical description is: A beautiful, romantic, and melancholic jazz piano solo, suitable for romantic settings such as weddings.

[0250] In 504, the electronic device constructs a second sound effect dataset for training a sound effect generation model based on the sound effect data included in the first sound effect dataset and the target sound effect description text.

[0251] In this embodiment, the second sound effect dataset includes not only the sound effect data from the first sound effect dataset but also additional sound effect data generated during the data augmentation step. Each sound effect data segment in the second sound effect dataset corresponds to a sound effect description text. Based on the second sound effect dataset, a sound effect generation model can be trained to perform sound effect generation tasks (TTA tasks, VTA tasks, and individual or combined TVTA tasks).

[0252] In summary, the data processing solution provided in this disclosure can construct large-scale and high-quality audio effect data by performing data mining, data preprocessing, and data cleaning on open-source shared audio effect data and unlabeled data. In other words, this solution can obtain clean and high-quality audio effect data based on multi-source heterogeneous data, realizing the expansion of the data volume of audio effect data. This provides a strong guarantee for the subsequent training of high-quality audio effect generation models.

[0253] In addition to introducing multimodal large models and expert models to classify clean sound effect data and generate initial captions, this solution also incorporates a data augmentation step. During the data augmentation stage, considering the potential for various user input formats in the application phase, this solution proposes multiple data augmentation methods, which significantly improves the accuracy of the final output caption. Therefore, in the application phase, regardless of the user's input format, the sound effect generation model can output sound effects tailored to the user's needs, ensuring the generation quality of the sound effect generation model.

[0254] Furthermore, during the data augmentation phase, this solution supports fine-grained control over the caption, enabling detailed control over the caption. Ultimately, the data processing system can output structured captions, which provide a comprehensive and multi-faceted detailed description of the sound effect data. After data augmentation, and following model training based on the obtained training data, the resulting sound effect generation model can generate more accurate and tailored sound effects based on user requests during the application phase.

[0255] In conclusion, this solution makes the sound effects generated in the sound effect generation task more accurate and better meet user needs, thereby improving the human-computer interaction effect.

[0256] Figure 7 This is a flowchart illustrating a sound effect generation method according to an exemplary embodiment. For example... Figure 7 As shown, this model training method is applied to electronic devices, such as... Figure 1 The model training method, as shown in server 102, includes the following steps.

[0257] 701. During the model training phase, the electronic device acquires the second sound effect dataset and trains the model based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task.

[0258] The second sound effect dataset is a dataset obtained based on any one of the implementation methods described in the above data processing method embodiments. Furthermore, the electronic device used to construct the second sound effect dataset and the electronic device used to train the sound effect generation model can be the same electronic device or different electronic devices; this disclosure does not limit this.

[0259] In this embodiment of the disclosure, the second sound effect dataset refers to Figure 4 The training data obtained after data augmentation consists of audio-text, audio-video, and audio-video-text modal combinations, which can be used for training TTA, VTA, and TVTA tasks individually or in combination. That is, the aforementioned sound effect generation task can be a TTA task, VTA task, or TVTA task individually or in combination.

[0260] Furthermore, the embodiments disclosed herein do not limit the specific architecture of the sound effect generation model or the model training method. The model can be trained using any model architecture or any model training method based on the second sound effect dataset.

[0261] 702. During the sound effect activation stage, the electronic device acquires user input data and calls the sound effect generation model to generate sound effects based on the user input data, thus obtaining sound effects that match the user input data.

[0262] As described above, the aforementioned sound effect generation task can be a TTA task, a VTA task, or a combination of TVTA tasks. Therefore, the aforementioned user input data can be text, video, or a combination of text and video, and this disclosure does not limit it in this regard.

[0263] In summary, this embodiment of the disclosure trains the model based on high-quality sound effect data constructed during the data processing stage. Specifically, considering the various possible input formats from users during the sound effect generation request stage, this solution introduces a data augmentation step during the data processing stage to enable the model to learn the corresponding knowledge during training. That is, this solution augments the sound effect description text previously generated for each sound effect data segment based on the possible input formats from users during the sound effect generation request stage. After data augmentation, high-quality sound effect data is obtained for training the sound effect generation model. Since data augmentation yields more accurate sound effect description text, regardless of the user's input format during the sound effect generation request stage, the sound effect generation model trained based on the constructed high-quality sound effect data can output sound effects that match the user's needs, ensuring the generation quality of the sound effect generation model. Because this solution makes the generated sound effects in the sound effect generation task more accurate and better meet user needs, it improves the human-computer interaction effect.

[0264] Figure 8 This is a block diagram illustrating a data processing apparatus according to an exemplary embodiment. (Refer to...) Figure 8 The device includes the following modules.

[0265] The first building module 801 is configured to build the first sound effects dataset;

[0266] The text generation module 802 is configured to generate initial sound effect description text for the sound effect data included in the first sound effect dataset;

[0267] Data augmentation module 803 is configured to perform data augmentation on the initial sound effect description text based on the possible input forms of the user during the request generation sound effect stage, so as to obtain the target sound effect description text;

[0268] The second construction module 804 is configured to construct a second sound effect dataset for training a sound effect generation model based on the sound effect data included in the first sound effect dataset and the target sound effect description text.

[0269] For sound effect generation scenarios, the data processing scheme provided in this disclosure can construct high-quality sound effect data for training a sound effect generation model. Specifically, considering the multiple possible input formats from users during the sound effect generation request stage, this scheme introduces a data augmentation step to enable the model to learn the corresponding knowledge during training. In the data augmentation stage, this scheme can augment the sound effect description text previously generated for each sound effect data segment based on the possible input formats of the user during the sound effect generation request stage. After data augmentation, training data for training the sound effect generation model is obtained. Since data augmentation yields more accurate sound effect description text, regardless of the user's input format during the sound effect generation request stage, the sound effect generation model trained based on the above training data can output sound effects adapted to the user's needs, ensuring the generation quality of the sound effect generation model.

[0270] In conclusion, this solution makes the sound effects generated in the sound effect generation task more accurate and better meet user needs, thereby improving the human-computer interaction effect.

[0271] In some embodiments, the data enhancement module is configured to:

[0272] Under given constraints, based on the possible input forms of the user during the request to generate sound effects, the initial sound effect description text is augmented with data to obtain the target sound effect description text.

[0273] The constraints are used to control the elements included in the target sound effect description text.

[0274] In other embodiments, the constraints include one or more of the following:

[0275] The output audio effect description text should include an audio quality description for each audio effect data segment;

[0276] The output sound effect description text should include an audio category description for each sound effect data segment;

[0277] The output sound effect description text should include a detailed description of the sound effect for each sound effect data segment;

[0278] The output sound effect description text should include a detailed description of the background music for each sound effect data segment.

[0279] In other embodiments, the input format includes keyword input;

[0280] The data enhancement module is configured to add keyword descriptions to the initial sound effect description text.

[0281] In other embodiments, the input format further includes timing input; the data enhancement module is configured to:

[0282] In the first sound effect dataset, filter first type of sound effect data and second type of sound effect data; wherein, the first type of sound effect data is used to describe a first type of event with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe a second type of event with a sound effect duration less than a second threshold; the first threshold is greater than the second threshold;

[0283] Each time, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data respectively;

[0284] After combining the two selected sound effect data segments, a sound effect description text including temporal relationships is generated for the combined sound effect data based on the initial sound effect description text of the two sound effect data segments.

[0285] In other embodiments, the two audio effect data segments include first audio effect data and second audio effect data; the first audio effect data comes from the first type of audio effect data, and the second audio effect data comes from the second type of audio effect data;

[0286] The data enhancement module is also configured to:

[0287] The first sound effect data is divided into multiple time periods, and the second sound effect data is inserted into different time periods of the first sound effect data to obtain multiple combined sound effect data.

[0288] Add the combined audio data of the multiple segments to the first audio data set.

[0289] In other embodiments, the input format further includes no text input; the data enhancement module is configured to:

[0290] Under the condition of meeting the preset duration limit, at least two audio effect data segments in the first audio effect dataset are spliced ​​together to obtain spliced ​​audio effect data.

[0291] Based on the initial sound effect description text of the at least two sound effect data segments, a sound effect description text including temporal relationships is generated for the concatenated sound effect data.

[0292] In other embodiments, the first building module is configured as follows:

[0293] The first sound effect dataset is constructed based on open-source shared sound effect data and unlabeled sound effect data;

[0294] The types of sound effect data in the first sound effect dataset include:

[0295] Audio data including sound effects;

[0296] Video data including sound effects.

[0297] In other embodiments, the first building module is configured as follows:

[0298] If the shared sound effect data is website data, candidate videos are filtered from the website data based on preset tag text to obtain a candidate video set;

[0299] If the shared sound effect data comes from multiple open-source datasets, establish a label mapping rule, and convert the original labels of the sound effect data included in the multiple open-source datasets into labels under the label mapping rule to obtain the label-converted sound effect data;

[0300] The unlabeled sound effect data is subjected to sound effect recognition. Based on the obtained sound effect recognition results, the data in the unlabeled sound effect data that does not contain sound effects is filtered out to obtain the target data.

[0301] Based on the extracted audio data, the sound effect data converted from the tags, and the target data, the first sound effect dataset is constructed.

[0302] In other embodiments, the first building module is configured as follows:

[0303] The extracted audio data, the labeled sound effect data, and the target data are preprocessed to obtain a third sound effect dataset;

[0304] The sound effect data included in the third dataset is cleaned to obtain the first sound effect dataset;

[0305] The data preprocessing includes one or more of the following:

[0306] Perform quality filtering on video data, including audio effects;

[0307] Transcode audio data, including sound effects;

[0308] The video data and the audio data are sliced.

[0309] In other embodiments, the first building module is configured to perform one or more of the following:

[0310] The audio data is subjected to quality filtering; wherein the audio data subjected to quality filtering includes the original audio data in the third sound effect dataset and the audio data extracted from the video data included in the third sound effect dataset;

[0311] Filter out data in the third sound effect dataset that does not include sound effects;

[0312] Filter out the audio effect data with incorrect labels in the third audio effect dataset.

[0313] In other embodiments, the generation module is configured to:

[0314] The sound effect data included in the first sound effect dataset is classified to obtain the classification label of each sound effect data in the first sound effect dataset;

[0315] For any segment of sound effect data in the first sound effect dataset, a model matching the sound effect data is determined based on at least one of the classification label, label quality, or video correlation of the sound effect data; based on the model matching the sound effect data, the initial sound effect description text is generated for the sound effect data.

[0316] The video correlation is used to indicate whether the audio data is video data that includes audio effects.

[0317] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0318] Figure 9 This is a block diagram illustrating a sound effect generation apparatus according to an exemplary embodiment. See also... Figure 9 The device includes the following modules.

[0319] Training module 901 is configured to acquire a second sound effect dataset during the model training phase; wherein the second sound effect dataset is a dataset obtained based on any of the above data processing methods; and to train the model based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task.

[0320] The sound effect generation module 902 is configured to, during the sound effect activation stage, acquire user input data, call the sound effect generation model to generate sound effects based on the user input data, and obtain sound effects that match the user input data.

[0321] In summary, this embodiment of the disclosure trains the model based on high-quality sound effect data constructed during the data processing stage. Specifically, considering the various possible input formats from users during the sound effect generation request stage, this solution introduces a data augmentation step during the data processing stage to enable the model to learn the corresponding knowledge during training. That is, this solution augments the sound effect description text previously generated for each sound effect data segment based on the possible input formats from users during the sound effect generation request stage. After data augmentation, high-quality sound effect data is obtained for training the sound effect generation model. Since data augmentation yields more accurate sound effect description text, regardless of the user's input format during the sound effect generation request stage, the sound effect generation model trained based on the constructed high-quality sound effect data can output sound effects that match the user's needs, ensuring the generation quality of the sound effect generation model. Because this solution makes the generated sound effects in the sound effect generation task more accurate and better meet user needs, it improves the human-computer interaction effect.

[0322] It should be noted that the data processing apparatus provided in the above embodiments is only illustrated by the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the data processing apparatus and data processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0323] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0324] In some embodiments, when the electronic device is provided as a server Figure 10 This is a block diagram illustrating a server 1000 according to an exemplary embodiment. The server 1000 can vary considerably depending on its configuration or performance, and includes one or more Central Processing Units (CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one line of program code, which is loaded and executed by the processor 1001 to implement the image processing methods provided in the various method embodiments described above. Of course, the server also has wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 also includes other components for implementing device functions, which will not be elaborated upon here.

[0325] In some embodiments, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, which are executed by a processor of an electronic device to implement the data processing method described above. In other embodiments, the computer-readable storage medium is a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0326] In some embodiments, a computer program product is also provided, including a computer program that, when executed by a processor of an electronic device, implements the data processing method described above.

[0327] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0328] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: Construct a first sound effect dataset and generate initial sound effect description text for the sound effect data included in the first sound effect dataset; Based on the possible input formats of the user during the request to generate sound effects, the initial sound effect description text is augmented to obtain the target sound effect description text. Based on the sound effect data included in the first sound effect dataset and the target sound effect description text, a second sound effect dataset is constructed for training the sound effect generation model; The step of performing data augmentation on the initial sound effect description text based on the possible input forms of the user during the request sound effect generation stage to obtain the target sound effect description text includes: If the input format is keyword input, add keyword description to the initial sound effect description text; If the input is a temporal input, a first type of sound effect data and a second type of sound effect data are filtered from the first sound effect dataset; wherein, the first type of sound effect data is used to describe a first type of event with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe a second type of event with a sound effect duration less than a second threshold; the first threshold is greater than the second threshold; each time, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data respectively; after combining the two selected segments of sound effect data, a sound effect description text including temporal relationship is generated for the combined sound effect data based on the initial sound effect description text of the two segments of sound effect data; If the input is a textless input, under the condition of satisfying the preset duration limit, at least two audio effect data segments in the first audio effect dataset are concatenated to obtain concatenated audio effect data; based on the initial audio effect description text of the at least two audio effect data segments, an audio effect description text including temporal relationship is generated for the concatenated audio effect data. The two audio effect data segments include first audio effect data and second audio effect data; the first audio effect data comes from the first type of audio effect data, and the second audio effect data comes from the second type of audio effect data; the method further includes: The first sound effect data is divided into multiple time periods, and the second sound effect data is inserted into different time periods of the first sound effect data to obtain multi-segment combined sound effect data; the multi-segment combined sound effect data is added to the first sound effect dataset.

2. The data processing method according to claim 1, characterized in that, The process of data augmenting the initial sound effect description text based on possible user input formats during the request-to-generate sound effect phase to obtain the target sound effect description text includes: Under given constraints, based on the possible input forms of the user during the request to generate sound effects, the initial sound effect description text is augmented with data to obtain the target sound effect description text. The constraints are used to control the elements included in the target sound effect description text.

3. The data processing method according to claim 2, characterized in that, The constraints include one or more of the following: The output audio effect description text should include an audio quality description for each audio effect data segment; The output sound effect description text should include an audio category description for each sound effect data segment; The output sound effect description text should include a detailed description of the sound effect for each sound effect data segment; The output sound effect description text should include a detailed description of the background music for each sound effect data segment.

4. The data processing method according to any one of claims 1 to 3, characterized in that, The construction of the first sound effect dataset includes: The first sound effect dataset is constructed based on open-source shared sound effect data and unlabeled sound effect data; The types of sound effect data in the first sound effect dataset include: Audio data including sound effects; Video data including sound effects.

5. The data processing method according to claim 4, characterized in that, The first sound effect dataset is constructed based on open-source shared sound effect data and unlabeled sound effect data, including: If the shared audio data is website data, candidate videos are filtered from the website data based on preset tag text to obtain a candidate video set; audio data is extracted from the candidate video set. If the shared sound effect data comes from multiple open-source datasets, establish a label mapping rule, and convert the original labels of the sound effect data included in the multiple open-source datasets into labels under the label mapping rule to obtain the label-converted sound effect data; The unlabeled sound effect data is subjected to sound effect recognition. Based on the obtained sound effect recognition results, the data in the unlabeled sound effect data that does not contain sound effects is filtered out to obtain the target data. Based on the extracted audio data, the sound effect data converted from the tags, and the target data, the first sound effect dataset is constructed.

6. The data processing method according to claim 5, characterized in that, The first sound effect dataset is constructed based on the extracted audio data, the labeled sound effect data, and the target data, including: The extracted audio data, the labeled sound effect data, and the target data are preprocessed to obtain a third sound effect dataset; The sound effect data included in the third sound effect dataset is cleaned to obtain the first sound effect dataset; The data preprocessing includes one or more of the following: Perform quality filtering on video data, including audio effects; Transcode audio data, including sound effects; The video data and the audio data are sliced.

7. The data processing method according to claim 6, characterized in that, The first sound effect dataset is obtained by cleaning the sound effect data included in the third sound effect dataset, and includes one or more of the following: The audio data is subjected to quality filtering; wherein the audio data subjected to quality filtering includes the original audio data in the third sound effect dataset and the audio data extracted from the video data included in the third sound effect dataset; Filter out data in the third sound effect dataset that does not include sound effects; Filter out the audio effect data with incorrect labels in the third audio effect dataset.

8. The data processing method according to any one of claims 1 to 3, characterized in that, The step of generating initial sound effect description text for the sound effect data included in the first sound effect dataset includes: The sound effect data included in the first sound effect dataset is classified to obtain the classification label of each sound effect data in the first sound effect dataset; For any segment of sound effect data in the first sound effect dataset, a model matching the sound effect data is determined based on at least one of the classification label, label quality, or video correlation of the sound effect data; based on the model matching the sound effect data, the initial sound effect description text is generated for the sound effect data. The video correlation is used to indicate whether the audio data is video data that includes audio effects.

9. A method for generating sound effects, characterized in that, The method includes: During the model training phase, a second sound effect dataset is obtained; wherein the second sound effect dataset is a dataset obtained based on any one of the methods of claims 1 to 8; the model is trained based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task; During the sound effect activation phase, user input data is acquired, and the sound effect generation model is invoked to generate sound effects based on the user input data, thereby obtaining sound effects that match the user input data.

10. A data processing apparatus, characterized in that, The device includes: The first building module is configured to build the first sound effects dataset; The text generation module is configured to generate initial sound effect description text for the sound effect data included in the first sound effect dataset; The data augmentation module is configured to augment the initial sound effect description text based on the possible input forms of the user during the request generation sound effect stage, so as to obtain the target sound effect description text. The second construction module is configured to construct a second sound effect dataset for training a sound effect generation model based on the sound effect data included in the first sound effect dataset and the target sound effect description text. The data enhancement module is configured to add keyword descriptions to the initial sound effect description text if the input is in the form of keyword input; The data enhancement module is configured to, if the input is a temporal input, filter a first type of sound effect data and a second type of sound effect data from the first sound effect dataset; wherein, the first type of sound effect data is used to describe a first type of event with a sound effect duration greater than a first threshold; the second type of sound effect data is used to describe a second type of event with a sound effect duration less than a second threshold; the first threshold is greater than the second threshold; each time, a segment of sound effect data is selected from the first type of sound effect data and the second type of sound effect data respectively; after combining the two selected segments of sound effect data, a sound effect description text including temporal relationship is generated for the combined sound effect data based on the initial sound effect description text of the two segments of sound effect data; The data augmentation module is configured to, if the input is no text input, concatenate at least two audio effect data segments from the first audio effect dataset to obtain concatenated audio effect data, provided that a preset duration limit is met; and generate audio effect description text including temporal relationships for the concatenated audio effect data based on the initial audio effect description text of the at least two audio effect data segments. The two audio effect data segments include first audio effect data and second audio effect data; the first audio effect data comes from the first type of audio effect data, and the second audio effect data comes from the second type of audio effect data. The data enhancement module is further configured to divide the first sound effect data into multiple time periods, insert the second sound effect data into different time periods of the first sound effect data, and obtain multi-segment combined sound effect data; and add the multi-segment combined sound effect data to the first sound effect dataset.

11. A sound effect generation device, characterized in that, The device includes: The training module is configured to acquire a second sound effect dataset during the model training phase; wherein the second sound effect dataset is a dataset obtained based on any one of the methods of claims 1 to 8; and to train the model based on the second sound effect dataset to obtain a sound effect generation model for performing the sound effect generation task. The sound effect generation module is configured to, during the sound effect activation phase, acquire user input data, call the sound effect generation model to generate sound effects based on the user input data, and obtain sound effects that match the user input data.

12. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the data processing method as described in any one of claims 1 to 8; or the sound effect generation method as described in claim 9.

13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the data processing method as described in any one of claims 1 to 8; Alternatively, the sound effect generation method as described in claim 9.

14. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor of an electronic device, implements the data processing method as described in any one of claims 1 to 8; or, the sound effect generation method as described in claim 9.

Citation Information

Patent Citations

  • Sound effect adjustment method and device of sound box, equipment and storage medium

    CN117130576A

  • Model generation method, sound effect description generation method, equipment, medium and product

    CN119274588A