Method and device for adding sound effect in audio, electronic equipment and storage medium

By automatically embedding and filtering sound effect identifiers through a pre-trained language model, the problem of low efficiency and poor accuracy in adding audio effects in existing technologies is solved, achieving efficient and accurate sound effect matching. The generated audio is more consistent with the text content, thus improving the user experience.

CN121545489APending Publication Date: 2026-02-17SHANGHAI ZHONG YUAN NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511655203.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing technologies rely on manual operation during the audio effect addition process, resulting in low efficiency and poor accuracy. In particular, the label matching deviation is serious in complex emotions or mixed scenes, making it difficult to meet the needs of large-scale, high-efficiency audio processing.

Method used

An automated method based on pre-trained language models is adopted. The first pre-trained language model embeds sound effect identifiers, and combines semantic similarity and time axis information to accurately match candidate sound effect information in the sound effect library. The second pre-trained language model further filters the target sound effect information and finally generates the target audio.

Benefits of technology

It achieves automated and precise sound effect addition, reduces reliance on professional skills, improves processing efficiency, and generates audio that better matches the text content, providing a superior listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545489A_ABST
    Figure CN121545489A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for adding a sound effect in audio, electronic equipment and a storage medium, and relates to the technical field of audio production and multimedia, and the method comprises the steps: converting an original audio to obtain an original text with a time axis, and automatically embedding a predicted sound effect identifier through a first pre-training language model, subsequent candidate sound effect matching, target sound effect screening and audio synthesis are all automatically completed, dependence on professional ability of audio engineers is avoided, and the processing efficiency is improved; the first pre-training language model is formed by training based on a sample text and a corresponding sound effect identifier, semantic, emotion and scene features of the text can be deeply understood, a sound effect insertion position and sound effect features are accurately predicted, and deviation caused by a fixed tag and a rule is avoided; the candidate sound effect information with the semantic similarity reaching the standard is matched from the sound effect library, and the target sound effect information with the semantic similarity reaching the standard with the intermediate text is screened by the second pre-training language model, so that double guarantees of coarse screening and fine selection are formed, and the manual correction requirement is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of audio production and multimedia, and in particular to a method and device for adding sound effects in audio, electronic equipment and storage medium. BACKGROUND

[0002] In the field of digital reading and audio content production, with the rapid development of mobile Internet technology and the continuous improvement of user demand for immersive content experience, as a content form that combines text information and auditory experience, the market size of audiobooks continues to expand, and the industry's requirements for content production efficiency and quality are increasingly stringent.

[0003] Currently, the mainstream technical path for adding audio sound effects in the industry mainly includes the following two kinds: the first is a manual adding method, specifically, an audio engineer manually selects and adds appropriate sound effects in a digital audio workstation according to the content of the manuscript to be processed, and adjusts the position and duration of the sound effects on the time axis to realize the matching of the sound effects and the manuscript content. This method completely relies on manual operation, which not only requires high professional ability of the audio engineer, but also takes a long time, resulting in extremely low processing efficiency, which is difficult to meet the demand of large-scale and high-efficiency audio processing. The second is a semi-automatic method assisted by labels, which includes the following steps: pre-labeling the manuscript content with emotion labels and scene type labels, and setting corresponding trigger rules based on the labels; in the process of adding sound effects, the system calls the matching sound effects from the sound effect library according to the label information of the manuscript through the preset rules. However, this method relies too much on the accuracy of the preset rules and labels, and when the manuscript content involves complex emotions or mixed scenes, the label matching deviation is easy to occur, resulting in inaccurate sound effect calling, which still needs a lot of manual intervention for correction. SUMMARY

[0004] In view of the above problems, the present application provides a method, device, electronic equipment and storage medium for adding sound effects in audio, which overcomes the above problems or at least partially solves the above problems, and the technical solution is as follows: This application provides a method for adding sound effects to audio. The method includes: converting the original audio into original text, the original text containing timeline information of the text sequence; inputting the original text into a first pre-trained language model, and processing it by the first pre-trained language model to obtain intermediate text that embeds at least one predicted sound effect identifier into the original text; wherein, the first pre-trained language model is trained based on sample text and corresponding labeled sound effect identifiers; the sound effect identifier is used to identify the sound effect insertion position in the text and its corresponding sound effect features; matching candidate sound effect information corresponding to the predicted sound effect identifier from a sound effect library; and determining the semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier. The candidate sound effect information includes multiple sound effect identifiers and sound effect addresses corresponding to the multiple sound effect identifiers; the candidate sound effect information corresponding to the predicted sound effect identifiers and the intermediate text are input into the second pre-trained language model, and the second pre-trained language model selects the target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifiers, whose semantic feature similarity to the intermediate text is greater than or equal to the second preset similarity, and modifies the predicted sound effect identifiers in the intermediate text to the target sound effect information to obtain the target text; wherein, the target text contains the time axis information of the text sequence and the target sound effect information; the target audio is generated according to the target text and the sound effect file corresponding to the target sound effect information.

[0005] This application provides an optional implementation method, in which the original text is input into a first pre-trained language model, and the original text is processed by the first pre-trained language model to obtain intermediate text after embedding at least one predicted sound effect tag into the original text. The method includes: inputting the original text into the first pre-trained language model, and the original text is processed by the first pre-trained language model to obtain at least one predicted sound effect tag and at least one time point of the predicted sound effect tag; the time point is the time point when the predicted sound effect tag appears in the original text; and embedding at least one predicted sound effect tag into the original text according to the timeline information and the time point to obtain intermediate text.

[0006] This application provides an optional implementation method, in which the original text is input into a first pre-trained language model, and the intermediate text after being processed by the first pre-trained language model is obtained by embedding at least one predicted sound effect identifier into the original text. The method includes: generating a first sound effect cue word related to the scene information or emotional information of the original text, wherein the first sound effect cue word is used to describe the auditory attributes of the sound effect and includes adjectives defining sound attributes and verbs defining sound dynamics; inputting the original text and the first sound effect cue word into the first pre-trained language model, and the intermediate text after being processed by the first pre-trained language model is obtained by embedding at least one predicted sound effect identifier into the original text.

[0007] This application provides an optional implementation method for matching candidate sound effect information corresponding to predicted sound effect identifiers from a sound effect library, including: converting the association between the predicted sound effect identifier and any sound effect identifier in the sound effect library into a quantitative index; filtering out sound effect identifiers whose quantitative index meets a preset threshold, and determining their corresponding candidate sound effect information.

[0008] This application provides an optional implementation method for matching candidate sound effect information corresponding to a predicted sound effect identifier from a sound effect library, including: calculating the similarity between the predicted sound effect identifier and any sound effect identifier in the sound effect library; summing the first text length of the predicted sound effect identifier and the second text length of any sound effect identifier to obtain the total text length; performing a ratio calculation on the similarity and the total text length to obtain the matching degree; and obtaining the candidate sound effect information corresponding to any sound effect identifier when the matching degree is greater than or equal to a preset matching degree threshold.

[0009] This application provides an optional implementation method, which involves inputting candidate sound effect information corresponding to the predicted sound effect identifier and intermediate text into a second pre-trained language model. After processing by the second pre-trained language model, target sound effect information with a semantic feature similarity greater than or equal to a second preset similarity is selected from the candidate sound effect information corresponding to the predicted sound effect identifier. The predicted sound effect identifier in the intermediate text is then modified to the target sound effect information to obtain the target text. This includes: generating a second sound effect prompt word that semantically matches the intermediate text based on the candidate sound effect information corresponding to the predicted sound effect identifier. The second sound effect prompt word contains constraints for selecting the target sound effect information from the candidate sound effect information. The intermediate text and the second sound effect prompt word are then input into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information according to the constraints and modifies the predicted sound effect identifier in the intermediate text to the target sound effect information to obtain the target text.

[0010] This application provides an optional implementation method, in which candidate sound effect information corresponding to the predicted sound effect identifier and intermediate text are input into a second pre-trained language model. The second pre-trained language model selects target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier, whose semantic feature similarity to the intermediate text is greater than or equal to a second preset similarity. After modifying the predicted sound effect identifier in the intermediate text to obtain the target sound effect information, and before generating the target audio based on the target text, the sound effect file corresponding to the target sound effect information, and the timeline information, the method further includes: searching for the corresponding sound effect address in the sound effect library based on the target sound effect information; and obtaining the sound effect file corresponding to the target sound effect information based on the sound effect address.

[0011] This application provides an apparatus for adding sound effects to audio, the apparatus comprising: The speech-to-text module is used to convert the original audio into the original text, which contains the timeline information of the text sequence. The intermediate text generation module is used to input the original text into the first pre-trained language model, and the first pre-trained language model processes the original text to obtain intermediate text after embedding at least one predicted sound effect label into the original text; wherein, the first pre-trained language model is trained based on sample text and corresponding labeled sound effect labels; the sound effect labels are used to identify the sound effect insertion position in the text and its corresponding sound effect features. The matching module is used to match candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library; the semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier is greater than or equal to the first preset similarity; the candidate sound effect information includes multiple sound effect identifiers and the sound effect addresses corresponding to the multiple sound effect identifiers. The target text generation module is used to input the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier, and the semantic feature similarity with the intermediate text is greater than or equal to the second preset similarity. The predicted sound effect identifier in the intermediate text is then modified to the target sound effect information to obtain the target text. The target text includes the time axis information of the text sequence and the target sound effect information. The audio generation module is used to generate target audio based on the target text and the corresponding sound effect file.

[0012] This application provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements any of the above-described methods for adding sound effects to audio.

[0013] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for adding sound effects to audio.

[0014] This application first obtains the original text with a timeline through raw audio conversion, and then automatically embeds predicted sound effect tags by a first pre-trained language model. Subsequent candidate sound effect matching, target sound effect selection, and audio synthesis are all completed automatically by the model and system, eliminating the reliance on the professional skills of audio engineers, improving processing efficiency, and adapting to large-scale audio processing needs. The semantic understanding capability of the pre-trained language model replaces fixed labels and rules. The first pre-trained language model is trained based on sample text and corresponding sound effect tags, and can deeply understand the semantic, emotional, and scene features of the text, accurately predicting the sound effect insertion position and sound effect features, avoiding problems caused by labeling bias. At the same time, by matching candidate sound effect information that meets the semantic similarity standard from the sound effect library, and further filtering target sound effect information that meets the semantic similarity standard with the intermediate text by a second pre-trained language model, a double guarantee of coarse screening and fine selection is formed. Even in the face of complex emotions or mixed scenes, accurate matching can be achieved through the model's deep analysis of text details, reducing the need for manual correction. The generated target audio is more in line with the text content, bringing users a better listening experience.

[0015] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating a method for adding sound effects to audio according to an embodiment of this application; Figure 2 This is a schematic diagram of a device for adding sound effects to audio according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] To more clearly illustrate the embodiments of this application, the technical terms used in the embodiments will be briefly introduced below: TTS (Text-to-Speech) is a technology that converts written text into audible speech. It falls under the core category of speech synthesis, and its working principle typically includes the following key steps: 1. Text Analysis: Processing the input text, including word segmentation, grammatical analysis, and punctuation recognition, to ensure accurate understanding of the text structure. 2. Prosodic Modeling: Determining the prosodic features of the speech, such as intonation, rate of speech, and pauses, to make the generated speech more natural and conform to language habits. 3. Speech Generation: Converting the processed text information into a speech signal using an acoustic model, ultimately outputting playable audio.

[0018] Speech-to-Text (STT) is a technology that uses algorithms to convert human speech signals (audio) into corresponding written text, and is also often referred to as speech recognition.

[0019] Subtitle Timing / Subtitle Syncing is a post-production process that precisely matches the timeline information of subtitle text, that is, determines the start and end times of each subtitle's display, so that the subtitles are synchronized with the audio, actions, or scene changes in the picture.

[0020] The standard definition of a Large Language Model (LLM) is a deep learning model trained on large-scale text data. It learns the syntax, semantics, logic, and knowledge of human language through a complex neural network structure (usually with Transformer as the core architecture), and is able to understand and generate human text to complete a variety of natural language processing tasks.

[0021] MoviePy is a Python library for video editing that allows users to perform various processing operations on video and audio files programmatically. The library enables video trimming, splicing, adding subtitles, audio processing (such as merging, splitting, and volume adjustment), and video effects (such as transitions, scaling, and rotation). Its design aims to provide developers with a simple yet powerful interface for quickly completing video editing tasks.

[0022] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0023] To address the inefficiency and inaccuracy issues caused by the over-reliance on manual intervention in mainstream audio sound effect addition techniques, this application provides a method for adding sound effects to audio, such as... Figure 1 As shown, Figure 1 This is a schematic flowchart illustrating a method for adding sound effects to audio according to an embodiment of this application. The method includes the following steps S100~S104: S100: Obtain the original text based on the original audio conversion.

[0024] The original text contains timeline information about the text sequence, indicating the playback time of each character in the original text sequence, including the start and end times. The original text can be obtained through speech-to-text technology.

[0025] In some embodiments, the original audio is first acquired and then converted into original text, which contains timeline information of the text sequence. The original audio can be the audio narration of a video or the original audio of an audiobook.

[0026] Specifically, during the process of converting the original audio into original text, the playback time of each character in the original text is marked.

[0027] For example, the original text containing the timeline information of the text sequence is shown below:

[0028] "text" is the complete text sequence, and "timestamp" is a list of word-by-word timestamps.

[0029] S101. Input the original text into the first pre-trained language model, and after processing by the first pre-trained language model, obtain the intermediate text after embedding at least one predicted sound effect identifier into the original text.

[0030] The first pre-trained language model can be a pre-trained large language model used to infer suitable sound effects for the text and embed sound effect identifiers into the text, outputting text containing sound effect identifiers. The predicted sound effect identifiers are used to identify the sound effect prediction insertion positions in the original text and their corresponding sound effect features.

[0031] The first pre-trained language model is used to understand the semantics of the original text and initially match sound effect identifiers that match the scene / emotion. Based on its understanding of the text's semantics, emotional tone, and scene atmosphere, the first pre-trained language model inserts predicted sound effect identifiers into the original text. The predicted sound effect identifiers are inserted into the original text in the form of "[sound effect name]", such as "[script sound]" or "[thunder]".

[0032] In some embodiments, when performing step S101, the original text is directly input into the first pre-trained language model, and the first pre-trained language model generates at least one predicted sound effect identifier corresponding to the original text. The generated sound effect identifier is more in line with the context and details of the original text. Then, at least one predicted sound effect identifier is embedded into the original text to obtain intermediate text.

[0033] For example, suppose the original text is "Chen rushed to her apartment, only to find it in a mess, with drawers ransacked." The intermediate text output by the first pre-trained language model is "Chen rushed to her apartment [script sound], only to find it in a mess, with drawers ransacked [thunder]." Here, "[script sound]" and "[thunder]" are predicted sound effect identifiers inferred by the first pre-trained language model.

[0034] The above embodiments embed predicted sound effect identifiers into the original text, tightly binding the sound effect identifiers to the text description. This ensures a more accurate correlation between sound effects and the text scene, avoiding a disconnect between sound effects and the scene. This embedding method lays a more accurate foundation for subsequent sound effect matching and audio generation, improving the synergy between sound effects and text content throughout the entire audio production process, and enabling the final generated audio to better echo the scene and atmosphere described in the text.

[0035] It should be noted that the predicted sound effect identifier is a suggestion type and does not need to be unique in the sound effect library. It only provides a directional anchor for subsequent accurate matching, so as to avoid the inefficient dilemma of aimless traversal in subsequent sound effect selection.

[0036] In some embodiments, when performing step S101, a first sound effect cue word related to the scene information or emotional information of the original text is first generated; then the original text and the first sound effect cue word are input into a first pre-trained language model, so that the first pre-trained language model generates at least one predicted sound effect identifier based on the original text and the first sound effect cue word, and embeds the at least one predicted sound effect identifier into the original text to obtain intermediate text.

[0037] The sound effects library is pre-built and includes multiple sound effect information entries. Each entry contains multiple sound effect identifiers and their corresponding sound effect addresses. The sound effect address represents the storage location of the sound effect file. The sound effects library can be categorized, including but not limited to animal sound effects, environmental sound effects, vehicle sound effects, and science fiction sound effects. Each category includes multiple sound effect identifiers; for example, animal sound effects might include "dog barking" or "bird chirping," while environmental sound effects might include "wind sounds," "rain sounds," or "forest sounds."

[0038] The first sound effect cue word is related to the scene or emotional information of the original text and is used to describe the auditory attributes of the sound effect. It includes adjectives that define the sound attributes and verbs that define the sound dynamics. The first sound effect cue word initially anchors the sound effect type based on the semantics of the original text and carries a vague sound effect identifier, without needing to consider the matching restrictions of the sound effect library. The first sound effect cue word can include the original text (including the timeline information of the text sequence), but does not need to be associated with sound effect library data. The constraints of the first sound effect cue word only focus on semantic adaptation and do not restrict whether the sound effect name exists in the sound effect library.

[0039] Following the previous example, the first sound effect prompt can be: "Based on the following original text, please insert a predicted sound effect identifier (format: [Sound Effect Name]) at the emotional / scene transition point. The predicted sound effect identifier must match the semantics of the text. Original text: 'Mr. Chen rushed to her apartment, only to find it in a mess, with drawers ransacked.' The sound effect type should preferably be scene-atmosphere type (such as ambient sound, emotional auxiliary sound)." The sound effect name in "[Sound Effect Name]" can be vague, such as "[Script Sound]" or "[Thunder]", and does not need to be unique.

[0040] In the above embodiments, the first sound effect prompt word provides a clear sound effect matching direction for the first pre-trained language model, avoiding the generation of sound effect labels that are disconnected from the text due to vague requirements. This improves the adaptability of sound effect labels to text semantics and solves the problem of poor sound effect matching in traditional large-scale model processing. The first sound effect prompt word does not need to be associated with sound effect library data or restricted to whether the sound effect name is in the library. It only focuses on semantic adaptation, which saves the time cost of retrieving sound effect library data in the early stage and avoids the limitations of model output caused by sound effect library data constraints. At the same time, the prompt word can be automatically generated directly by combining the original text and timeline information, without the need for manual annotation of sound effect types or rules. This eliminates the dependence on professionals, reduces the cost of manual intervention, provides efficiency support for large-scale audiobook sound effect processing, and alleviates the industry pain points of low processing efficiency and difficulty in meeting large-scale needs.

[0041] Subsequently, intermediate text with fuzzy sound effect identifiers is generated based on the first sound effect cue word. This achieves an initial transformation from text to sound effect direction, providing clear sound effect type anchors for the subsequent rough matching of staking points and the sound effect library, avoiding aimless full-scale traversal in subsequent matching. Furthermore, the fuzziness of the sound effect identifiers retains the flexibility for subsequent filtering, preventing the omission of potentially compatible sound effects due to prematurely limiting the sound effect library. This design of first anchoring the direction and then precisely filtering effectively connects the initial model inference and the sound effect library matching stage, providing crucial transitional support for the efficiency and accuracy of the entire sound effect addition process, while avoiding the problems of cue word overload and slow production speed caused by directly calling the entire sound effect library.

[0042] Based on the above embodiments, when executing step S101, the original text is first input into the first pre-trained language model. After processing by the first pre-trained language model, at least one predicted sound effect identifier and at least one time stump of the predicted sound effect identifier are obtained. Then, according to the timeline information of the original text and the time stump of the at least one predicted sound effect identifier, at least one predicted sound effect identifier is embedded into the original text to obtain intermediate text. Here, the time stump is the time point at which the predicted sound effect identifier should appear in the original text.

[0043] The first pre-trained language model can infer the timing of sound effect tags embedded in the text. The original text is input into the first pre-trained language model, which predicts the appropriate sound effect tags and the timing when these sound effect tags should appear in the original text. Based on the timeline information of the text sequence in the original text and the timing points of these predicted sound effect tags, the predicted sound effect tags are embedded into the original text according to the timing points to obtain the intermediate text.

[0044] The above embodiments acquire the original audio and convert it into original text with timeline information, marking the playback time point for each character. This is equivalent to establishing a time coordinate system for the entire audio generation process. Based on this, a first pre-trained language model is used to determine the time markers for predicted sound effect icons, ensuring that the sound effect icons not only fit the text scene but also accurately correspond to the specific time position of the dubbing. This achieves precise synchronization between sound effects and dubbing in the time dimension, greatly enhancing the immersive experience of the audio. Furthermore, the intermediate text based on timeline information makes the subsequent generation process of sound effect files and dubbing more controllable, reducing repeated adjustments caused by time matching issues and further improving the efficiency and quality stability of audio generation. In addition, the clear time markers provide an orderly time reference for the superposition of multiple sound effects in complex scenes, avoiding the chaos caused by stacking sound effects.

[0045] In some embodiments, the first pre-trained language model can also infer the duration of sound effects. The original text is input into the pre-trained language model, which predicts appropriate sound effect identifiers and their time points, as well as the appropriate playback duration of the sound effects. Based on the timeline information of the original text, the first pre-trained language model embeds at least one predicted sound effect identifier into the original text according to the corresponding time points and marks the duration of the sound effects to obtain intermediate text.

[0046] The above embodiments allow the first pre-trained language model to directly predict the time points and playback durations of sound effect identifiers based on the timeline information of the original text, and embed them into the original text according to the time points and label the durations to form intermediate text. This means that sound effects can not only accurately correspond to specific time points in the dubbing, but also control the playback duration according to the time span of the text scene. From a technical perspective, based on the accurate synchronization of time points, the duration adaptation makes the time dimension of the sound effects completely consistent with the text description, avoiding scene distortion caused by sound effects that are too long or too short; the intermediate text contains both time points and duration information, providing more detailed parameter guidance for subsequent audio generation, reducing editing adjustments caused by duration mismatch during the generation process, and further improving production efficiency; for text scenes containing continuous actions or time spans, sound effects can start and end naturally with the development of the plot, significantly enhancing the narrative fluency and auditory immersion of the audio.

[0047] S102. Match candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library.

[0048] The predicted sound effect identifier is generated by the first pre-trained language model and may not be unique in the sound effect library. Therefore, this application searches for matching sound effect identifiers in the sound effect library as candidate sound effect information. The semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier is greater than or equal to the first preset similarity.

[0049] In some embodiments, when performing step S102, the correlation between the predicted sound effect identifier and any sound effect identifier in the sound effect library is first converted into a quantitative index. Then, sound effect identifiers whose quantitative index meets a preset threshold are selected, and their corresponding candidate sound effect information is determined.

[0050] Specifically, preprocessing can first eliminate format differences (such as character cleaning, format unification, and redundant information filtering) to ensure consistent matching foundations. Then, specific technical means (such as edit distance, semantic vectors, and keyword weights) can be used to transform the association between any sound effect identifier S0 in the sound effect library and the predicted sound effect identifier S1 into quantifiable indicators (such as similarity value, matching score, and probability value), realizing the transformation from fuzzy association to precise measurement. Finally, by setting and executing thresholds, sound effect identifiers S0 that meet the association criteria are selected as candidate sound effects. This avoids the waste of resources and inefficiency caused by directly entering the entire sound effect library into subsequent processing, while ensuring that the candidate sound effects have potential adaptability, providing high-quality input for accurate matching of subsequent pre-trained language models.

[0051] In some embodiments, when performing step S102, the similarity between the predicted sound effect identifier and any sound effect identifier in the sound effect library is first calculated. Simultaneously, the text length of the predicted sound effect identifier (labeled as the first text length for easy distinction) and the text length of this sound effect identifier in the sound effect library (labeled as the second text length) are obtained. Then, the first text length of the predicted sound effect identifier and the second text length of any sound effect identifier are summed to obtain the total text length; the similarity and the total text length are then compared to obtain the matching degree. If the matching degree is greater than or equal to a preset matching degree threshold, the candidate sound effect information for this sound effect identifier is obtained.

[0052] Specifically, for any sound effect identifier in the sound effect library, its second text length is determined, and the similarity (e.g., edit distance) between this sound effect identifier and the predicted sound effect identifier is calculated. Then, the first text length of the predicted sound effect identifier and the second text length of any sound effect identifier are summed to obtain the total text length; the similarity and the total text length are then compared to obtain the matching score. If the matching score is greater than or equal to a preset matching score threshold, candidate sound effect information for this sound effect identifier is obtained. If the matching score is less than the preset matching score threshold, the above operation is performed on another sound effect identifier in the sound effect library.

[0053] The preset matching threshold is determined by combining the size of the sound effects library and business needs, and needs to take into account both semantic relevance and filtering efficiency, such as 0.4-0.6.

[0054] For example, assuming a predicted sound effect identifier S1 and a sound effect identifier S0 in the sound effect library, the similarity d between sound effect identifier S0 and predicted sound effect identifier S1 is calculated. Then, based on the first text length d1 of predicted sound effect identifier S1, the second text length d0 of sound effect identifier S0, and the similarity d, the matching degree t = d / (d1+d0) is calculated. When the matching degree t is greater than or equal to a preset matching degree threshold, the sound effect information of sound effect identifier S0 is obtained as candidate sound effect information. For example, although the keywords "heavy rain sound" and "stormy rain sound" are not completely identical, their semantic similarity is high. After combining the text length calculation, the matching degree may exceed the threshold, thus being identified as candidate sound effect information.

[0055] Following the previous example, the candidate sound effects information for

script sound

rapid script sound, slow script sound, light script sound

thunder

booming thunder, soft thunder, distant muffled thunder

[0056] If a predicted sound effect identifier does not match a corresponding candidate sound effect, that identifier is removed from the intermediate text. Removing invalid identifiers avoids the awkward situation of having no corresponding sound effect during the generation process, ensuring the coherence and integrity of the audio.

[0057] The above embodiments calculate the similarity between predicted sound effect identifiers and identifiers in the sound effect library, combine the text lengths of both to calculate the matching degree, and set a matching degree threshold to filter candidate sound effect information, achieving more accurate semantic-level matching. By calculating the matching degree in multiple dimensions, the accuracy of matching sound effect identifiers with sound effect library resources is improved, making the called sound effects more in line with the needs of the text scene and reducing the problem of sound effects not matching the content due to matching deviations. This precise filtering mechanism reduces the necessity of manual intervention, improves the automation and efficiency of audio generation, optimizes the sound effect adaptability of the final generated audio, and further enhances the user's auditory immersion.

[0058] Optionally, a semantic similarity-based filtering scheme uses natural language processing techniques to understand the deep semantics of sound effect names, rather than relying solely on the edit distance of text characters. This is suitable for scenarios where sound effect names have "literal differences but similar meanings," such as "script sound" and "scene transition script sound." First, the stationary sound effect identifier S1 and the sound effect library identifier S0 undergo text preprocessing, including word segmentation (e.g., splitting "urgent script sound" into "urgent" and "script sound") and stop word filtering (removing meaningless words such as "of" and "type"). Then, a pre-trained language model (such as BERT or Sentence-BERT) is used to convert both into fixed-dimensional semantic vectors. Subsequently, the cosine similarity between the vectors is calculated, and finally, a cosine similarity threshold (usually 0.6-0.8) is set to filter out sound effects with similarity higher than the threshold as candidate sound effects. For example, when S1 is "thunder" and S0 is "rumbling thunder", the cosine similarity of their semantic vectors may reach 0.85, far exceeding the threshold, and thus be included in the candidate list; while when S0 is "raindrop sound", the cosine similarity may only be 0.2, and it will be filtered out.

[0059] The advantage of the above solution is that it can break through the limitations of literal characters and accurately capture the semantic relationship of sound effect names, which is especially suitable for scenarios where there are a large number of descriptive words and diverse semantic expressions in the sound effect library.

[0060] Optionally, the keyword matching and weighted scoring-based screening scheme focuses more on the direct association of core information in the sound effect name. It is suitable for scenarios where the sound effect name structure is well-organized and the core keywords are clear. For example, the sound effect identifier adopts the format of "attribute + core type", such as "urgent - script sound" and "slight - thunder". First, core keywords need to be labeled for each sound effect identifier S0 in the sound effect library. This can be done manually or by extracting keywords according to rules. For example, "rumbling" and "thunder" can be extracted from "rumbling thunder" as keywords, and "thunder" can be given a higher weight because it is the core identifier of the sound effect type. At the same time, keywords should be extracted for the pile point sound effect identifier S1. For example, "thunder" can be extracted from "thundering" as the core keyword. Then, keyword matching rules should be constructed to count the number of matches and matching weights of S1 keywords in S0 keywords. For example, "core keyword matching gets 3 points, attribute keyword matching gets 1 point". If the S1 keyword "thunder" matches the core keyword "thunder" of S0 "rumbling thunder", it can get 3 points. If S1 has the attribute keyword "urgent" and matches "urgent" of S0 "urgent script sound", it gets an extra 1 point. Finally, a total score threshold should be set to filter out S0 with a total score higher than the threshold as candidate sound effects. For example, when S1 is "script sound", the core keyword "script sound" is extracted. If it matches S0 "urgent script sound", it gets 3 points, which exceeds the threshold (e.g., 2 points) and becomes a candidate sound effect; if it does not match the keyword S0 "light footsteps", it gets 0 points and is filtered out.

[0061] The advantages of the above solution are its simplicity, high computational efficiency, and lack of complex models. It is suitable for small-scale sound effect libraries (less than 10,000 sound effects) or scenarios with high requirements for screening speed.

[0062] S103. Input the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information corresponding to the predicted sound effect information identifier. The target sound effect information is selected from the target sound effect information. The predicted sound effect identifier in the intermediate text is modified to the target sound effect information to obtain the target text.

[0063] The second pre-trained language model is a pre-trained large language model, which can be the same as or different from the first pre-trained language model. The target text is the text obtained by modifying the predicted sound effect identifiers in the intermediate text to the target sound effect information. The target sound effect information is selected from the candidate sound effect information corresponding to the predicted sound effect identifiers. The semantic feature similarity between the target sound effect information and the intermediate text is greater than or equal to the second preset similarity.

[0064] In some embodiments, for candidate sound effect information and intermediate text, the predicted sound effect identifiers that match the candidate sound effect information in the intermediate text are first replaced, while the predicted sound effect identifiers that do not match the candidate sound effect information are deleted, avoiding problems caused by invalid identifiers in subsequent audio generation. Then, the processed intermediate text is input into a second pre-trained language model, which infers more suitable target sound effect information and outputs target text containing the target sound effect information. It should be noted that the target sound effect information is obtained based on candidate sound effect information existing in the sound effect library, and the target sound effect information also exists in the sound effect library, ensuring the feasibility of subsequent calls.

[0065] The above embodiment inputs the processed intermediate text into a second pre-trained language model, allowing the model to infer more suitable target sound effect information based on existing candidate sound effect information in the sound effect library. The model combines the context and emotional details of the original text to select the most fitting option (i.e., the target sound effect information) from the candidate sound effect information. From a technical perspective, on the one hand, through secondary inference by the second pre-trained language model, the sound effect information is accurately adapted, improving the fit between the sound effect and the text scene, allowing the audio to better convey the details and emotions of the text; on the other hand, it avoids arbitrary selection caused by multiple available options in the candidate sound effect information, making the selection of sound effects more targeted, and further enhancing the immersiveness and expressiveness of the audio.

[0066] In some embodiments, when performing step S103, a second sound effect cue word that semantically matches the intermediate text is first generated based on the candidate sound effect information corresponding to the predicted sound effect identifier. The second sound effect cue word contains constraints for selecting the target sound effect information from the candidate sound effect information. Then, the intermediate text and the second sound effect cue word are input into a second pre-trained language model, which processes the data to obtain the target text. The second sound effect cue word combines the scene details of the intermediate text to clarify the features required for the sound effect, such as intensity, style, and duration.

[0067] Specifically, based on the candidate sound effect information corresponding to the predicted sound effect identifier and the semantics of the intermediate text, constraints for selecting the target sound effect information from the candidate sound effect information are determined, and a second sound effect prompt word containing these constraints is generated. The candidate sound effect information corresponding to the predicted sound effect identifier is within the sound effect library.

[0068] The constraints require that the semantic feature similarity between the target sound effect information and the intermediate text is greater than or equal to the second preset similarity, and that the semantic depth is adapted (e.g., "hurried" needs to be matched with an emergency-style sound effect). They also require that the target sound effect identifier contained in the target sound effect information is unique in the sound effect library, so that the second pre-trained language model can select the uniquely adapted target sound effect information from the limited candidate sound effect information, and avoid the second pre-trained language model generating sound effects outside the sound effect library.

[0069] Then, the intermediate text and the second sound effect prompt are input into the second pre-trained language model so that the second pre-trained language model can infer the appropriate target sound effect information. It should be noted that this is actually a screening of candidate sound effect information, retaining the appropriate candidate sound effect information as the target sound effect information, so as to output the target text.

[0070] Following the previous example, the intermediate text is "Ms. Chen rushed to her apartment with a [script sound], only to find the apartment in a mess, with drawers ransacked [thunder]". The second sound effect cue can be: "Please select the unique target sound effect information that matches the predicted sound effect identifier in the intermediate text from the following candidate sound effect information: Intermediate text 'Ms. Chen rushed to her apartment with a [script sound]...', candidate sound effect information: [urgent script sound, slow script sound, booming thunder, light thunder], constraint: sound effect intensity matches text emotion ('rushed' corresponds to urgency, 'finds the apartment in a mess' corresponds to strong conflict)". Subsequently, the model filters the candidate sound effect information based on the intermediate text and the second sound effect cue, retaining the identifier that best matches the requirements of the second sound effect cue as the target sound effect information. "Script sound" is refined to "urgent script sound", and "thunder" is refined to "booming thunder". The target text is "Mr. Chen rushed to her apartment, only to find it in a mess, with drawers ransacked and in disarray."

[0071] In the above embodiments, the second sound effect prompt explicitly incorporates constraints for semantic depth adaptation with the intermediate text. This requires the model to select sound effects based on text details, enabling the model to upgrade from fuzzy type matching to refined semantic matching. This avoids sound effect style mismatches caused by the model's misunderstanding of text details, solving the problem of adaptation bias caused by fuzzy labels in traditional semi-automatic methods. Simultaneously, the constraint that the target sound effect identifier is unique in the sound effect library directly ensures that the model output corresponds to the actual sound effect files in the library, eliminating the need for subsequent manual secondary matching and further improving the accuracy and practicality of sound effect addition.

[0072] The second set of sound effect prompts limits the model's selection to only the range of candidate sound effect information. This avoids the problem of prompt word overload caused by the model calling the entire sound effect library, eliminating the need to include a large amount of irrelevant sound effect information from the library as input. This reduces the amount of data processing during model inference, improves response speed, and alleviates the pain point of slow production speed when large models process sound effects in the industry. In addition, the explicit constraints reduce the model's decision uncertainty, reduce invalid inference caused by an overly broad output range, and allow the model to focus on the core matching task, further optimizing inference efficiency and adapting to the batch processing needs of large-scale audiobook sound effects.

[0073] The constraint of the second sound effect prompt word to avoid generating sound effects outside the sound effect library eliminates invalid sound effects that cannot be implemented from the source, ensuring that each output target sound effect identifier can be directly associated with the sound effect file address, providing directly usable materials for the subsequent sound effect synthesis link, and achieving seamless connection from model inference to audio generation. This design not only ensures the closed-loop nature of the technical solution, but also reduces the risk of process interruption caused by missing sound effects. At the same time, it avoids the additional cost of manually correcting invalid sound effects, making the entire sound effect addition process more in line with the stability and reliability requirements of industrial production.

[0074] S104. Generate the target audio according to the sound effect file corresponding to the target text and the target sound effect information.

[0075] Among them, the target text contains the time axis information of the text sequence and the target sound effect information. Here, the time axis information includes the playback time points of each word in the text sequence and the playback time points of the sound effect file corresponding to the target sound effect information. The playback time point includes the start time and the end time. The generated target audio can accurately play each word and add sound effects at the appropriate position, realizing the optimization of the original audio.

[0076] The target text contains both the逐字 playback time points of the text sequence and the playback time points of the sound effect file, and both of them clearly define the start and end times, which can ensure that the sound effect embedding position is completely synchronized with the text content. For example, the playback time of the character "陈" in the text is 0 - 0.39 seconds, and the character "某" is 0.39 - 0.57 seconds. The corresponding [urgent script sound] is set to 0.39 - 0.89 seconds (covering the core text segment of the "急忙" action). It will neither be inserted too early, resulting in the disconnection between the sound effect and the text, nor be inserted too late, missing the plot atmosphere. It completely solves the problem of time-consuming time axis adjustment in traditional manual addition or the deviation of sound effect insertion position in semi-automated tagging, realizes the precise alignment of sound effects and text, and ensures the coherence and logic of the audio content.

[0077] The generated target audio can accurately play each word and add sound effects at the appropriate position. The presence of sound effects will not only interfere with the normal listening of the text content, but also enhance the scene substitution and emotional transmission power of the audiobook through the auditory attributes of the sound effects (such as the "urgent script sound" strengthening the sense of action tension and the "roaring thunder" setting off the conflict atmosphere). For example, inserting the "roaring thunder" at the plot of "finding the apartment in a mess" can magnify the character's shocked emotion through the strong impact of the sound effect, making the listener more easily immersed in the story scene. Compared with the original audio without optimization or the audio with mismatched sound effects, it improves the user's auditory experience and content acceptance.

[0078] This step generates the final audio based on the target text and target sound effect files output from the previous steps, requiring no additional manual intervention. It achieves end-to-end automation from raw audio to text conversion, sound effect matching, and target audio generation. This closed-loop design avoids efficiency losses caused by process interruptions and ensures that each step's output directly serves the final result, meeting the high-efficiency and standardized requirements of large-scale audiobook production. Simultaneously, the generated target audio is a direct optimization of the original audio, which can be directly applied to actual business scenarios such as platform launch and content distribution without secondary processing, reducing the cost of implementing the technical solution and enhancing its commercial application value.

[0079] In some embodiments, after executing step S103 and before executing step S104, the corresponding sound effect address is first searched in the sound effect library according to the target sound effect information, and then the sound effect file corresponding to the target sound effect information is obtained according to the sound effect address.

[0080] The sound effects library stores various sound effect identifiers and their corresponding sound effect addresses. Target sound effect information exists in the library; its corresponding sound effect address can be found based on the target sound effect information, and then the corresponding sound effect file can be retrieved by indexing the corresponding location. For the content in the target text other than the target sound effect information, text-to-speech technology is used to obtain candidate audio, which is then combined with the sound effect file corresponding to the target sound effect information to generate the target audio. This optimizes the sound effect effects of the text. Target audio can be generated using Moviepy.

[0081] In the above embodiments, the sound effect library stores sound effect identifiers and corresponding sound effect addresses. Target sound effect information can directly look up the sound effect address and index to obtain the sound effect file based on this. This process is fast and accurate, avoiding delays and errors when calling sound effect files. For audio generation, candidate audio is first obtained through text-to-speech technology, and then generated by combining the sound effect file corresponding to the target sound effect information and timeline information, making the combination of dubbing and sound effects more targeted. Since the target sound effect information is precisely selected, its corresponding sound effect file highly matches the text content, allowing for better integration with candidate audio during generation. In summary, this application provides a method for adding sound effects to audio. This method addresses the problems of manual addition relying on human intervention, low efficiency, and difficulty in meeting large-scale needs. It achieves end-to-end automated processing from raw audio to target audio without human intervention. First, the original text with a timeline is obtained through raw audio conversion. Then, the predicted sound effect labels are automatically embedded by the first pre-trained language model. Subsequent candidate sound effect matching, target sound effect selection, and audio synthesis are all completed automatically by the model and system, completely eliminating the reliance on the professional skills of audio engineers and significantly improving processing efficiency. It can efficiently adapt to large-scale audio processing needs. Addressing the problems of label-assisted semi-automation methods over-relying on preset rules and label accuracy, large sound effect matching deviations in complex scenarios, and the need for extensive manual correction, this method replaces fixed labels and rules with the semantic understanding capabilities of the pre-trained language model. Trained based on sample text and corresponding sound effect labels, it can deeply understand the semantic, emotional, and scene features of the text, accurately predict the insertion position and features of sound effects, and avoid the problems caused by labeling bias. At the same time, by matching candidate sound effects with semantic similarity standards from the sound effect library, and further filtering target sound effects with semantic similarity standards with intermediate text by a second pre-trained language model, a double guarantee of coarse screening and fine selection is formed. Even when facing complex emotions or mixed scenes, it can achieve accurate matching through the model's deep analysis of text details, significantly reducing the need for manual correction and fundamentally making up for the core defects of the two mainstream methods.

[0082] In the text processing and sound effect generation stages, a pre-trained language model processes the original text, combines it with sound effect library information to generate and embed predicted sound effect tags with time markers and durations. First and second sound effect prompts guide the model, ensuring the sound effect tags fit the text context and match the sound effect library resources, improving the adaptability of the sound effect tags to the text and the sound effect library. In the sound effect matching stage, the similarity between the predicted sound effect tags and the sound effect library tags, along with the text length, is calculated to determine the matching degree and filter candidate sound effect information, removing invalid tags that do not match. This mechanism makes sound effect calls more precise, avoids problems caused by invalid tags, and improves the accuracy of sound effect matching. Regarding time synchronization, the model predicts the time markers and durations of the sound effects, combines them with the embedded tags on the original text timeline, ensuring the sound effects appear at the appropriate time and for the appropriate duration, accurately corresponding to the moments described in the text, enhancing audio coherence. In the final generation stage, the files are quickly obtained based on the sound effect addresses corresponding to the target sound effect information in the sound effect library. The candidate audio and sound effect files are accurately synthesized. Because the sound effect information has been optimized in multiple rounds, the sound effects and dubbing in the generated audio are highly coordinated, which improves the overall expressiveness and immersion.

[0083] The training process for the pre-trained language model in this application includes: First, a massive amount of audiobook manuscripts and corresponding sound effect annotation data were collected, including sound effect insertion positions, sound effect types, and applicable scenarios, to construct a training dataset. The manuscripts needed to cover different themes (such as suspense, romance, and science) and different emotional tones (such as tension, warmth, and conflict). Sound effect annotations should clearly define the correspondence between "text fragment - sound effect type - sound effect name," for example, "rapid action description - script sound - rapid script sound." Simultaneously, the training dataset underwent preprocessing, including manuscript word segmentation, redundant information filtering, and sound effect tag standardization to ensure a consistent data format.

[0084] Subsequently, a basic pre-trained language model (such as BERT or GPT series models) was selected as the foundation, and fine-tuning was performed based on the constructed training dataset. The training objective can be set as "to enable the model to learn the mapping relationship between text semantics and sound effect type and sound effect identifier," that is, when given text with a timeline, the model can output the appropriate sound effect insertion position and corresponding sound effect identifier. During training, the error between the model's prediction results and the labeled data can be calculated using the cross-entropy loss function. The model parameters can be continuously optimized until the accuracy of sound effect matching and the accuracy of sound effect insertion position on the validation set reach the preset standards, thus completing the training.

[0085] In addition, business rules for audiobook sound effect synthesis can be incorporated into the training process, such as matching the duration of sound effects with the text playback time and ensuring that the volume of sound effects does not cover the character's voice, so as to ensure that the model output meets the needs of actual applications.

[0086] The reasoning process of a pre-trained language model includes: During inference, the first pre-trained language model takes the original text from step 2 (including the timeline information of the text sequence, i.e., the timestamps for word-by-word playback) as input, combined with the first sound effect cue words (explicitly requiring the model to insert predicted sound effect identifiers (format: [sound effect name]) at appropriate positions based on text semantics, emotion, and plot transitions; these sound effect names need to fit the scene but do not need to be unique). After receiving the input, the first pre-trained model first parses the semantic logic and emotional tone of the original text (e.g., "hurriedly" corresponds to a tense scene, "finding the apartment in a mess" corresponds to a conflict scene), and then, based on the "text-sound effect" mapping relationship learned during training, identifies the key positions for sound effect insertion, generates fuzzy predicted sound effect identifiers, and finally outputs the intermediate text embedded with the predicted sound effect identifiers. Throughout the process, the first pre-trained language model does not need to associate with a sound effect library; it only focuses on the initial matching of text semantics and sound effect types. The core is to quickly anchor the direction of sound effect insertion, providing a foundation for subsequent rough matching.

[0087] The second pre-trained language model takes intermediate text and candidate sound effect information corresponding to predicted sound effect identifiers as input for inference, along with optimized second sound effect cue words. These cue words explicitly require the model to select a uniquely matching target sound effect from the candidate sound effect information, satisfying both semantic fit and a unique correspondence with the sound effect library. After receiving the input, the model first deeply analyzes the text details corresponding to each predicted sound effect identifier in the intermediate text (e.g., the intensity of the action "hurriedly"), then compares the features of each sound effect in the candidate sound effect information (e.g., the rhythmic features of "urgent script sound" and the intensity features of "rumbling thunder"). Combining this with the refined matching logic learned during training, the model selects a uniquely matching target sound effect for each predicted sound effect identifier, ultimately outputting target text with the unique target sound effect information. In this process, the model is limited to making decisions within the scope of candidate sound effect information, avoiding cue word inflation and ensuring that the output sound effect can be directly associated with the file address in the sound effect library, providing accurate input for subsequent audio synthesis. From the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Figure 2 As shown, embodiments of this application also provide an apparatus for adding sound effects to audio, the apparatus comprising: The speech-to-text module 200 is used to convert the original audio into the original text, which contains the timeline information of the text sequence. The intermediate text generation module 201 is used to input the original text into the first pre-trained language model, and after processing by the first pre-trained language model, to obtain the intermediate text after embedding at least one predicted sound effect label into the original text; wherein, the first pre-trained language model is trained based on sample text and corresponding labeled sound effect labels; the sound effect labels are used to identify the sound effect insertion position in the original text and its corresponding sound effect features. Matching module 202 is used to match candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library; the semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier is greater than or equal to the first preset similarity; the candidate sound effect information includes multiple sound effect identifiers and sound effect addresses corresponding to the multiple sound effect identifiers. The target text generation module 203 is used to input the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier, whose semantic feature similarity to the intermediate text is greater than or equal to the second preset similarity, and modifies the predicted sound effect identifier in the intermediate text to the target sound effect information to obtain the target text. The target text includes the time axis information of the text sequence and the target sound effect information. The audio generation module 204 is used to generate target audio based on the target text and the sound effect file corresponding to the target sound effect information.

[0088] This application provides an optional implementation method, in which the intermediate text generation module 201 is specifically used to: input the original text into a first pre-trained language model, and after processing by the first pre-trained language model, obtain at least one predicted sound effect identifier and at least one time point of the predicted sound effect identifier; the time point is the time point when the predicted sound effect identifier appears in the original text; and embed at least one predicted sound effect identifier into the original text according to the timeline information and the time point to obtain intermediate text.

[0089] This application provides an optional implementation method, in which the intermediate text generation module 201 is specifically used to: generate a first sound effect cue word related to the scene information or emotional information of the original text, wherein the first sound effect cue word is used to describe the auditory attributes of the sound effect and includes adjectives defining sound attributes and verbs defining sound dynamics; input the original text and the first sound effect cue word into a first pre-trained language model, and obtain intermediate text after processing by the first pre-trained language model by embedding at least one predicted sound effect identifier into the original text.

[0090] This application provides an optional implementation method, in which the matching module 202 is specifically used to: convert the association between the predicted sound effect identifier and any sound effect identifier in the sound effect library into a quantitative index; filter out the sound effect identifiers whose quantitative index meets a preset threshold, and determine their corresponding candidate sound effect information.

[0091] This application provides an optional implementation method, in which the matching module 202 is specifically used to: calculate the similarity between the predicted sound effect identifier and any sound effect identifier in the sound effect library; sum the first text length of the predicted sound effect identifier and the second text length of any sound effect identifier to obtain the total text length; perform a ratio operation on the similarity and the total text length to obtain the matching degree; and, if the matching degree is greater than or equal to a preset matching degree threshold, obtain the candidate sound effect information corresponding to any sound effect identifier.

[0092] This application provides an optional implementation method. The target text generation module 203 is specifically used to: generate a second sound effect prompt word that semantically matches the intermediate text based on the candidate sound effect information corresponding to the predicted sound effect identifier, wherein the second sound effect prompt word contains constraints for selecting target sound effect information from the candidate sound effect information; input the intermediate text and the second sound effect prompt word into a second pre-trained language model, wherein the second pre-trained language model selects target sound effect information from the candidate sound effect information according to the constraints, and modifies the predicted sound effect identifier in the intermediate text to the target sound effect information to obtain the target text.

[0093] This application provides an optional implementation method, in which the audio generation module 204 is further configured to: search for the corresponding sound effect address in the sound effect library based on the target sound effect information; and obtain the sound effect file corresponding to the target sound effect information based on the sound effect address. Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0094] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0095] For example, such as Figure 3 As shown, the electronic device includes a memory 301 and a processor 302. The memory 301 stores executable program code, and the processor 302 is used to call and execute the executable program code to perform a method of adding sound effects to audio.

[0096] This embodiment can divide the electronic device into functional modules based on the above method example. For example, each function can be assigned to a separate module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. When dividing each functional module according to its corresponding function, the electronic device may include: an intermediate text generation module, a matching module, a target text generation module, and an audio generation module, etc. It should be noted that all relevant content of each step involved in the above method embodiment can be referenced to the functional description of the corresponding functional module, and will not be repeated here.

[0097] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the method for adding sound effects to audio when it is run.

[0098] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0099] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the method for adding sound effects to audio.

[0100] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the method for adding sound effects to audio.

[0101] The beneficial effects of the above embodiments can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0102] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0103] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0104] In the description of this application, it should be understood that if the terms "upper", "lower", "front", "rear", "left" and "right" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the position or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0106] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for adding sound effects to audio, characterized in that, include: The original text is obtained by converting the original audio, and the original text contains timeline information of the text sequence. The original text is input into a first pre-trained language model, and after processing by the first pre-trained language model, intermediate text is obtained that embeds at least one predicted sound effect label into the original text; wherein, the first pre-trained language model is trained based on sample text and corresponding labeled sound effect labels; the sound effect labels are used to identify the sound effect insertion position in the text and its corresponding sound effect features. Match candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library; the semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier is greater than or equal to a first preset similarity; the candidate sound effect information includes multiple sound effect identifiers and sound effect addresses corresponding to the multiple sound effect identifiers; The candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text are input into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier. The target sound effect information has a semantic feature similarity to the intermediate text that is greater than or equal to a second preset similarity. The predicted sound effect identifier in the intermediate text is then modified to the target sound effect information to obtain the target text. The target text includes the time axis information of the text sequence and the target sound effect information. The target audio is generated based on the target text and the sound effect file corresponding to the target sound effect information.

2. The method according to claim 1, characterized in that, The step of inputting the original text into a first pre-trained language model, and obtaining intermediate text by processing it with at least one predicted sound effect identifier embedded in the original text, includes: The original text is input into the first pre-trained language model, and after processing by the first pre-trained language model, at least one predicted sound effect identifier and the time stipulation of the at least one predicted sound effect identifier are obtained; the time stipulation is the time point when the predicted sound effect identifier appears in the original text. Based on the timeline information and the time markers, the at least one predicted sound effect identifier is embedded into the original text to obtain the intermediate text.

3. The method according to claim 1, characterized in that, The step of inputting the original text into a first pre-trained language model, and obtaining intermediate text by processing it with at least one predicted sound effect identifier embedded in the original text, includes: Generate a first sound effect cue word related to the scene information or emotional information of the original text. The first sound effect cue word is used to describe the auditory attributes of the sound effect and includes adjectives that define the sound attributes and verbs that define the sound dynamics. The original text and the first sound effect prompt are input into the first pre-trained language model, and the first pre-trained language model processes them to obtain intermediate text after embedding at least one predicted sound effect identifier into the original text.

4. The method according to claim 1, characterized in that, The step of matching the candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library includes: The association between the predicted sound effect identifier and any sound effect identifier in the sound effect library is converted into a quantitative indicator; Sound effect identifiers that meet the preset threshold of the quantitative indicators are selected, and their corresponding candidate sound effect information is determined.

5. The method according to claim 1, characterized in that, The step of matching the candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library includes: Calculate the similarity between the predicted sound effect identifier and any sound effect identifier in the sound effect library; The total text length is obtained by summing the first text length of the predicted sound effect identifier and the second text length of any sound effect identifier. The matching degree is obtained by calculating the ratio between the similarity score and the total text length. If the matching degree is greater than or equal to a preset matching degree threshold, obtain the candidate sound effect information corresponding to any sound effect identifier.

6. The method according to claim 1, characterized in that, The step of inputting the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into a second pre-trained language model, processing it, and then selecting target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier with a semantic feature similarity greater than or equal to a second preset similarity with the intermediate text, and modifying the predicted sound effect identifier in the intermediate text to the target sound effect information to obtain the target text includes: Based on the candidate sound effect information corresponding to the predicted sound effect identifier, a second sound effect prompt word that semantically matches the intermediate text is generated. The second sound effect prompt word contains constraints for selecting target sound effect information from the candidate sound effect information. The intermediate text and the second sound effect prompt are input into the second pre-trained language model. The second pre-trained language model selects the target sound effect information from the candidate sound effect information according to the constraints, and modifies the predicted sound effect identifier in the intermediate text to the target sound effect information to obtain the target text.

7. The method according to claim 1, characterized in that, The method further includes the following steps: inputting the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into a second pre-trained language model; selecting target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier with a semantic feature similarity greater than or equal to a second preset similarity with the intermediate text; modifying the predicted sound effect identifier in the intermediate text to obtain the target sound effect information; and generating the target audio based on the target text, the sound effect file corresponding to the target sound effect information, and the timeline information. Based on the target sound effect information, the corresponding sound effect address is searched from the sound effect library; The sound effect file corresponding to the target sound effect information is obtained based on the sound effect address.

8. A device for adding sound effects to audio, characterized in that, include: The speech-to-text module is used to convert the original audio into original text, which contains timeline information of the text sequence. An intermediate text generation module is used to input the original text into a first pre-trained language model, and the first pre-trained language model processes the original text to obtain intermediate text after embedding at least one predicted sound effect label; wherein, the first pre-trained language model is trained based on sample text and corresponding labeled sound effect labels; the sound effect labels are used to identify the sound effect insertion position in the text and its corresponding sound effect features. The matching module is used to match candidate sound effect information corresponding to the predicted sound effect identifier from the sound effect library; the semantic feature similarity between the candidate sound effect information and the predicted sound effect identifier is greater than or equal to a first preset similarity; the candidate sound effect information includes multiple sound effect identifiers and sound effect addresses corresponding to the multiple sound effect identifiers. The target text generation module is used to input the candidate sound effect information corresponding to the predicted sound effect identifier and the intermediate text into a second pre-trained language model. The second pre-trained language model selects target sound effect information from the candidate sound effect information corresponding to the predicted sound effect identifier, whose semantic feature similarity to the intermediate text is greater than or equal to a second preset similarity. The predicted sound effect identifier in the intermediate text is then modified to the target sound effect information to obtain the target text. The target text includes the time axis information of the text sequence and the target sound effect information. The audio generation module is used to generate target audio based on the target text and the sound effect file corresponding to the target sound effect information.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the method for adding sound effects to audio as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for adding sound effects to audio as described in any one of claims 1 to 7.