Audio editing method and device, electronic equipment, readable storage medium and program product
By obtaining the difference between audio and text, and using a pre-trained encoder to process audio, the problems of difficulty and inefficiency in existing audio editing techniques are solved, and a user-friendly audio editing method is realized.
Patent Information
- Application Number
- CN202510316676.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
AI Technical Summary
The existing audio editing technology is difficult to get started and is inefficient, so users need to master professional skills.
By obtaining the difference between the first text and the second text, the audio is processed using a pre-trained audio encoder and a text encoder to generate an audio effect that meets the user's expectations.
Users can generate audio effects that meet expectations without mastering audio editing technology, achieving flexible editing and efficiency improvement.
Smart Images

Figure CN120279880A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an audio editing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of computer technology and audio technology, audio editing technology has emerged. Through audio editing, conventional audio processing such as cutting, copying, pasting, multi-file merging, and mixing can be achieved.
[0003] In the related art of audio editing, users need to operate through professional audio editing software, which has a high learning curve and a relatively high threshold, reducing the audio editing efficiency.
[0004] Therefore, in the related art, there is a problem of low audio editing efficiency. Summary of the Invention
[0005] Based on this, it is necessary to provide an audio editing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can improve audio editing efficiency for the above technical problems.
[0006] In a first aspect, this application provides an audio editing method, including:
[0007] Obtain a first text and a second text; the first text is a text describing a first audio; the second text is a text describing a desired audio, and the desired audio is the audio that the first audio is expected to be modified into;
[0008] Process the first audio according to the difference between the first text and the second text to obtain a second audio.
[0009] In one embodiment, the process of processing the first audio according to the difference between the first text and the second text to obtain a second audio includes:
[0010] Input the first audio into a pre-trained audio encoder to obtain a first audio feature;
[0011] Adjust the first audio feature according to the difference between the first text and the second text to obtain a second audio feature;
[0012] Generate the second audio according to the second audio feature.
[0013] In one embodiment, the process of adjusting the first audio feature according to the difference between the first text and the second text to obtain a second audio feature includes:
[0014] Obtain a first text feature corresponding to the first text and a second text feature corresponding to the second text through a pre-trained text encoder; both the first text feature and the second text feature are in the same feature space as the first audio feature;
[0015] Adjust the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature.
[0016] In one embodiment, the adjusting the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature includes:
[0017] Obtain the difference between the first text feature and the second text feature;
[0018] Adjust the first audio feature according to the difference to obtain the second audio feature.
[0019] In one embodiment, the pre-trained audio encoder and the pre-trained text encoder adopt the following training method:
[0020] Obtain first training sample data; the first training sample data includes a first sample audio and a corresponding sample text; the sample text is the text describing the first sample audio;
[0021] Iteratively train the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text;
[0022] When the trained text encoder and the trained audio encoder meet a preset first training end condition, obtain the pre-trained text encoder and the pre-trained audio encoder.
[0023] In one embodiment, the iteratively training the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text includes:
[0024] Input the first sample audio into the audio encoder to be trained to obtain a first sample audio feature;
[0025] Input the sample text corresponding to the first sample audio into the text encoder to be trained to obtain a sample text feature;
[0026] Obtain a preset loss function, and determine the loss value of the loss function according to the difference between the first sample audio feature and the sample text feature;
[0027] Adjust the model parameters of the text encoder to be trained and the model parameters of the audio encoder to be trained according to the loss value, and continue iterative training until the trained text encoder and the trained audio encoder meet the first training end condition, then stop training.
[0028] In one embodiment, the generating the second audio according to the second audio feature includes:
[0029] Input the second audio feature into a pre-trained audio decoder to obtain the second audio;
[0030] Wherein, the pre-trained audio decoder adopts the following training method:
[0031] Obtain second training sample data; the second training sample data includes a second sample audio and a corresponding second sample audio feature; the second sample audio feature is obtained by encoding the second sample audio with the pre-trained audio encoder;
[0032] Input the second sample audio feature corresponding to the second sample audio into an audio decoder to be trained to obtain a new sample audio;
[0033] According to the difference between the second sample audio and the new sample audio, adjust the model parameters of the audio decoder to be trained and continue iterative training until the trained audio decoder meets the second training end condition, then stop training to obtain the pre-trained audio decoder.
[0034] In a second aspect, the present application further provides an audio editing device, including:
[0035] An acquisition module, configured to acquire a first text and a second text; the first text is a text describing a first audio; the second text is a text describing a desired audio, and the desired audio is the audio that the first audio is expected to be modified into;
[0036] A processing module, configured to process the first audio according to the difference between the first text and the second text to obtain a second audio.
[0037] In a third aspect, the present application further provides an electronic device. The electronic device includes a memory and a processor, and the memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0039] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps of the above-described method.
[0040] For the above audio editing method, apparatus, electronic device, computer-readable storage medium, and computer program product, a first text and a second text are obtained; the first text is a text describing a first audio; the second text is a text describing a desired audio, where the desired audio is the audio that the first audio is desired to be modified into; and the first audio is processed according to the difference between the first text and the second text to obtain a second audio.
[0041] In this way, by obtaining the first text for describing the first audio and the second text for describing the audio that the first audio is desired to be modified into, the first audio can be processed according to the difference between the first text and the second text, so that the obtained second audio is more matched with the second text. Since the second text is used to describe the audio that the first audio is desired to be modified into, the actually generated second audio can more meet the audio effect expected by the user. The user does not need to master relevant audio editing technical content. Just by using the text for describing the audio, an audio that meets the audio effect expected by the user can be generated, satisfying the user's audio editing needs, realizing flexible editing of the audio, and improving the audio editing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is an application environment diagram of an audio editing method in an embodiment;
[0044] Figure 2 It is a schematic flowchart of an audio editing method in an embodiment;
[0045] Figure 3 It is an iterative training schematic diagram of a to-be-trained audio encoder and a to-be-trained text encoder in an embodiment;
[0046] Figure 4 It is a training process schematic diagram of a to-be-trained audio decoder in an embodiment;
[0047] Figure 5 It is a schematic flowchart of an audio editing method in another embodiment;
[0048] Figure 6 It is a structural block diagram of an audio editing device in an embodiment;
[0049] Figure 7 It is an internal structure diagram of an electronic device in an embodiment. Specific embodiments
[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0052] In one embodiment, Figure 1 It can be a schematic diagram of the implementation environment involved in the audio editing method. The implementation environment at least includes a user terminal 110, a smart device 130, a server end 170, and a network device. In Figure 1 it, the network device includes a gateway 150 and a router 190, and this is not a specific limitation here.
[0053] Among them, the user terminal 110, which can also be regarded as a user side or a terminal, can deploy (also understood as install) a client associated with the smart device 130. This user terminal 110 can be an electronic device such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart control panel, or other devices with display and control functions, and is not limited here.
[0054] Among them, the client is associated with the intelligent device 130. Essentially, the user registers an account in the client and configures the intelligent device 130 in the client. For example, the configuration includes adding a device identifier to the intelligent device 130, etc., so that when the client runs in the user terminal 110, functions such as device display and device control of the intelligent device 130 can be provided to the user. This client can be in the form of an application or a web page. Correspondingly, the interface for the client to perform device display can be in the form of a program window or a web page, and this is not limited here either.
[0055] The intelligent device 130 is deployed in the gateway 150 and communicates with the gateway 150 through its own configured communication module, and thus is controlled by the gateway 150. It should be understood that the intelligent device 130 generally refers to one of multiple intelligent devices 130. Only the intelligent device 130 is taken as an example in the embodiments of the present application. That is, the embodiments of the present application do not limit the number and device type of the intelligent devices deployed in the gateway 150. In an application scenario, the intelligent device 130 accesses the gateway 150 through a local area network and is thus deployed in the gateway 150. The process by which the intelligent device 130 accesses the gateway 150 through a local area network includes: the gateway 150 first establishes a local area network, and the intelligent device 130 joins the local area network established by the gateway 150 by connecting to the gateway 150. This local area network includes but is not limited to: ZIGBEE or Bluetooth. Among them, the intelligent device 130 can be but is not limited to various smart home devices (or Internet of Things devices), such as smart printers, smart fax machines, smart cameras, smart air conditioners, smart door locks, smart lights, or human body sensors, door and window sensors, temperature and humidity sensors, water immersion sensors, natural gas alarms, smoke alarms, wall switches, wall sockets, wireless switches, wireless wall sticker switches, magic cube controllers, curtain motors, millimeter wave radars, etc. configured with communication modules.
[0056] The interaction between the user terminal 110 and the intelligent device 130 can be achieved through a local area network or a wide area network. In an application scenario, the user terminal 110 establishes a communication connection with the gateway 150 through the router 190 in a wired or wireless manner, etc. For example, the wired or wireless manner includes, but is not limited to, WIFI, etc., so that the user terminal 110 and the gateway 150 are deployed in the same local area network, and thus the user terminal 110 can interact with the intelligent device 130 through the local area network path. In another application scenario, the user terminal 110 establishes a communication connection with the gateway 150 through the server 170 in a wired or wireless manner, etc. For example, the wired or wireless manner includes, but is not limited to, 2G, 3G, 4G, 5G, WIFI, etc., so that the user terminal 110 and the gateway 150 are deployed in the same wide area network, and thus the user terminal 110 can interact with the intelligent device 130 through the wide area network path.
[0057] Among them, the server 170 can also be considered as the cloud, cloud platform, platform side, service side, etc. This server 170 can be a single server, or a server cluster composed of multiple servers, or a cloud computing center composed of multiple servers.
[0058] In an application scenario, the user terminal 110 obtains a first text and a second text; among them, the first text is the text describing the first audio; among them, the second text is the text describing the expected audio, and the expected audio is the audio that the first audio is expected to be modified into; the user terminal 110 processes the first audio according to the difference between the first text and the second text to obtain the second audio.
[0059] In one embodiment, as Figure 2 shown, an audio editing method is provided. In this embodiment, the method is exemplified by being applied to an electronic device. The electronic device can be Figure 1 the user terminal 110, server 170, gateway 150, intelligent device 130, etc. in. In this embodiment, the method includes the following steps:
[0060] Step S210, obtain a first text and a second text.
[0061] Among them, the first text is the text describing the first audio.
[0062] Among them, the first audio refers to the audio data to be processed, and the processing specifically refers to editing the audio data.
[0063] Among them, the first text is used to describe the first audio. Specifically, it can be the text for describing the audio content of the first audio. For example, if the first audio is the chirping sound of birds in a park, the first text can be "There is the chirping sound of birds in the park."
[0064] In some other embodiments, the first text may include various elements in the audio, such as sound type, rhythm, mood, scene, and many other aspects. For example, if the first audio is the chirping sound of birds in a forest, the first text can be "This audio depicts the scene of the forest in the early morning, with the clear chirping sound of birds, the pecking sound of woodpeckers, and the calls of various birds intertwined, with a distinct rhythm. The overall atmosphere is peaceful and full of vitality, giving people a feeling of being deep in a quiet forest."
[0065] In practical applications, the text description of the first audio can be automatically generated by an audio understanding model to obtain the first text. Specifically, by inputting the first audio into the audio understanding model, the obtained text description is used as the first text. In some embodiments, the text descriptions output by multiple audio understanding models for the first audio can be obtained, and the text descriptions output by the multiple audio understanding models are aggregated and fused by a large language model to obtain the first text. In some other embodiments, the first text can also be obtained through manual description. This embodiment does not make specific limitations on the acquisition method of the first text.
[0066] Among them, the second text is the text for describing the expected audio.
[0067] Among them, the expected audio is the audio that the first audio is expected to be modified into. Among them, the modification of the first audio may include but is not limited to at least one operation such as adding sound elements, adjusting the sound ratio, changing the sound characteristics, changing the overall atmosphere, and reducing sound elements. It can be understood that the expected audio is not an actually existing audio, and the expected audio is used to represent the actual audio editing needs of the user.
[0068] Specifically, the second text may include the text information of one or more conditions that the audio to be generated by this audio processing is required to meet. Among them, the specific text content of the second text can be determined according to the actual audio editing needs of the user, and the second text is used to indicate the content expressed by the audio that the user expects to be modified. In practical applications, the user can input the text for describing the audio that the first audio is expected to be modified into on the audio editing interaction interface according to the actual audio editing needs of the individual, so that the electronic device can obtain the second text.
[0069] For example, the first audio is a recording of forest birdsong. The user expects the first audio to be modified into an audio with more sound elements, such as adding the sound of a stream, the sound of water hitting stones, and the subtle sound of a squirrel scratching the tree trunk. Correspondingly, the second text can be: "This audio depicts the scene of a forest in the early morning, with clear birdcalls, the pecking sound of a woodpecker, as well as the gurgling sound of a stream, the clear sound of water hitting stones, and the subtle sound of a squirrel's claws scratching the tree trunk as it jumps among the branches. All these sounds are intertwined with a distinct rhythm. The overall atmosphere is serene and full of vitality, giving a feeling of being deep in a quiet forest."
[0070] In some other embodiments, continuing with the above example, the user expects the first audio to be modified into an audio with fewer sound elements, such as removing the pecking sound of the woodpecker. Correspondingly, the second text can be: "This audio depicts the scene of a forest in the early morning, with clear birdcalls and a distinct rhythm. The overall atmosphere is serene and full of vitality, giving a feeling of being deep in a quiet forest."
[0071] In some other embodiments, the user expects to increase the volume of the pecking sound of the woodpecker in the first audio. Correspondingly, the second text can be: "This audio depicts the scene of a forest in the early morning, with clear birdcalls and the pecking sound of a woodpecker. The calls of various birds are intertwined with a distinct rhythm, and among them, the pecking sound is more prominent. The overall atmosphere is serene and full of vitality, giving a feeling of being deep in a quiet forest."
[0072] Step S220: Process the first audio according to the differences between the first text and the second text to obtain the second audio.
[0073] Among them, the difference between the first text and the second text can refer to the text difference.
[0074] For example, the first text is: "This audio depicts the scene of a forest in the early morning, with clear birdcalls and the pecking sound of a woodpecker. The calls of various birds are intertwined with a distinct rhythm. The overall atmosphere is serene and full of vitality, giving a feeling of being deep in a quiet forest."; the second text is: "This audio depicts the scene of a forest in the early morning, with clear birdcalls and a distinct rhythm. The overall atmosphere is serene and full of vitality, giving a feeling of being deep in a quiet forest." Then the difference between the first text and the second text is that the second text lacks the text descriptions "and the pecking sound of a woodpecker. The calls of various birds are intertwined".
[0075] For another example, continuing from the previous example, if the second text is "This audio depicts the scene of a forest in the early morning, with the clear chirping of birds, the pecking sound of woodpeckers, as well as the gurgling sound of a stream, the crisp sound of water hitting stones, and the faint sound of a squirrel's claws scratching the tree trunk as it jumps between branches. All these sounds are intertwined with a distinct rhythm. The overall atmosphere is peaceful and full of vitality, giving a feeling of being deep in a quiet forest.", then the difference between the first text and the second text is that the second text adds the text descriptions of "as well as the gurgling sound of a stream, the crisp sound of water hitting stones, and the faint sound of a squirrel's claws scratching the tree trunk as it jumps between branches".
[0076] In a specific implementation, the electronic device can obtain the difference between the first text and the second text, and based on these differences, process the first audio to obtain the second audio.
[0077] Furthermore, the difference between texts can be described by feature vectors; in this way, the electronic device processes the first audio based on the difference between the text features corresponding to the first text and the text features corresponding to the second text to obtain the second audio.
[0078] Furthermore, the electronic device can map the first text, the second text, and the first audio to the same feature space, and process the audio features corresponding to the first audio based on the difference between the text features corresponding to the first text and the text features corresponding to the second text to obtain the second audio. Since the second text refers to the text describing the audio that the first audio is expected to be modified into, and the second audio is obtained by processing the first audio based on the difference between the first text and the second text, the actually generated second audio can better meet the user's audio editing requirements for the first audio.
[0079] In the above audio editing method, the first text and the second text are obtained; the first text is the text describing the first audio; the second text is the text describing the expected audio, and the expected audio is the audio that the first audio is expected to be modified into; based on the difference between the first text and the second text, the first audio is processed to obtain the second audio.
[0080] Thus, by obtaining the first text for describing the first audio and the second text for describing the audio that the first audio is expected to be modified into, the first audio can be processed according to the difference between the first text and the second text, so that the processed second audio better matches the second text. Since the second text is used to describe the audio that the first audio is expected to be modified into, the actually generated second audio can better meet the audio effect expected by the user. The user does not need to master relevant audio editing technical content. Just by using the text for describing the audio, the user can generate the audio that meets the expected audio effect, which meets the user's audio editing needs, realizes flexible editing of the audio, and improves the audio editing efficiency.
[0081] In one embodiment, processing the first audio according to the difference between the first text and the second text to obtain a second audio includes: inputting the first audio into a pre-trained audio encoder to obtain first audio features; adjusting the first audio features according to the difference between the first text and the second text to obtain second audio features; and generating a second audio according to the second audio features.
[0082] Wherein, the first audio features refer to the audio features obtained by encoding the first audio.
[0083] Wherein, the second audio features refer to the audio features obtained by adjusting the first audio features.
[0084] Wherein, the pre-trained audio encoder can be constructed based on a neural network (such as neural networks like ResNet (Residual Network) and Transformer).
[0085] In a specific implementation, when the electronic device processes the first audio according to the difference between the first text and the second text to obtain a second audio, the electronic device can input the first audio into the pre-trained audio encoder, and the pre-trained audio encoder encodes the first audio to obtain the audio features corresponding to the first audio as the first audio features.
[0086] Thus, the electronic device can adjust the first audio features according to the difference between the first text and the second text, and use the adjusted audio features as the second audio features. Specifically, the electronic device can map the first text, the second text, and the first audio to the same feature space, and adjust the first audio features according to the difference between the text features corresponding to the first text and the text features corresponding to the second text to obtain the second audio features. Then, a second audio can be generated based on the second audio features. Specifically, by inputting the second audio features into the pre-trained audio decoder and decoding the second audio features, a second audio is obtained.
[0087] In the technical solution of this embodiment, by inputting the first audio into a pre-trained audio encoder, the first audio feature is obtained; according to the difference between the first text and the second text, the first audio feature is adjusted to obtain the second audio feature; and according to the second audio feature, the second audio is generated. In this way, the first audio feature corresponding to the first audio can be accurately obtained through the pre-trained audio encoder, and the first audio feature is adjusted according to the difference between the first text describing the first audio and the second text, so that the adjusted second audio feature matches the second text better, ensuring the audio effect of the second audio generated based on the second audio feature.
[0088] In one embodiment, adjusting the first audio feature according to the difference between the first text and the second text to obtain the second audio feature includes: obtaining the first text feature corresponding to the first text and the second text feature corresponding to the second text through a pre-trained text encoder; both the first text feature and the second text feature are in the same feature space as the first audio feature; and adjusting the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature.
[0089] Wherein, the first text feature refers to the text feature obtained by encoding the first text.
[0090] Wherein, the second text feature refers to the text feature obtained by encoding the second text.
[0091] Wherein, the pre-trained text encoder can be constructed based on a neural network (for example, neural networks such as ResNet (Residual Network) and Transformer).
[0092] Wherein, the pre-trained text encoder and the pre-trained audio encoder can be optimized through the same loss function, so that the text feature encoded by the pre-trained text encoder and the audio feature encoded by the pre-trained audio encoder can be mapped to the same feature space.
[0093] In a specific implementation, when the electronic device adjusts the first audio feature according to the difference between the first text and the second text to obtain the second audio feature, the first text and the second text can be respectively input into the pre-trained text encoder to perform encoding processing on the first text and the second text, so as to obtain the first text feature corresponding to the first text and the second text feature corresponding to the second text, and both the first text feature and the second text feature are in the same feature space as the first audio feature. Thus, the electronic device can adjust the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature.
[0094] In the technical solution of this embodiment, through a pre-trained text encoder, a first text feature corresponding to a first text and a second text feature corresponding to a second text are obtained; both the first text feature and the second text feature are in the same feature space as a first audio feature; according to the difference between the first text feature and the second text feature, the first audio feature is adjusted to obtain a second audio feature. In this way, through the pre-trained text encoder, the first text feature corresponding to the first text and the second text feature corresponding to the second text can be accurately obtained, and both the first text feature and the second text feature are in the same feature space as the first audio feature. Therefore, the adjustment process of the first audio feature can be more flexibly and accurately guided by the difference between the first text feature and the second text feature, making the entire adjustment logic more interpretable and realizing flexible audio editing operations according to text differences.
[0095] In one embodiment, adjusting the first audio feature according to the difference between the first text feature and the second text feature to obtain a second audio feature includes: obtaining a difference between the first text feature and the second text feature; and adjusting the first audio feature according to the difference to obtain a second audio feature.
[0096] Wherein, the first text feature, the second text feature, and the first audio feature have the same feature dimension.
[0097] In a specific implementation, when the electronic device adjusts the first audio feature according to the difference between the first text feature and the second text feature to obtain a second audio feature, the first text feature, the second text feature, and the first audio feature have the same feature dimension. By obtaining the difference between the first text feature and the second text feature and then adjusting the first audio feature, a second audio feature that better matches the second text can be obtained.
[0098] Specifically, since the first text feature, the second text feature, and the first audio feature have the same feature dimension, the numbers in the same positions of the first text feature, the second text feature, and the first audio feature can be added or subtracted. For example, the electronic device can obtain the difference obtained by subtracting the second text feature from the first text feature, and then subtract this difference from the first audio feature to obtain the second audio feature. Or, the electronic device can obtain the difference obtained by subtracting the first text feature from the second text feature, and then add this difference to the first audio feature to obtain the second audio feature.
[0099] The technical solution of this embodiment is to obtain the difference between the first text feature and the second text feature; and adjust the first audio feature according to the difference to obtain the second audio feature. In this way, both the first text feature and the second text feature are in the same feature space as the first audio feature. By using the difference between the first text feature and the second text feature to adjust the first audio feature corresponding to the first text feature, the adjusted second audio feature can be made to match the second text better.
[0100] In one embodiment, the pre-trained audio encoder and the pre-trained text encoder are trained as follows: Obtain the first training sample data; the first training sample data includes the first sample audio and the corresponding sample text; the sample text is the text describing the first sample audio; Iteratively train the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text; When the trained text encoder and the trained audio encoder meet the preset first training end condition, the pre-trained text encoder and the pre-trained audio encoder are obtained.
[0101] Among them, the first training sample data refers to the training sample data used to train the text encoder to be trained and the audio encoder to be trained.
[0102] Among them, the electronic device can use a large amount of audio data stored locally or a large amount of audio data downloaded from the network as the sample audio.
[0103] Among them, the first training sample data includes the first sample audio and the corresponding sample text, and the first sample audio can refer to the sample audio used to train the text encoder and the audio encoder.
[0104] Among them, the sample text corresponding to the first sample audio is the text describing the first sample audio. The acquisition method of the sample text corresponding to the first sample audio is the same as that of the first text, and will not be elaborated here.
[0105] Among them, the first training end condition can refer to the training end conditions of the text encoder to be trained and the audio encoder to be trained.
[0106] Among them, a text encoder and an audio encoder can be randomly initialized as the text encoder to be trained and the audio encoder to be trained.
[0107] Among them, the text encoder and the audio encoder can use common neural networks, such as ResNet, Transformer, etc.
[0108] In a specific implementation, the following training method is adopted for the pre-trained audio encoder and the pre-trained text encoder: The electronic device obtains first training sample data, which includes text-audio data pairs formed by multiple first sample audios and corresponding sample texts. In this way, the text encoder to be trained and the audio encoder to be trained can be iteratively trained based on the first sample audio and the corresponding sample text. When the trained text encoder and the trained audio encoder meet the preset first training end condition, the trained text encoder is used as the pre-trained text encoder, and the trained audio encoder is used as the pre-trained audio encoder.
[0109] In the technical solution of this embodiment, by obtaining first training sample data; the first training sample data includes a first sample audio and a corresponding sample text; the sample text is the text describing the first sample audio; based on the first sample audio and the corresponding sample text, the text encoder to be trained and the audio encoder to be trained are iteratively trained; when the trained text encoder and the trained audio encoder meet the preset first training end condition, the pre-trained text encoder and the pre-trained audio encoder are obtained. In this way, through the paired first sample audio and the sample text describing the first sample audio, the text encoder to be trained and the audio encoder to be trained are iteratively trained until the preset first training end condition is met, so that the trained text encoder and the trained audio encoder can extract text features and audio features more accurately and deeply.
[0110] In one embodiment, iteratively training the text encoder to be trained and the audio encoder to be trained based on the first sample audio and the corresponding sample text includes: inputting the first sample audio into the audio encoder to be trained to obtain first sample audio features; inputting the sample text corresponding to the first sample audio into the text encoder to be trained to obtain sample text features; obtaining a preset loss function, and determining the loss value of the loss function according to the difference between the first sample audio features and the sample text features; adjusting the model parameters of the text encoder to be trained and the model parameters of the audio encoder to be trained according to the loss value and continuing the iterative training until the trained text encoder and the trained audio encoder meet the first training end condition and stop training.
[0111] Among them, the first sample audio features refer to the audio features obtained by encoding the first sample audio by the audio encoder to be trained.
[0112] Among them, the sample text features refer to the text features obtained by encoding the sample text corresponding to the first sample audio by the text encoder to be trained.
[0113] Among them, the first training end condition may include that the loss value is lower than a preset threshold, or the number of iterations reaches a preset number.
[0114] Among them, the preset loss function can use a distance loss function, such as: Euclidean distance loss function, cosine similarity loss function, etc.
[0115] In specific implementation, as Figure 3 shown, a schematic diagram of iterative training of a to-be-trained audio encoder and a to-be-trained text encoder is provided. During the process of the electronic device iteratively training the to-be-trained text encoder and the to-be-trained audio encoder according to the first sample audio and the corresponding sample text, the first sample audio can be input into the to-be-trained audio encoder, and the to-be-trained audio encoder encodes the first sample audio to obtain the first sample audio feature; the sample text corresponding to the first sample audio is input into the to-be-trained text encoder, and the to-be-trained text encoder encodes the sample text to obtain the sample text feature. The electronic device can combine the preset loss function to measure the difference between the first sample audio feature and the sample text feature, determine the loss value of the loss function, adjust the model parameters of the to-be-trained text encoder and the model parameters of the to-be-trained audio encoder according to the loss value and continue the iterative training, so that the sample text feature and the first sample audio feature become more and more similar, and stop training until the trained text encoder and the trained audio encoder meet the first training end condition.
[0116] Furthermore, in the case where the first training end condition is that the loss value is lower than the preset threshold, when the electronic device determines the loss value of the loss function according to the difference between the first sample audio feature and the sample text feature, the electronic device can determine whether the loss value is lower than the preset threshold; when the loss value is lower than the preset threshold, it means that the model parameters of the text encoder and the model parameters of the audio encoder converge at this time, and the electronic device uses the text encoder and the audio encoder at this time as the pre-trained text encoder and the pre-trained audio encoder.
[0117] When the loss value is greater than or equal to a preset threshold, the electronic device can determine the model parameter update gradient of the text encoder and the model parameter update gradient of the audio encoder according to the model loss value, update the model parameters of the text encoder in the reverse direction based on the model parameter update gradient of the text encoder, update the model parameters of the audio encoder in the reverse direction based on the model parameter update gradient of the audio encoder, use the updated text encoder as the text encoder to be trained, use the updated audio encoder as the audio encoder to be trained, and repeat the steps of inputting the first sample audio into the audio encoder to be trained to obtain the first sample audio feature; inputting the sample text corresponding to the first sample audio into the text encoder to be trained to obtain the sample text feature, so as to continuously update the model parameters of the text encoder and the model parameters of the audio encoder until the loss value of the loss function is less than the preset threshold.
[0118] In the technical solution of this embodiment, by inputting the first sample audio into the audio encoder to be trained, the first sample audio feature is obtained; by inputting the sample text corresponding to the first sample audio into the text encoder to be trained, the sample text feature is obtained; a preset loss function is obtained, and the loss value of the loss function is determined according to the difference between the first sample audio feature and the sample text feature; the model parameters of the text encoder to be trained and the model parameters of the audio encoder to be trained are adjusted according to the loss value and iterative training is continued until the trained text encoder and the trained audio encoder meet the first training end condition and then the training is stopped.
[0119] In this way, by using the same preset loss function to measure the difference between the sample text feature output by the text encoder to be trained and the first sample audio feature output by the audio encoder to be trained, the model parameters of the text encoder to be trained and the model parameters of the audio encoder to be trained are continuously optimized, so that the text feature output by the trained text encoder for the sample text is more and more similar to the audio feature output by the trained audio encoder for the first sample audio, thereby improving the matching degree between the text feature and the audio feature respectively output by the text encoder and the audio encoder for paired text audio data.
[0120] In one embodiment, generating a second audio according to a second audio feature includes: inputting the second audio feature into a pre-trained audio decoder to obtain the second audio; wherein, the pre-trained audio decoder adopts the following training method: obtaining second training sample data; the second training sample data includes a second sample audio and a corresponding second sample audio feature; the second sample audio feature is obtained by encoding the second sample audio with a pre-trained audio encoder; inputting the second sample audio feature corresponding to the second sample audio into the audio decoder to be trained to obtain a new sample audio; adjusting the model parameters of the audio decoder to be trained according to the difference between the second sample audio and the new sample audio and continuing iterative training until the trained audio decoder meets the second training end condition and then stopping training to obtain the pre-trained audio decoder.
[0121] Wherein, the second training sample data refers to the training sample data for training the audio decoder to be trained.
[0122] Wherein, the electronic device can use a large amount of audio data stored locally or downloaded from the network as the sample audio.
[0123] Wherein, the second training sample data includes a second sample audio and a corresponding second sample audio feature.
[0124] Wherein, the second sample audio may refer to the sample audio for training the audio decoder to be trained.
[0125] Wherein, the second sample audio feature corresponding to the second sample audio can be obtained by encoding the second sample audio with the pre-trained audio encoder obtained in the previous step.
[0126] In a specific implementation, when the electronic device generates the second audio according to the second audio feature, it can input the second audio feature into the pre-trained audio decoder, and the pre-trained audio decoder decodes the second audio feature to obtain the second audio.
[0127] Wherein, such as Figure 4As shown in the figure, a schematic diagram of the training process of an audio decoder to be trained is provided. The pre-trained audio decoder adopts the following training method: Input the second sample audio (real audio) into the pre-trained audio encoder. The pre-trained audio encoder performs encoding processing on the second sample audio to obtain the second sample audio features corresponding to the second sample audio. Randomly initialize an audio decoder (the audio decoder can use common neural networks such as ResNet, Transformer, etc.) as the audio decoder to be trained. Input the second sample audio features corresponding to the second sample audio into the audio decoder to be trained. The audio decoder to be trained performs decoding processing on the second sample audio features to generate a new sample audio. According to the difference between the second sample audio and the new sample audio, adjust the model parameters of the audio decoder to be trained and continue iterative training until the trained audio decoder meets the second training end condition and then stop training to obtain the pre-trained audio decoder.
[0128] Among them, in the process of adjusting the model parameters of the audio decoder to be trained according to the difference between the second sample audio and the new sample audio and continuing iterative training, a preset loss function can be combined (a distance loss function can be used, such as: Manhattan distance loss function, Euclidean distance loss function, etc.). According to the difference between the second sample audio and the new sample audio, calculate the loss value. According to the loss value, adjust the model parameters of the audio decoder to be trained and continue iterative training until the trained audio decoder meets the second training end condition and then stop training.
[0129] Among them, the second training end condition can include that the loss value is lower than the preset threshold, or the number of iterations reaches the preset number.
[0130] Furthermore, in the case where the second training end condition is that the loss value is lower than the preset threshold, when the electronic device determines the loss value according to the difference between the second sample audio and the new sample audio, the electronic device can judge whether the loss value is lower than the preset threshold; when the loss value is lower than the preset threshold, it indicates that the model parameters of the audio decoder at this time have converged, and the electronic device takes the audio decoder at this time as the pre-trained audio decoder.
[0131] When the loss value is greater than or equal to a preset threshold, the electronic device may determine the model parameter update gradient of the audio decoder according to the model loss value, update the model parameters of the audio decoder in reverse based on the model parameter update gradient of the audio decoder, use the updated audio decoder as the audio decoder to be trained, and repeat the step of inputting the second sample audio feature corresponding to the second sample audio into the audio decoder to be trained, so as to continuously update the model parameters of the audio decoder, making the new sample audio obtained after decoding more and more similar to the second sample audio until the loss value of the loss function is less than the preset threshold.
[0132] In the technical solution of this embodiment, the second audio feature is input into a pre-trained audio decoder to obtain a second audio; wherein, the pre-trained audio decoder adopts the following training method: obtaining second training sample data; the second training sample data includes a second sample audio and the corresponding second sample audio feature; the second sample audio feature is obtained by encoding the second sample audio with a pre-trained audio encoder; inputting the second sample audio feature corresponding to the second sample audio into the audio decoder to be trained to obtain a new sample audio; adjusting the model parameters of the audio decoder to be trained according to the difference between the second sample audio and the new sample audio and continuing iterative training until the trained audio decoder meets the second training end condition and then stopping training to obtain the pre-trained audio decoder.
[0133] In this way, during the training process of the pre-trained audio decoder, by encoding the second sample audio with the pre-trained audio encoder, the audio feature corresponding to the second sample audio, that is, the second sample audio feature, can be obtained more accurately. Then, by inputting the second sample audio feature into the audio decoder to be trained to obtain a new sample audio, and based on the difference between the second sample audio and the new sample audio, adjusting the model parameters of the audio decoder to be trained and continuing iterative training, making the second sample audio and the new sample audio more and more similar, that is, enabling the original second sample audio to be decoded as accurately as possible after encoding, and further enabling the pre-trained audio decoder to decode the audio feature more accurately.
[0134] In another embodiment, as Figure 5 shown, a flowchart of an audio editing method is provided. Taking the application of this method to an electronic device as an example, the method specifically includes:
[0135] (1) Obtain a pre-trained text encoder, a pre-trained audio encoder, and a pre-trained audio decoder; obtain a first audio (original audio), a first text (original audio text description), and a second text (desired modified text description);
[0136] (2) Through the pre-trained text encoder, obtain the first text feature (original audio text feature) corresponding to the first text and the second text feature (expected modified text feature) corresponding to the second text. Input the first audio into the pre-trained audio encoder to obtain the first audio feature (original audio feature);
[0137] (3) Obtain the difference between the second text feature and the first text feature, and then add this difference to the first audio feature to obtain the second audio feature (modified audio feature);
[0138] (4) Input the second audio feature into the pre-trained audio decoder to obtain the second audio (modified audio).
[0139] It should be noted that the specific limitations of the above steps can refer to the specific limitations of an audio editing method described above.
[0140] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0141] Based on the same inventive concept, the embodiments of the present application also provide an audio editing device for implementing the above-mentioned audio editing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following audio editing device can refer to the limitations of the audio editing method in the above text, and will not be repeated here.
[0142] In an exemplary embodiment, as Figure 6 shown, an audio editing device is provided, including: an acquisition module 610 and a processing module 620, where:
[0143] The acquisition module 610 is used to acquire the first text and the second text; the first text is the text describing the first audio; the second text is the text describing the expected audio, and the expected audio is the audio that the first audio is expected to be modified into.
[0144] A processing module 620, configured to process the first audio according to the difference between the first text and the second text to obtain a second audio.
[0145] In one embodiment, the processing module 620 is specifically configured to input the first audio into a pre-trained audio encoder to obtain first audio features; adjust the first audio features according to the difference between the first text and the second text to obtain second audio features; and generate the second audio according to the second audio features.
[0146] In one embodiment, the processing module 620 is specifically configured to obtain a first text feature corresponding to the first text and a second text feature corresponding to the second text through a pre-trained text encoder; both the first text feature and the second text feature are in the same feature space as the first audio features; and adjust the first audio features according to the difference between the first text feature and the second text feature to obtain the second audio features.
[0147] In one embodiment, the processing module 620 is specifically configured to obtain the difference between the first text feature and the second text feature; and adjust the first audio features according to the difference to obtain the second audio features.
[0148] In one embodiment, the apparatus further includes: a training module, configured to obtain first training sample data; the first training sample data includes a first sample audio and a corresponding sample text; the sample text is a text describing the first sample audio; perform iterative training on the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text; and obtain the pre-trained text encoder and the pre-trained audio encoder when the trained text encoder and the trained audio encoder meet a preset first training end condition.
[0149] In one embodiment, the training module is specifically configured to input the first sample audio into the audio encoder to be trained to obtain first sample audio features; input the sample text corresponding to the first sample audio into the text encoder to be trained to obtain sample text features; obtain a preset loss function, and determine a loss value of the loss function according to the difference between the first sample audio features and the sample text features; adjust model parameters of the text encoder to be trained and model parameters of the audio encoder to be trained according to the loss value and continue iterative training until the trained text encoder and the trained audio encoder meet the first training end condition and then stop training.
[0150] In one embodiment, the processing module 620 is specifically configured to input the second audio feature into a pre-trained audio decoder to obtain the second audio; the training module is further configured to obtain second training sample data; the second training sample data includes a second sample audio and a corresponding second sample audio feature; the second sample audio feature is obtained by encoding the second sample audio using the pre-trained audio encoder; input the second sample audio feature corresponding to the second sample audio into the audio decoder to be trained to obtain a new sample audio; adjust the model parameters of the audio decoder to be trained according to the difference between the second sample audio and the new sample audio and continue iterative training until the trained audio decoder meets the second training end condition, and then stop training to obtain the pre-trained audio decoder.
[0151] Each module in the above audio editing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0152] In an exemplary embodiment, an electronic device is provided. The electronic device can be a server, and its internal structure diagram can be as Figure 7 shown. The electronic device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store the model parameters of the pre-trained audio encoder, the model parameters of the pre-trained text encoder, and the model parameters of the pre-trained audio decoder. The input / output interface of the electronic device is used for the processor to exchange information with external devices. The communication interface of the electronic device is used for communicating with external terminals through a network connection. The computer program, when executed by the processor, implements an audio editing method.
[0153] Those skilled in the art can understand, Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0154] In one embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0155] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0156] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., without limitation.
[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0160] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An audio editing method, characterized in that, The method includes: Obtaining a first text and a second text; the first text is a text describing a first audio; the second text is a text describing a desired audio, where the desired audio is the audio that the first audio is expected to be modified into; Processing the first audio according to the difference between the first text and the second text to obtain a second audio.
2. The method according to claim 1, wherein The processing the first audio according to the difference between the first text and the second text to obtain a second audio includes: Inputting the first audio into a pre-trained audio encoder to obtain a first audio feature; Adjusting the first audio feature according to the difference between the first text and the second text to obtain a second audio feature; Generating the second audio according to the second audio feature.
3. The method according to claim 2, wherein The adjusting the first audio feature according to the difference between the first text and the second text to obtain a second audio feature includes: Obtaining a first text feature corresponding to the first text and a second text feature corresponding to the second text through a pre-trained text encoder; both the first text feature and the second text feature are in the same feature space as the first audio feature; Adjusting the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature.
4. The method according to claim 3, wherein The adjusting the first audio feature according to the difference between the first text feature and the second text feature to obtain the second audio feature includes: Obtaining the difference between the first text feature and the second text feature; Adjusting the first audio feature according to the difference to obtain the second audio feature.
5. The method according to claim 3, wherein The pre-trained audio encoder and the pre-trained text encoder adopt the following training method: Obtaining first training sample data; the first training sample data includes a first sample audio and a corresponding sample text; the sample text is a text describing the first sample audio; Iteratively training the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text; When the trained text encoder and the trained audio encoder meet a preset first training end condition, obtaining the pre-trained text encoder and the pre-trained audio encoder.
6. The method according to claim 5, characterized in that The iteratively training the text encoder to be trained and the audio encoder to be trained according to the first sample audio and the corresponding sample text includes: Inputting the first sample audio into the audio encoder to be trained to obtain a first sample audio feature; Inputting the sample text corresponding to the first sample audio into the text encoder to be trained to obtain a sample text feature; Obtaining a preset loss function, and determining the loss value of the loss function according to the difference between the first sample audio feature and the sample text feature; Adjust the model parameters of the text encoder to be trained and the model parameters of the audio encoder to be trained according to the loss value, and continue iterative training until the trained text encoder and the trained audio encoder meet the first training end condition, then stop training.
7. The method according to claim 2, wherein The generating the second audio according to the second audio feature includes: Input the second audio feature into a pre-trained audio decoder to obtain the second audio; Among them, the pre-trained audio decoder adopts the following training method: Obtain second training sample data; the second training sample data includes a second sample audio and a corresponding second sample audio feature; the second sample audio feature is obtained by encoding the second sample audio using the pre-trained audio encoder; Input the second sample audio feature corresponding to the second sample audio into the audio decoder to be trained to obtain a new sample audio; According to the difference between the second sample audio and the new sample audio, adjust the model parameters of the audio decoder to be trained and continue iterative training until the trained audio decoder meets the second training end condition, then stop training to obtain the pre-trained audio decoder.
8. An audio editing device, characterized in that, The device includes: An acquisition module, configured to acquire a first text and a second text; the first text is a text describing a first audio; the second text is a text describing an expected audio, and the expected audio is an audio that the first audio is expected to be modified into; A processing module, configured to process the first audio according to the difference between the first text and the second text to obtain a second audio.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.