Audio editing method, device and electronic equipment

By segmenting audio into audio granular data and merging them based on voiceprint similarity, the problems of speech fluency and low editing efficiency after long audio clipping are solved, achieving efficient audio editing.

CN119785765BActive Publication Date: 2026-05-29CHINA MOBILE INTERNET CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE INTERNET CO LTD
Filing Date
2025-01-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In scenarios such as meetings, the quality of long audio recordings is poor, and when cut into multiple short audio segments, it is easy to cut off the words of the same person, resulting in low editing efficiency and the need for frequent proofreading.

Method used

The audio to be edited is divided into multiple audio particle data. The similarity of the voiceprints of adjacent audio particle data is obtained through voiceprint recognition. The similarity is then merged to generate the target audio segment, which is then assigned to the editing object for editing.

Benefits of technology

It avoids the unreasonable segmentation of the same speaker's speech, ensures the fluency of speech, reduces the proofreading time of the editing object for audio segments, and improves editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785765B_ABST
    Figure CN119785765B_ABST
Patent Text Reader

Abstract

The application discloses an audio editing method and device and electronic equipment, and belongs to the technical field of multi-person audio editing. The method comprises the following steps: cutting a target audio to be edited into a plurality of audio particle data, and acquiring voiceprints of the plurality of audio particle data; determining voiceprint similarity between the voiceprints of two adjacent audio particle data in the plurality of audio particle data; performing merging processing on the plurality of audio particle data according to the voiceprint similarity, to obtain a target audio segment; and performing editing on the target audio segment by assigning the target audio segment to an editing object, to obtain an edited audio corresponding to the target audio. In this way, the speech of the same speaker can be avoided from being cut into unreasonable audio segments, and the fluency of the speech is ensured, so that the editing object does not need to spend too much time on checking the fluency of the audio segment, and the efficiency of audio editing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multi-person audio editing technology, and in particular to an audio editing method, apparatus and electronic device. Background Technology

[0002] In scenarios such as meetings, users may record long periods of target audio. However, the quality of the recorded target audio is often lacking and requires further editing in order to create a high-quality file for other users.

[0003] To improve editing efficiency, long audio files can be cut into multiple short audio segments, each segment distributed to different users for editing, and then finally combined. However, current techniques for cutting long audio files into multiple short audio segments rely heavily on user experience, which can easily lead to the misinterpretation of continuous speech by the same person, potentially splitting a complete sentence into two short audio segments. This necessitates frequent re-proofing of the short audio segments during editing, resulting in low overall audio editing efficiency. Summary of the Invention

[0004] This application provides an audio editing method, apparatus, and electronic device that can at least solve the problems of lack of fluency and low editing efficiency after long audio is cut.

[0005] To solve the above-mentioned technical problems, this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide an audio editing method, which includes: dividing a target audio to be edited into multiple audio particle data, and obtaining the voiceprints of the multiple audio particle data; determining the voiceprint similarity between the voiceprints of two adjacent audio particle data in the multiple audio particle data; merging the multiple audio particle data according to the voiceprint similarity to obtain a target audio segment; and obtaining the edited audio corresponding to the target audio by assigning the target audio segment to an editing object for editing.

[0007] Secondly, embodiments of this application provide an audio editing apparatus, which includes: a first acquisition module, configured to segment a target audio to be edited into multiple audio particle data and acquire the voiceprints of the multiple audio particle data; a determination module, configured to determine the voiceprint similarity between the voiceprints of two adjacent audio particle data in the multiple audio particle data; a merging module, configured to merge the multiple audio particle data according to the voiceprint similarity to obtain a target audio segment; and a second acquisition module, configured to obtain the edited audio corresponding to the target audio by assigning the target audio segment to an editing object for editing.

[0008] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and when the program or instructions are executed by the processor, they implement the steps of the method described in the first aspect above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect above.

[0010] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of the method described in the first aspect above.

[0011] The technical solution provided in this application may include the following beneficial effects:

[0012] In this embodiment, the target audio to be edited is segmented into multiple audio particle data, and the voiceprints of the multiple audio particle data are obtained. Then, the voiceprint similarity between the voiceprints of two adjacent audio particle data is determined. Subsequently, the multiple audio particle data are merged according to the voiceprint similarity to obtain the target audio segment. Finally, the target audio segment is assigned to the editing object for editing to obtain the edited audio corresponding to the target audio. Through the above method, it is possible to avoid cutting the speech of the same speaker into unreasonable audio segments, ensuring the fluency of speech. Thus, the editing object does not need to spend too much time proofreading the fluency of the audio segment, improving the efficiency of audio editing.

[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0015] Figure 1 A flowchart illustrating an audio editing method provided in an embodiment of this application is shown;

[0016] Figure 2 A schematic diagram of the structure of the acoustic network provided in an embodiment of this application is shown;

[0017] Figure 3 A schematic diagram of the structure of the speech synthesis model provided in an embodiment of this application is shown;

[0018] Figure 4 This illustration shows a structural schematic diagram of an audio editing device provided in an embodiment of this application;

[0019] Figure 5 This illustration shows a structural schematic diagram of an electronic device provided in an embodiment of this application;

[0020] Figure 6 A schematic diagram of the structure of another electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0022] In scenarios such as meetings, users may record long audio clips. However, the quality of these long audio clips is often lacking, requiring subsequent editing such as cutting, deleting, noise reduction, volume adjustment, and text proofreading to create high-quality files for other users.

[0023] To improve editing efficiency, the current practice is to cut long audio files into multiple short audio segments of equal length, distribute each segment to different users for editing, and then compile them together at the end.

[0024] In related technologies, cutting long audio files into multiple short audio segments mainly relies on user experience and is a blind operation. This can easily lead to the misinterpretation of continuous speech by the same person, resulting in a single complete sentence being split into two short audio segments. This necessitates frequent re-proofreading of the short audio segments, increasing workload. Furthermore, the length of short audio segments is still relatively long compared to a single sentence. Editing short audio requires users to repeatedly listen to the audio and frequently drag the progress bar within a still relatively long short audio segment, which is cumbersome, time-consuming, and inefficient. In addition, cutting long audio files into multiple short audio segments and editing each segment are all offline operations, making the management of long audio files difficult.

[0025] Figure 1 This illustration shows a flowchart of an audio editing method provided in an exemplary embodiment of this application. This method can be executed by an electronic device. The electronic device can be a terminal such as a mobile phone or computer, or it can be a server. Figure 1As shown, the method mainly includes the following steps:

[0026] S101: Divide the target audio to be edited into multiple audio particle data, and obtain the voiceprints of multiple audio particle data.

[0027] In this embodiment, the target audio to be edited can be segmented into multiple audio granular data, and then the voiceprints of these multiple audio granular data can be obtained. In practical applications, there are usually multiple speakers who take turns speaking, and the recorded target audio includes the content spoken by each speaker in turn. After recording, the user can upload the target audio to a cloud drive for storage. At this time, the cloud drive server can segment the target audio into multiple audio granular data with smaller granularity. Since in practical applications, the length of each speaker's speech varies, and the duration of their speech is highly random, it is unknown which speaker spoke how much before segmentation. Therefore, in this embodiment, to cover the speaker's speech content as much as possible, an initial small scale can be set. This scale can be smaller than the length of a single word. This scale can be represented as n frames of speech signal in the target audio, where n is a variable. In practical applications, n frames of speech signal can be sequentially segmented from the target audio along the time axis as audio granular data, with any remaining speech signal less than n frames treated as one audio granular data. In this embodiment, the target audio can be segmented into multiple smaller audio particle data according to this scale, which facilitates the coverage of the speech content as much as possible. Each speaker has the spectral characteristics of the sound waves emitted, and a unique voiceprint can be generated for each speaker, thereby using the voiceprint to identify the speaker's identity. In this embodiment, the voiceprint of the speaker corresponding to the audio particle data can be obtained, which facilitates the merging of audio particle data through voiceprint.

[0028] In one optional implementation, acquiring the voiceprint of the plurality of audio particle data includes:

[0029] By inputting multiple audio particle data into a voiceprint network for voiceprint recognition, the voiceprints of the multiple audio particle data are obtained.

[0030] In this embodiment, audio granular data can be input into a voiceprint network for voiceprint recognition to obtain voiceprints from multiple audio granular data. The voiceprint network in this embodiment involves a Spearkr Verification (SV) task, which verifies whether an input speech segment belongs to a specific speaker. This involves two concepts: enrollment utterance and verification utterance. The former can be understood as a reserved "voiceprint," while the latter is the speech used for verification. This embodiment primarily uses enrollment utterance and does not require verification utterance. The voiceprint network in this embodiment can use a simple two-layer neural network model for voiceprint recognition, sacrificing some accuracy to improve processing speed. Inputting audio granular data into the voiceprint network for voiceprint recognition can improve the speed of obtaining voiceprints from audio granular data and shorten the time required to obtain voiceprints. In practical applications, other methods can also be used for voiceprint recognition (for example, a voiceprint recognition model consisting of a Long Short-Term Memory (LSTM) network model and a generalized end-to-end loss function (GE2E) can be used for voiceprint recognition). This application does not impose specific limitations on the embodiments.

[0031] In the embodiments of this application, the voiceprint network includes a two-layer structure. The receptive field can be increased by alternately using convolution and a Long Short-Term Memory (LSTM) network with self-attention mechanism. The weight of local audio signals can also be enhanced to achieve a balance between the two-layer structure, improve the robustness of the system, and have a significant inhibitory effect on the generation of bad cases, thereby improving the computation speed while ensuring the voiceprint quality.

[0032] In an optional implementation, the step of obtaining the voiceprint of the multiple audio particles by inputting them into a voiceprint network for voiceprint recognition may include the following steps:

[0033] Step 1: By performing a Fourier transform on the multiple audio particle data, the Mel spectrograms corresponding to the multiple audio particle data are obtained. In this embodiment, a Fourier transform can be performed on the multiple audio particle data to obtain the Mel spectrograms (Mel spectrograms are characteristic data of sound, a spectrogram at the mel scale) corresponding to the multiple audio particle data. In practical applications, the Fast Fourier Transform (FFT) algorithm can be used to extract the Mel spectrogram from each audio signal in the audio particle data, or other methods can be used to obtain the Mel spectrograms corresponding to the audio particle data. This embodiment does not specifically limit the methods.

[0034] Step 2: By inputting the Mel spectrum into the first layer of the speaker network, the first spectral feature is obtained. The first layer includes PreNet, Convolutional Network (Conv), Long Short-Term Memory (LSTM) network with Self-Attention, and LSTM. In this embodiment, the Mel spectrum can be input into the first layer of the speaker network and processed sequentially through PreNet, Conv, LSTM with Self-Attention, and LSTM to obtain the first spectral feature.

[0035] Step 3: By inputting the first spectral feature into the second layer structure of the voiceprint network, the second spectral feature is obtained, wherein the second layer structure is the same as the first layer structure. In this embodiment, the first spectral feature is input into the second layer structure of the voiceprint network and processed sequentially through PreNet, Conv, LSTM with Self Attention, and LSTM to obtain the second spectral feature.

[0036] Step 4: By inputting the second spectral feature into the fully connected layer of the voiceprint network, voiceprints of multiple audio particle data are obtained. In this embodiment, inputting the second spectral feature into the fully connected layer of the voiceprint network can be mapped to the speaker's voiceprint, thereby obtaining voiceprints of multiple audio particle data.

[0037] In the embodiments of this application, such as Figure 2As shown, the PreNet in the voiceprint network consists of two layers, each including a fully connected (FC) layer, a rectified linear unit (ReLU), and dropout structures. The first PreNet performs a non-linear transformation on the Mel Spectrogram, and the second PreNet performs a non-linear transformation on the first spectral feature output by the (LSTM). Conv, taking the input two-dimensional vector sequence (i.e., the Mel Spectrogram mentioned above), performs multiple 1D convolution operations to effectively model the current two-dimensional vector sequence and contextual information, outputting low-dimensional spectral features. The first LSTM extracts temporal spectral features using a self-attention mechanism. The self-attention mechanism means that the LSTM's output and input are processed by a self-attention module, maintaining a monotonic order between features and attention, without skipping input sequences, thus enhancing the robustness of feature extraction. The self-attention mechanism and the LSTM output are concatenated to form new features. The second LSTM extracts the temporal contextual features of the sequence normally, obtaining the spectral features.

[0038] In one alternative implementation, the voiceprint network can be determined in the following manner:

[0039] Construct a speech synthesis model, wherein the speech synthesis model includes an acoustic model and a vocoder, and the acoustic model and the vocoder are cascaded together.

[0040] The trained speech synthesis model is obtained by training the speech synthesis model.

[0041] The voiceprint network is determined based on the acoustic model in the trained speech synthesis model.

[0042] In practical applications, voiceprints are essentially the acoustic features of a speaker, and voiceprint networks are feature extraction networks. Because it's difficult to label features with tags, voiceprint networks are challenging to train in a supervised manner. Therefore, unsupervised training can be used. Unsupervised training means that the input itself does not carry tags; the input is used as a sample to fit the loss function. Under unsupervised training, the voiceprint is an intermediate output. In this embodiment, to obtain the voiceprint network, a speech synthesis model can be constructed. This model includes an acoustic model with the same structure as the voiceprint network and a vocoder. The acoustic model and vocoder are cascaded, and the overall structure is as follows: Figure 3As shown. Then, the speech synthesis model is trained to obtain a trained speech synthesis model; finally, the voiceprint network is determined based on the acoustic model in the trained speech synthesis model. In practical applications, the acoustic model and vocoder can form a classic architecture in Text-to-Speech (TTS). The vocoder can include HiFi-GAN (a Generative Adversarial Networks (GAN) model for efficient, high-quality speech synthesis), WaveNet (a text-to-speech system), etc. In this embodiment, the specific structure of the vocoder is not specifically limited. In this embodiment, the speech synthesis model is not constructed for actual speech synthesis, but rather to train an acoustic model with the same structure as the voiceprint network. Finally, upon completion of training, the voiceprint network can inherit the parameters of the acoustic model, thereby achieving unsupervised training of the acoustic network.

[0043] In one optional implementation, the step of training the speech synthesis model to obtain the trained speech synthesis model may include the following steps:

[0044] Step 1: Divide multiple sample audios into sample audio particle data; In this embodiment of the application, some audio data can be selected as sample audios, and the sample audios can be divided into multiple sample audio particle data according to different scales.

[0045] Step 2: Perform Fourier transform on the sample audio particle data to obtain the first Mel spectrum; In this embodiment of the application, Fourier transform can be performed on the sample audio particle data to obtain the first Mel spectrum. In practical applications, the FFT algorithm can be used to extract the first Mel spectrum from each audio signal of the sample audio particle data.

[0046] Step 3: By inputting the first Mel spectrum into the acoustic model, the voiceprint of the sample audio particle is obtained. In this embodiment, the acoustic model is responsible for understanding and compressing all relevant contextual information of the input Mel Spectrogram (first Mel spectrum), encoding the input first Mel spectrum into a new, dense vector representation. The encoded information can typically be highly compressed and capture the key features and semantic information of the first Mel spectrum, so that the vocoder can generate a new target sequence based on these decoded features. Inputting the first Mel spectrum into the acoustic model can obtain a representation of the voiceprint of the sample audio particle data.

[0047] Step 4: By inputting the voiceprint of the sample audio particles into the vocoder, the second Mel spectrum is obtained. In this embodiment, the vocoder can receive the hidden representation or feature vector from the acoustic model, and then remap the abstract representation obtained from the acoustic model back to the original data space, converting it into the form of the original data to obtain a new MelSpectrogram (second Mel spectrum).

[0048] Step 5: Update the training parameters of the acoustic model and the vocoder by calculating the loss value between the first Mel spectrum and the second Mel spectrum. In this embodiment, the loss value between the first Mel spectrum and the second Mel spectrum can be calculated, and the training parameters of the acoustic model and the vocoder can be updated based on the loss value.

[0049] Step 6: If the training iterations meet the preset number of training iterations or the loss value meets the preset conditions, obtain the trained speech synthesis model. In this embodiment, training is considered complete when the training iterations reach the preset number of training iterations or the loss value stabilizes over multiple consecutive rounds; otherwise, return to step 3 above until training is complete. Upon completion of training, the speaker network can inherit the parameters of the acoustic model, thereby achieving unsupervised training of the acoustic network and making the performance of the speaker network more inclined to resolve speaker-distinguishing features.

[0050] In this embodiment of the application, for step 5 above, the loss value between the first Mel spectrum and the second Mel spectrum can be calculated using loss functions such as cross-entropy.

[0051] For example, the formula for calculating the loss value is as follows:

[0052]

[0053] in, The number of frames is equal to the number of frames in the first Mel spectrum and the second Mel spectrum. For the first Mel spectrum, This is the second Mel spectrum.

[0054] In practical applications, direction propagation can be performed using optimization methods such as stochastic gradient descent based on the loss value to update the parameters of the acoustic model and the vocoder.

[0055] S102: Determine the voiceprint similarity between the voiceprints of two adjacent audio particle data in the plurality of audio particle data.

[0056] In this embodiment, the voiceprint similarity between two adjacent audio particle data points can be determined based on the voiceprint of the audio particle data. In practical applications, algorithms such as cosine angle can be used in the temporal direction to calculate the voiceprint similarity between two sorted adjacent audio particle data points. Determining the voiceprint similarity helps in merging audio particle data based on voiceprint similarity, which can improve the accuracy of subsequent audio particle data merging.

[0057] S103: Based on the voiceprint similarity, merge multiple audio particle data to obtain the target audio segment.

[0058] In this embodiment, multiple audio particle data can be merged based on voiceprint similarity to obtain a target audio segment. The resulting target audio segments are similar in voiceprint, indicating they were spoken by the same person. This not only avoids splitting a single person's speech into two segments but also prevents merging speech from different people, thus improving the accuracy of target audio segmentation.

[0059] In one optional implementation, merging multiple audio particle data based on the voiceprint similarity to obtain the target audio segment may include the following steps:

[0060] Step 1: Based on the voiceprint similarity, merge multiple audio particle data to obtain a first initial audio segment. In this embodiment, a threshold can be preset. To improve resistance to background noise, the threshold can be set lower, for example, 0.6. If the voiceprint similarity of two adjacent audio particle data is greater than or equal to the threshold, the two audio particle data can be concatenated together. This process is repeated until all audio particle data is traversed, ultimately yielding multiple first initial audio segments.

[0061] Step 2: Based on the length of the first initial audio segment and the voiceprint similarity, merge the first initial audio segments to obtain a second initial audio segment. In practical applications, since the voiceprint network and the threshold in Step 1 may have a certain false detection rate, post-processing can be performed on the first initial audio segments to smooth out the fluctuations between them and improve the accuracy of the final target audio segment. In this embodiment, the first initial audio segments can be merged based on their length and voiceprint similarity to obtain a second initial audio segment. This can compensate for potential errors and help obtain an accurate target audio segment.

[0062] Step 3: Organize the second initial audio segment to obtain the target audio segment. The organization includes merging and segmenting the second initial audio segment. In this embodiment, the second initial audio segment can be organized according to preset organization rules to obtain the target audio segment. This allows for more efficient editing by the user.

[0063] In this embodiment of the application, for step 2 above, in an optional implementation, merging the first initial audio segment according to the length of the first initial audio segment and the voiceprint similarity to obtain the second initial audio segment may include:

[0064] Get the length of the first initial audio segment;

[0065] If the length of the first initial audio segment is less than the first length, the initial audio segments are merged based on the voiceprint similarity between the last audio particle data and the first audio particle data of adjacent initial audio segments to obtain a second initial audio segment.

[0066] In this embodiment, the length of each first initial audio segment can be obtained. If the length of the first initial audio segment A is less than a first length (which can be the length of a single character, an adjustable empirical value), algorithms such as cosine similarity can be used to calculate the voiceprint similarity between the last audio particle data in the first initial audio segment B and the first audio particle data in the first initial audio segment C. If the voiceprint similarity is greater than or equal to a threshold, the first initial audio segments A, B, and C can be merged into a second initial audio segment. The first initial audio segment B is ordered before the first initial audio segment A, and the first initial audio segment C is ordered after the first initial audio segment A. In practical applications, this situation is often due to errors in the voiceprint network; the first initial audio segments A, B, and C belong to the same speaker's speech.

[0067] In an alternative implementation, after obtaining the second initial audio segment, the method further includes:

[0068] By performing speech recognition on the second initial audio segment, the text information corresponding to the second initial audio segment can be obtained;

[0069] Based on the text information corresponding to the second initial audio segment, identify whether the second initial audio segment is a complete sentence;

[0070] If the second initial audio segment is an incomplete sentence, the target audio is re-segmented to obtain a new second initial audio segment.

[0071] In this embodiment, speech recognition can be performed independently on the second initial audio segment to obtain the corresponding text information. This text information, besides verifying the integrity of the second initial audio segment, can also be reused for subsequent editing. That is, the editing object can add speaker information (such as name, identity, etc.) to the text information and correct the text information itself, speaking time, etc., ultimately forming the meeting minutes. In practical applications, a speaker speaks sentence by sentence. Therefore, full syntactic analysis techniques in Natural Language Processing (NLP) (such as N-grams (a language model), Neural Nets, etc.) can be used to analyze whether each piece of text information is one or more complete sentences. If the second initial audio segment is an incomplete sentence (i.e., the text information of the second initial audio segment is not a complete sentence), the segmentation scale can be reduced, the target audio can be re-segmented, new audio particle data can be obtained, and a new second initial audio segment can be obtained. Optionally, the method for reducing the segmentation scale can be: n = nx * m frames, where x is the number of reductions and m is the step size. This ensures that the final target audio segment is a complete sentence. In this embodiment, if the text information of a second initial audio segment is empty, it indicates that the second initial audio segment is a silence point, and the second initial audio segment can be deleted to save storage space.

[0072] In one optional implementation, the step of processing the second initial audio segment to obtain the target audio segment may include the following steps:

[0073] Step 1-1: Determine the target quantity based on the number of editing objects. In this embodiment, the organization rule may include assuming the number of editing objects (users collaborating on editing) is P, then the number of target audio segments needs to reach 2P+Q or more. 2P ensures that target audio segments can be staggered and allocated to users subsequently, and Q ensures that there is a certain margin for each user to edit in turn.

[0074] Steps 1-2: If the number of second initial audio segments is less than the target number, starting from the longest second initial audio segment, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment; in this embodiment, if the number of second initial audio segments is less than the target number, starting from the longest second initial audio segment, some longer second initial audio segments are segmented according to the text information obtained from speech recognition, with a complete sentence as the node, until the number of second initial audio segments reaches 2P+Q, and the second initial audio segment can be determined as the target audio segment.

[0075] Step 2-1: Obtain the length of the second initial audio segment. In this embodiment, the sorting rules may further include that the length of the target audio segment is between [T1, T2], ensuring that the workload of a single edit is similar and providing a basis for fairness for collaborative editing by multiple people. During sorting, the length of the second initial audio segment can be obtained first.

[0076] Step 2-2: If the length of the second initial audio segment is less than the second length, merge the second initial audio segment with the adjacent second initial audio segment of shorter length to obtain the target audio segment. In this embodiment, if the length of the second initial audio segment is less than the second length (T1), compare the lengths of the previous and subsequent second initial audio segments, and splice the current second initial audio segment to the shorter second initial audio segment to obtain the target audio segment. This situation could be due to a slight pause during the same user's speech, or it could be due to users alternately saying a simple word, such as a meeting host saying "next".

[0077] Steps 2-3: If the length of the second initial audio segment is greater than the third length, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment, wherein the second length is less than the third length. In this embodiment, if the length of the second initial audio segment is greater than the third length (T2), the second audio segment can be segmented according to the text information of speech recognition, with a complete sentence as the node, until the length of each segmented second initial audio segment is between [T1, T2], and then the second initial audio segment is determined as the target audio segment.

[0078] S104: By assigning the target audio segment to the editing object for editing, the edited audio corresponding to the target audio is obtained.

[0079] In this embodiment, a target audio segment can be assigned to an editing object for editing, resulting in the edited audio corresponding to the target audio. In a specific implementation, the editing object, i.e., the user, can directly edit the target audio segment on a cloud drive. In practical applications, users can directly view the complete target audio on their respective cloud drives. The target audio can display the segmentation points of each target audio segment; silent audio segments are displayed as frozen and uneditable. However, users can choose to delete silent audio segments. This method effectively reduces the complexity of operations, saves time, and helps improve the efficiency of audio editing.

[0080] In an optional implementation, obtaining the edited audio corresponding to the target audio by assigning the target audio segment to an editing object for editing may include the following steps:

[0081] Step 1: Randomly assign the target audio segments to the editing objects for collaborative editing. In this embodiment, the target audio segments can be randomly assigned to each user for editing simultaneously, following the principle of "scattering." "Scattering" means that the target audio segments assigned to users are spaced at least one other target audio segment apart. This is because after editing the current target audio segment, users are more accustomed to editing adjacent target audio segments. This ensures the continuity of the content of the initially edited target audio segments. Initially scattering the target audio segments assigned to users aligns with user habits and improves initial editing efficiency. As users become more familiar with the content of the target audio data, skipping between target audio segments will not present significant obstacles.

[0082] Step 2: Lock the target audio segment. In this embodiment, an idle thread can be allocated to each user in the thread pool. This thread is responsible for locking, unlocking, and editing the target audio segment assigned to the user. The thread locks the target audio segment assigned to the user, allowing editing by the current user while preventing other users from editing it, thus avoiding interference with the target audio segment being edited.

[0083] Step 3: After the editing object completes the editing of the target audio segment, the locking of the target audio segment is removed and the review object is determined from the editing objects; In this embodiment of the application, when the current target audio segment is completed, the thread will unlock it and can select the qualified one from the editing objects as the review object, so that the review object can review the edited target audio segment.

[0084] Step 4: If the review object completes the review, the edited audio corresponding to the target audio is obtained. In this embodiment of the application, the final edited audio can be obtained after the review object completes the review.

[0085] In an optional implementation, after randomly assigning the multiple target audio segments to the multiple editing objects for collaborative editing, the method further includes:

[0086] In response to a request from the editing object, if the editing object does not have the target audio segment being edited and the target audio segment corresponding to the request is not locked, the target audio segment is assigned to the editing object.

[0087] In this embodiment, initially, a target audio segment can be assigned to the editing object. In subsequent processes, in response to the editing object's request, if the editing object does not have a target audio segment being edited and the requested target audio segment is not locked, the target audio segment can be assigned to the editing object. If the editing object itself has a target audio segment being edited and / or the requested target audio segment is locked, the editing object's request can be rejected, and the reason for rejection can be displayed. The aforementioned thread pool can also be responsible for locking, unlocking, and editing the target audio segment selected by the user. For the target audio segment selected by the user, the thread will lock it, allowing only the current user to edit while preventing other users from editing. When the current target audio segment is finished being edited, the thread will unlock it.

[0088] In an optional implementation, the method further includes:

[0089] During the editing process of the target audio segment, an operation time limit is added to the target audio segment;

[0090] If the editing object fails to complete the editing of the target audio segment within the operation time limit, an alarm signal is generated, wherein the alarm signal is used to indicate a reduction in the operation time limit;

[0091] If the operation time limit is lower than the preset time limit threshold, the editing of the target audio segment by the editing object is cancelled and the locking of the target audio segment is removed.

[0092] In this embodiment, to ensure editing efficiency, an operation time limit can be added to each target audio segment. If this time limit is exceeded, an alarm signal is generated, indicating that the operation time limit should be shortened. In practical applications, if the editing object fails to complete the editing operation before the time limit is reached, it is usually due to offline status or leaving the site. The operation time limit can be decayed in dynamically increasing steps, accelerating the decay process, repeating until the operation time limit reaches 0. At this point, the editing object's editing of the target audio segment is canceled, the target audio segment is unlocked, and other editing objects can edit it. This prevents editing objects from being offline or leaving the site from occupying the target audio segment, thereby improving the overall efficiency of audio data editing.

[0093] For example, after generating an alarm signal, the operation time limit can be calculated using the following formula:

[0094]

[0095] in, The time limit for the (t+1)th operation is... Let t be the time limit for the operation. It is a constant, and , The length of the target audio segment.

[0096] In one optional implementation, adding an operation time limit to the target audio segment can include the following two scenarios:

[0097] Scenario 1: If the frequency with which the editing object triggers editing of the target audio segment is less than a preset threshold, the operation time limit is determined based on the length of the target audio segment and a preset amplification factor. If the frequency with which the editing object triggers editing of the target audio segment is less than the preset threshold, the corresponding operation time limit can be calculated based on the length of the target audio segment and a preset amplification factor. This can be understood as amplifying based on the length of the target audio segment to ensure that the editing object has time to listen to the target audio segment (i.e., the editing time is at least equal to the length of the target audio segment), with the remaining time reserved for the editing object to perform the editing. For example, the formula can be:

[0098] , ;

[0099] in, The frequency at which editing of the target audio segment is triggered for the object being edited. For the preset threshold, For operation time limit, This is the magnification factor. The length of the target audio segment. In practical applications, the amplification factor can be... .

[0100] Scenario 2: When the frequency is greater than or equal to the preset threshold, the operation time limit is determined based on the length of the target audio segment, the amplification factor, the frequency, the preset threshold, and the time step. In this embodiment, if the frequency is greater than or equal to the preset threshold, the operation time limit can be calculated based on the length of the target audio segment, the amplification factor, the frequency, the preset threshold, and the time step. This is because some target audio segments are complex, potentially involving sound effects, unclear audio quality, or important speaker remarks, making editing tedious. When the frequency of the editing object triggering the editing operation on the target audio segment is high, the editing time can be increased proportionally to provide more ample editing time for the editing object and ensure completion of the editing. For example, the formula can be:

[0101] , ;

[0102] in, The frequency at which editing of the target audio segment is triggered for the object being edited. For the preset threshold, For operation time limit, This is the magnification factor. The length of the target audio segment, For the time step, in practical applications, the amplification factor can take the following values: .

[0103] In one alternative implementation, once all target audio segments have been edited, at least two editable objects can be selected as review objects from all editable objects, and review permissions can be granted to them to review the target audio segments edited by the other editable objects.

[0104] Optionally, the audit targets can be determined in the following ways:

[0105] Step 1: Calculate the average score of the most recent collaborative audio editing data of each editing object (user). Sort the users according to the average score from largest to smallest and select the top K (K>I) users. The average score, to a certain extent, represents the user's attitude and ability. Select users with better attitude and ability as the review objects.

[0106] Step 2: Sort the K users by the number of times they have revoked the editing permission for audio segments, from largest to smallest, and select the top M (M>I) users. Regardless of whether the user has network problems or is unable to edit, they are not suitable as review subjects and are removed.

[0107] Step 3: Count the number of audio segments edited by each of the M users. Sort the M users from largest to smallest number of audio segments edited and select the top I users as the review targets, granting them review permissions.

[0108] In practical applications, all target audio segments can be divided into J equal parts (J is a multiple of I), with each part containing multiple consecutive target audio segments, ensuring the continuity of the review process. Based on the principle of avoiding favoritism, if a segment of audio segments does not contain any target audio segments edited by the reviewer, that segment can be assigned to that reviewer. In extreme cases, if a segment contains audio segments edited by all reviewers, it should be split into multiple equal parts until one can be assigned to a reviewer. Reviewers can browse other users' edits to the audio segments and delete any errors or omissions, thus submitting the correct edit.

[0109] Once all audio segments have been reviewed by all reviewers, the final edited audio can be obtained, and subsequent business operations can be performed, such as submitting for approval or publishing to a specified range of users.

[0110] In an alternative implementation, after obtaining the edited audio, the method further includes:

[0111] The editing object is scored based on the operation time limit of the editing object, and the corresponding score result is obtained.

[0112] In this embodiment, the operation time limit is used not only to eliminate invalid cases occupying the target audio segment, but also to score the target audio segment edited by the editing object (user) and obtain the corresponding score result. The score result can be used as a reference for the workload of the editing object.

[0113] For example, the formula is as follows:

[0114]

[0115] in, The rating result corresponding to the edited object. , , and All are weights. The total duration of all target audio segments auditioned for the object being edited. This represents the total number of all editing operations triggered on this editable object. The total number of editing operations triggered on the edited object and retained (i.e., not deleted by the reviewed object). This refers to the number of times the operation time limit has been exceeded.

[0116] In this embodiment, the target audio to be edited can be segmented into multiple audio particle data, and the voiceprints of these multiple audio particle data can be obtained. Then, the voiceprint similarity between the voiceprints of two adjacent audio particle data is determined. Subsequently, based on the voiceprint similarity, the multiple audio particle data are merged to obtain the target audio segment. Finally, the target audio segment is assigned to an editing object for editing, resulting in the edited audio corresponding to the target audio. This method avoids segmenting the same speaker's speech into unreasonable audio segments, ensuring the fluency of the speech. Therefore, the editing object does not need to spend excessive time proofreading the fluency of the audio segments, improving the efficiency of audio editing.

[0117] The audio editing method provided in this application can be executed by an audio editing device. This application uses an audio editing device executing the audio editing method as an example to illustrate the audio editing device provided in this application.

[0118] Figure 4 This application shows a schematic diagram of the structure of an audio editing apparatus provided in an exemplary embodiment, which can achieve the following: Figure 1 The audio editing device, as shown in the embodiments, includes all or part of the following: a first acquisition module 401, a determination module 402, a merging module 403, and a second acquisition module 404.

[0119] In this embodiment, the first acquisition module 401 is used to segment the target audio to be edited into multiple audio particle data and acquire the voiceprints of the multiple audio particle data; the determination module 402 is used to determine the voiceprint similarity between the voiceprints of two adjacent audio particle data in the multiple audio particle data; the merging module 403 is used to merge the multiple audio particle data according to the voiceprint similarity to obtain a target audio segment; and the second acquisition module 404 is used to obtain the edited audio corresponding to the target audio by assigning the target audio segment to an editing object for editing.

[0120] In an optional implementation, the first acquisition module 401, when acquiring the voiceprint of the plurality of audio particle data, is specifically used for:

[0121] By inputting multiple audio particle data into a voiceprint network for voiceprint recognition, the voiceprints of the multiple audio particle data are obtained.

[0122] In an optional implementation, the first acquisition module 401, when used for acquiring the voiceprint of the multiple audio particles by inputting them into a voiceprint network for voiceprint recognition, specifically performs the following:

[0123] By performing Fourier transform on the multiple audio particle data, the Mel spectrum corresponding to the multiple audio particle data is obtained;

[0124] The first spectral feature is obtained by inputting the Mel spectrum into the first layer structure of the voiceprint network, wherein the first layer structure includes PreNet, Convolutional Conv, Long Short-Term Memory Network LSTM with Self Attention mechanism, and LSTM.

[0125] By inputting the first spectral feature into the second layer structure of the acoustic network, a second spectral feature is obtained, wherein the second layer structure is identical to the first layer structure.

[0126] By inputting the second spectral feature into the fully connected layer of the voiceprint network, voiceprints of multiple audio particle data are obtained.

[0127] In one optional implementation, the voiceprint network is determined according to the following method:

[0128] Construct a speech synthesis model, wherein the speech synthesis model includes an acoustic model and a vocoder, and the acoustic model and the vocoder are cascaded together.

[0129] The trained speech synthesis model is obtained by training the speech synthesis model.

[0130] The voiceprint network is determined based on the acoustic model in the trained speech synthesis model.

[0131] In one optional implementation, the step of training the speech synthesis model to obtain a trained speech synthesis model includes:

[0132] Multiple audio samples are segmented into granular audio data.

[0133] Perform a Fourier transform on the sample audio particle data to obtain the first Mel spectrum;

[0134] The acoustic signature of the sample audio particles is obtained by inputting the first Mel spectrum into the acoustic model;

[0135] The second Mel spectrum is obtained by inputting the voiceprint of the sample audio particles into the vocoder;

[0136] The training parameters of the acoustic model and the vocoder are updated by calculating the loss value between the first Mel spectrum and the second Mel spectrum.

[0137] If the number of training iterations meets the preset number of training iterations or the loss value meets the preset conditions, the trained speech synthesis model is obtained.

[0138] In an optional implementation, when the merging module 403 is used to merge multiple audio particle data according to the voiceprint similarity to obtain the target audio segment, it is specifically used for:

[0139] Based on the voiceprint similarity, multiple audio particle data are merged to obtain a first initial audio segment;

[0140] Based on the length of the first initial audio segment and the voiceprint similarity, the first initial audio segment is merged to obtain a second initial audio segment.

[0141] The second initial audio segment is processed to obtain the target audio segment, wherein the processing includes merging the second initial audio segment and segmenting the second initial audio segment.

[0142] In an optional implementation, when the merging module 403 performs the merging process on the first initial audio segment based on the length of the first initial audio segment and the voiceprint similarity to obtain the second initial audio segment, it is specifically used for:

[0143] Get the length of the first initial audio segment;

[0144] If the length of the first initial audio segment is less than the first length, the initial audio segments are merged based on the voiceprint similarity between the last audio particle data and the first audio particle data of the adjacent initial audio segments to obtain the second initial audio particle.

[0145] In an optional implementation, the device further includes: a recognition module, configured to obtain text information corresponding to the second initial audio segment by performing speech recognition on the second initial audio segment; to identify whether the second initial audio segment is a complete sentence based on the text information corresponding to the second initial audio segment; and, if the second initial audio segment is an incomplete sentence, to re-segment the target audio to obtain a new second initial audio segment.

[0146] In an optional implementation, when the merging module 403 is used to organize the second initial audio segment to obtain the target audio segment, it is specifically used for:

[0147] Determine the target quantity based on the number of objects to be edited;

[0148] If the number of the second initial audio segments is less than the target number, starting from the longest second initial audio segment, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment;

[0149] Get the length of the second initial audio segment;

[0150] If the length of the second initial audio segment is less than the second length, the second initial audio segment is merged with the adjacent second initial audio segment with the shorter length to obtain the target audio segment;

[0151] If the length of the second initial audio segment is greater than the third length, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment, wherein the second length is less than the third length.

[0152] In an optional implementation, the second acquisition module 404, when used to obtain the edited audio corresponding to the target audio by assigning the target audio segment to the editing object for editing, specifically performs the following:

[0153] The target audio segments are randomly assigned to multiple editing objects for collaborative editing.

[0154] The target audio segment is locked;

[0155] After the editing object completes the editing of the target audio segment, the locking of the target audio segment is removed and the review object is determined from the editing object;

[0156] Once the review object has completed the review, the edited audio corresponding to the target audio is obtained.

[0157] In an optional implementation, the second acquisition module 404 is further used for:

[0158] In response to a request from the editing object, if the editing object does not have the target audio segment being edited and the target audio segment corresponding to the request is not locked, the target audio segment is assigned to the editing object.

[0159] In an optional implementation, the second acquisition module 404 is further configured to: add an operation time limit to the target audio segment during the process of the editing object editing the target audio segment;

[0160] If the editing object fails to complete the editing of the target audio segment within the operation time limit, an alarm signal is generated, wherein the alarm signal is used to indicate a reduction in the operation time limit;

[0161] If the operation time limit is lower than the preset time limit threshold, the editing of the target audio segment by the editing object is cancelled and the locking of the target audio segment is removed.

[0162] In an optional implementation, the second acquisition module 404, when used to add an operation time limit to the target audio segment, specifically performs the following:

[0163] If the frequency of the editing object triggering the editing of the target audio segment is less than a preset threshold, the operation time limit is determined according to the length of the target audio segment and the preset amplification factor;

[0164] If the frequency is greater than or equal to the preset threshold, the operation time limit is determined based on the length of the target audio segment, the amplification factor, the frequency, the preset threshold, and the time step.

[0165] In an optional implementation, the second acquisition module 404 is further configured to acquire the score result corresponding to the editing object by scoring the editing object according to the operation time limit of the editing object.

[0166] The audio editing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0167] The audio editing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0168] The audio editing device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0169] Optionally, such as Figure 5 As shown in the illustration, this application also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the above-mentioned functions. Figure 1 The audio editing method shown here involves each step and achieves the same technical effect. To avoid repetition, it will not be described in detail here.

[0170] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0171] Figure 6 This illustration shows a structural block diagram of another electronic device 600 according to an exemplary embodiment of this application. The electronic device 600 can be implemented as a smartphone, tablet computer, laptop computer, desktop computer, smartwatch, and television, etc. The electronic device 600 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0172] Typically, electronic device 600 includes a processor 601 and a memory 602.

[0173] Processor 601 may include one or more processing cores, such as a quad-core processor or a deca-core processor. Processor 601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0174] Memory 602 may include one or more computer-readable storage media, which may be non-transitory. Memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 602 is used to store at least one instruction, which is executed by processor 601 to implement all or part of the steps in the audio editing method illustrated in the method embodiments of this application.

[0175] In some embodiments, the electronic device 600 may optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, memory 602, and peripheral device interface 603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0176] In some embodiments, the electronic device 600 further includes one or more sensors 609. The one or more sensors 609 include, but are not limited to, an accelerometer 610, a gyroscope 611, a pressure sensor 612, an optical sensor 613, and a proximity sensor 614.

[0177] Those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on the electronic device 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0178] This application also provides a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the various processes of the above-described audio editing method and achieve the same technical effect. To avoid repetition, these will not be described again here.

[0179] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0180] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-mentioned audio editing method and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0181] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0182] This application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, implement the above-described functionality. Figure 1 The steps of the audio editing method shown can achieve the same technical effect, and will not be repeated here to avoid repetition.

[0183] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0184] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An audio editing method, characterized in that, include: The target audio to be edited is divided into multiple audio particle data, and the voiceprints of multiple audio particle data are obtained; Determine the voiceprint similarity between two adjacent audio particle data in multiple audio particle data sets; Based on the voiceprint similarity, multiple audio particle data are merged to obtain the target audio segment; By assigning the target audio segment to an editing object for editing, the edited audio corresponding to the target audio is obtained; The step of merging multiple audio particle data based on the voiceprint similarity to obtain the target audio segment includes: Based on the voiceprint similarity, multiple audio particle data are merged to obtain a first initial audio segment; Based on the length of the first initial audio segment and the voiceprint similarity, the first initial audio segment is merged to obtain a second initial audio segment. The second initial audio segment is processed to obtain the target audio segment, wherein the processing includes merging the second initial audio segment and segmenting the second initial audio segment.

2. The method according to claim 1, characterized in that, The process of acquiring the voiceprint of multiple audio particle data includes: By inputting multiple audio particle data into a voiceprint network for voiceprint recognition, the voiceprints of the multiple audio particle data are obtained.

3. The method according to claim 2, characterized in that, The step of inputting multiple audio particles into a voiceprint network for voiceprint recognition to obtain the voiceprint of multiple audio particle data includes: By performing Fourier transform on the multiple audio particle data, the Mel spectrum corresponding to the multiple audio particle data is obtained; The first spectral feature is obtained by inputting the Mel spectrum into the first layer structure of the voiceprint network, wherein the first layer structure includes PreNet, Convolutional Conv, Long Short-Term Memory Network LSTM with Self Attention mechanism, and LSTM. By inputting the first spectral feature into the second layer structure of the acoustic network, a second spectral feature is obtained, wherein the second layer structure is identical to the first layer structure. By inputting the second spectral feature into the fully connected layer of the voiceprint network, voiceprints of multiple audio particle data are obtained.

4. The method according to claim 2, characterized in that, The voiceprint network is determined according to the following method: Construct a speech synthesis model, wherein the speech synthesis model includes an acoustic model with the same structure as the voiceprint network and a vocoder, and the acoustic model and the vocoder are cascaded together; The trained speech synthesis model is obtained by training the speech synthesis model. The voiceprint network is determined based on the acoustic model in the trained speech synthesis model.

5. The method according to claim 4, characterized in that, The step of training the speech synthesis model to obtain the trained speech synthesis model includes: Multiple audio samples are segmented into granular audio data. Perform a Fourier transform on the sample audio particle data to obtain the first Mel spectrum; The acoustic signature of the sample audio particles is obtained by inputting the first Mel spectrum into the acoustic model; The second Mel spectrum is obtained by inputting the voiceprint of the sample audio particles into the vocoder; The training parameters of the acoustic model and the vocoder are updated by calculating the loss value between the first Mel spectrum and the second Mel spectrum. If the number of training iterations meets the preset number of training iterations or the loss value meets the preset conditions, the trained speech synthesis model is obtained.

6. The method according to claim 1, characterized in that, The step of merging the first initial audio segment based on its length and the voiceprint similarity to obtain a second initial audio segment includes: Get the length of the first initial audio segment; If the length of the first initial audio segment is less than the first length, the first initial audio segment, the preceding audio segment, and the following audio segment are merged based on the voiceprint similarity between the last audio particle data of the preceding audio segment and the first audio particle data of the following audio segment to obtain the second initial audio particle.

7. The method according to claim 1 or 6, characterized in that, After obtaining the second initial audio segment, the method further includes: By performing speech recognition on the second initial audio segment, the text information corresponding to the second initial audio segment can be obtained; Based on the text information corresponding to the second initial audio segment, identify whether the second initial audio segment is a complete sentence; If the second initial audio segment is an incomplete sentence, the target audio is re-segmented to obtain a new second initial audio segment.

8. The method according to claim 7, characterized in that, The step of processing the second initial audio segment to obtain the target audio segment includes: Determine the target quantity based on the number of objects to be edited; If the number of the second initial audio segments is less than the target number, starting from the longest second initial audio segment, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment; Get the length of the second initial audio segment; If the length of the second initial audio segment is less than the second length, the second initial audio segment is merged with the adjacent second initial audio segment with the shorter length to obtain the target audio segment; If the length of the second initial audio segment is greater than the third length, the second initial audio segment is segmented according to the text information corresponding to the second initial audio segment to obtain the target audio segment, wherein the second length is less than the third length.

9. The method according to any one of claims 1 to 6, characterized in that, The step of assigning the target audio segment to an editing object for editing to obtain the edited audio corresponding to the target audio includes: The target audio segments are randomly assigned to multiple editing objects for collaborative editing. The target audio segment is locked; After the editing object completes the editing of the target audio segment, the locking of the target audio segment is removed and the review object is determined from the editing object; Once the review object has completed the review, the edited audio corresponding to the target audio is obtained.

10. The method according to claim 9, characterized in that, After randomly assigning multiple target audio segments to multiple editing objects for collaborative editing, the method further includes: In response to a request from the editing object, if the editing object does not have the target audio segment being edited and the target audio segment corresponding to the request is not locked, the target audio segment is assigned to the editing object.

11. The method according to claim 9, characterized in that, The method further includes: During the editing process of the target audio segment, an operation time limit is added to the target audio segment; If the editing object fails to complete the editing of the target audio segment within the operation time limit, an alarm signal is generated, wherein the alarm signal is used to indicate a reduction in the operation time limit; If the operation time limit is lower than the preset time limit threshold, the editing of the target audio segment by the editing object is cancelled and the locking of the target audio segment is removed.

12. The method according to claim 11, characterized in that, Adding the operation time limit to the target audio segment includes: If the frequency of the edited object triggering the editing of the target audio segment is less than a preset threshold, the operation time limit is determined according to the length of the target audio segment and the preset amplification factor; If the frequency is greater than or equal to the preset threshold, the operation time limit is determined based on the length of the target audio segment, the amplification factor, the frequency, the preset threshold, and the time step.

13. The method according to claim 9, characterized in that, After the editing object completes editing the target audio segment, the method further includes: The editing object is scored based on the operation time limit of the editing object, and the corresponding score result is obtained.

14. An audio editing device, characterized in that, include: The first acquisition module is used to divide the target audio to be edited into multiple audio particle data and acquire the voiceprints of the multiple audio particle data. The determination module is used to determine the voiceprint similarity between the voiceprints of two adjacent audio particle data in the plurality of audio particle data; The merging module is used to merge multiple audio particle data according to the voiceprint similarity to obtain a target audio segment; The second acquisition module is used to obtain the edited audio corresponding to the target audio by assigning the target audio segment to the editing object for editing. The step of merging multiple audio particle data based on the voiceprint similarity to obtain the target audio segment includes: Based on the voiceprint similarity, multiple audio particle data are merged to obtain a first initial audio segment; Based on the length of the first initial audio segment and the voiceprint similarity, the first initial audio segment is merged to obtain a second initial audio segment. The second initial audio segment is processed to obtain the target audio segment, wherein the processing includes merging the second initial audio segment and segmenting the second initial audio segment.

15. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the steps of the method as described in any one of claims 1 to 13.