Resource editing methods, apparatus, electronic devices and readable storage media

CN122554680APending Publication Date: 2026-08-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0019]根据本公开的又一方面,提供了一种存储有计算机指令的非瞬时计算机可读存储介质,所述计算机指令用于使所述计算机执行如上所述的方面和任一可能的实现方式的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554680A_ABST
    Figure CN122554680A_ABST
Patent Text Reader

Abstract

This disclosure provides a resource editing method, apparatus, electronic device, and readable storage medium, relating to the technical fields of audio and video processing and audio and video editing. The specific implementation involves: acquiring the resource to be edited and a specified target audio segment; the target audio segment being the audio segment to be deleted from the resource; the resource including audio; acquiring the target feature type of the target audio segment; determining a matching audio segment in the resource that matches the target audio segment based on the target feature type of the target audio segment; and removing the target audio segment and the matching audio segment from the resource.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of audio and video processing and audio and video editing, and in particular to a resource editing method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] With the popularization of short video creation, online education, podcast recording and other applications, video post-production editing has become an important part of content creation.

[0003] During video editing, creators often need to handle repetitive content in the audio, such as: the speaker's verbal tics like frequent use of interjections like "that" and "then," intermittent background noise in the recording environment like throat clearing or coughing, and recurring background music clips. Existing video editing software typically offers manual editing functions based on a timeline, requiring users to locate and delete unwanted audio segments one by one. Summary of the Invention

[0004] This disclosure provides a resource editing method, apparatus, electronic device, and readable storage medium.

[0005] According to one aspect of this disclosure, a resource editing method is provided, comprising:

[0006] Obtain the resource to be edited and the specified target audio segment; the target audio segment is the audio segment in the resource to be deleted; the resource includes audio.

[0007] Obtain the target feature type of the target audio segment;

[0008] Based on the target feature type of the target audio segment, determine the matching audio segment in the resource that matches the target audio segment;

[0009] Remove the target audio segment and the matching audio segment from the resource.

[0010] According to another aspect of this disclosure, a resource editing apparatus is provided, comprising:

[0011] The segment acquisition module is used to acquire the resource to be edited and the specified target audio segment; the target audio segment is the audio segment to be deleted from the resource; the resource includes audio.

[0012] The type acquisition module is used to acquire the target feature type of the target audio segment;

[0013] The determining module is used to determine the matching audio segments in the resource that match the target audio segment based on the target feature type of the target audio segment;

[0014] A removal module is used to remove the target audio segment and the matching audio segment from the resource.

[0015] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.

[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.

[0020] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.

[0021] According to the technology disclosed herein, the accuracy and efficiency of resource editing can be effectively improved.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0024] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;

[0025] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;

[0026] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;

[0027] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;

[0028] Figure 5 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0031] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.

[0032] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0033] The manual editing function of the timeline provided by existing video editing software is inefficient when dealing with a large number of repetitive audio clips, and it is difficult to ensure that all similar clips are accurately deleted, thus resulting in low video editing efficiency.

[0034] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown, this embodiment provides a resource editing method, which may specifically include the following steps:

[0035] S101. Obtain the resource to be edited and the specified target audio segment; the target audio segment is the audio segment in the resource to be deleted; the resource includes audio;

[0036] The resources in this embodiment include audio. For example, the resources in this embodiment can be audio resources that only include audio, or audio and video resources that include both audio and video.

[0037] In this embodiment, the target audio segment can be a user-specified audio segment from the resource to be edited. This specified target audio segment can be the audio segment that the user wants to remove during resource editing.

[0038] The execution subject of the resource editing method in this embodiment can be a resource editing device, which can be an electronic entity or a software-integrated device.

[0039] In this embodiment, the target audio segment to be deleted can be any specified audio segment from the resource to be edited.

[0040] S102. Obtain the target feature type of the target audio segment;

[0041] The target feature type of the target audio segment is used to identify the type of features that are suitable for extraction and can accurately and effectively identify the target audio segment during feature extraction.

[0042] In other words, for the target video segment, different features can be extracted during feature extraction for different feature types. Some types of features can accurately identify the target video segment, while others cannot. This step is used to obtain the target feature type corresponding to the features that can identify the target audio segment.

[0043] S103. Based on the target feature type of the target audio segment, determine the matching audio segment in the resource that matches the target audio segment;

[0044] S104. Remove the target audio segment and the matching audio segment from the resource.

[0045] Since the features corresponding to the target feature type can accurately identify the target audio segment, and by referring to the target feature type of the target audio segment, the matching audio segment in the resource that matches the target audio segment can be accurately and effectively determined. This allows for the accurate removal of the target audio segment and the matching audio segment that matches the target audio segment from the resource, thus enabling comprehensive and accurate editing of the resource.

[0046] The resource editing method of this embodiment can determine the matching audio segments in the resource that match the target audio segment based on the target feature type of the target audio segment, and then remove the target audio segment and the matching audio segments from the resource. Compared with the manual editing of the prior art, it can accurately and effectively obtain all matching audio segments and perform accurate and effective resource editing, which can effectively improve the accuracy and efficiency of resource editing.

[0047] Figure 2This is a schematic diagram based on the second embodiment of this disclosure; the resource editing method of this embodiment, in the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 2 As shown, the resource editing method in this embodiment may specifically include the following steps:

[0048] S201. Obtain the resource to be edited and the specified target audio segment; the target audio segment is the audio segment in the resource to be deleted;

[0049] For example, in this embodiment, when the resource to be edited is audio, the user can play the audio and select an audio segment to be deleted as the target audio segment. In practical applications, the duration of the target audio segment is relatively short, less than or equal to a preset duration threshold. For example, the preset duration threshold can be set to 2 seconds or 3 seconds according to actual needs.

[0050] Alternatively, when the resource to be edited is audio or video, the audio and video can be separated first, and then an audio segment to be deleted can be selected as the target audio segment. The rest of the implementation principle is the same as when the resource to be edited is audio.

[0051] S202. Using a pre-trained feature type recognition model, the feature type of the target audio segment is identified to obtain the target feature type of the target audio segment.

[0052] For example, in a practical implementation, the target audio segment is input into a feature type recognition model. This model analyzes the input target audio segment and outputs the target feature type of the audio segment. This feature type recognition model is a neural network model that has been pre-trained under supervised supervision.

[0053] In this embodiment, the target feature type can be a multi-classification model, and the identified target feature type can include a first type, a second type, or an uncertain type.

[0054] For example, if the target audio segment includes special pronunciations of specific words, such as pig noises, dog barks, or other animal sounds, or if it also includes special pronunciations of specific words by humans, then the target feature type of the target audio segment can be determined as the first type. Correspondingly, the feature extracted from the target audio segment can be the Mel Frequency Cestrum Coefficient (MFCC) feature, that is, the MFCC feature is a feature that can help identify the special pronunciations of specific words.

[0055] When the target audio segment contains verbal tics such as "this," "that," or "then," the target feature type of the target audio segment can be determined to be the second type. Correspondingly, the feature extracted from the target audio segment can be the spectral centroid feature, that is, the spectral centroid feature is a feature that can help identify verbal tics.

[0056] When it is impossible to determine whether the target feature type corresponding to the target audio segment is the first type or the second type, the target feature type of the target audio segment can be considered to be an uncertain type.

[0057] In this embodiment, a pre-trained feature type recognition model is used to accurately and effectively determine the target feature type of the target audio segment.

[0058] Alternatively, in practical applications, after selecting a specific target audio segment, the user can determine the target feature type of the target audio segment based on their personal experience and input it into the resource editing device. This method can also accurately obtain the target feature type of the target audio segment.

[0059] Optionally, in one embodiment of this disclosure, the target audio segment can also be analyzed based on other preset rules to determine the target feature type of the target audio segment, which will not be described in detail here.

[0060] S203. Segment the audio in the resource to obtain multiple audio segments to be matched;

[0061] In practice, the audio resource can be segmented based on the duration of the target audio segment to obtain multiple audio segments to be matched. However, this segmentation method does not overlap between adjacent audio segments to be matched, and it cannot cover audio segments including the latter half of the preceding audio segment and the first half of the following audio segment, resulting in incomplete coverage of the segmented audio segments to be matched.

[0062] Therefore, to improve the comprehensiveness and accuracy of the multiple audio segments to be matched obtained from segmentation, in this embodiment, the size of the sliding window can be set with reference to the duration of the target audio segment. Based on the sliding window, the audio in the resource is segmented by sliding to generate multiple audio segments to be matched; the sliding step size is smaller than the duration of the target audio segment, and can be set according to needs or experience. Preferably, the size of the sliding window can be equal to the duration of the target audio segment.

[0063] Optionally, in practical applications, the length of each audio segment to be matched may not be exactly the same as the length of the target audio segment; for example, it may be slightly longer or slightly shorter than the target audio segment. For instance, if the length of the target audio segment is 3 seconds, the length of each audio segment to be matched may be 0.29 seconds or 0.31 seconds, etc.

[0064] S204. Normalize the pitch features of the target audio segment and each audio segment to be matched.

[0065] In practice, the maximum and minimum pitch amplitudes of the target audio segment and multiple audio segments to be matched can be obtained. Furthermore, the pitches of the target audio segment and each audio segment to be matched are normalized to a range between the minimum and maximum pitch amplitudes. This step unifies all audio to the same reference pitch, improving the accuracy of subsequent matching.

[0066] S205. Based on the target feature type of the target audio segment, extract the features of the target audio segment and the features of each audio segment to be matched;

[0067] Referring to the above explanation of target feature types, it can be seen that if the target feature type of the target audio segment is the first type, the MFCC features of the target audio segment and each audio segment to be matched can be extracted separately.

[0068] If the target feature type of the target audio segment is the second type, extract the spectral centroid features of the target audio segment and each audio segment to be matched respectively;

[0069] If the target feature type of the target audio segment is uncertain, extract the MFCC features and spectral centroid features of the target audio segment and each audio segment to be matched.

[0070] For a target audio segment or each audio segment to be matched, an audio segment may include multiple frames, and each frame may include multiple MFCC features; therefore, the MFCC feature of an audio segment can be an M*N multidimensional vector, where M can be the number of frames in the audio segment, and N is the number of MFCC coefficients included in each frame. The acquisition of the MFCC coefficients for each frame can be found in relevant techniques, and will not be elaborated upon here.

[0071] For a target audio segment or various audio segments to be matched, an audio segment may include multiple frames. When calculating the spectral centroid of each frame, the audio frame can first be transformed to the frequency domain using a Fast Fourier Transform (FFT). Then, the amplitude or power at each frequency point is used as a weight, and a weighted average is calculated with the corresponding frequency value to obtain the spectral centroid of that frame. That is, each frame's spectral centroid feature includes only one element. Therefore, the spectral centroid feature of an audio segment can be an M*1 vector, where M can be the number of frames in the audio segment. The calculation of the spectral centroid of each frame can also refer to relevant techniques, which will not be elaborated here.

[0072] S206. Calculate the similarity between the features of each audio segment to be matched and the features of the target audio segment;

[0073] S207. Based on the similarity between the features of each audio segment to be matched and the features of the target audio segment, determine the matching audio segment that matches the target audio segment among the multiple audio segments to be matched.

[0074] For example, when the features of each audio segment to be matched and the features of the target audio segment are both MFCC features, the similarity between the MFCC features of each audio segment to be matched and the MFCC features of the target audio segment is calculated. Specifically, the DTW distance between the two audio segments can be calculated using the Dynamic Time Warping (DTW) algorithm. Then, based on the DTW distance, the similarity of the MFCC features of the two audio segments can be determined. The smaller the DTW distance between the two audio segments, the greater the similarity of their MFCC features; conversely, the larger the DTW distance, the smaller the similarity of their MFCC features. For detailed calculation methods, please refer to the implementation principle of the DTW algorithm, which will not be elaborated here.

[0075] The similarity of the MFCC features of two audio segments, based on the DTW distance, can be determined using the following formula: Similarity = 1 - d / D max Where d can be the normalized DTW distance between the two audio segments, D max A set distance threshold can be set. The normalized DTW distance d between two audio segments can be obtained by referencing the DTW distances between the MFCC features of each audio segment to be matched and the MFCC features of the target audio segment, and normalizing them according to a normalization strategy. For details, please refer to the relevant documentation on normalization strategies, which will not be elaborated here.

[0076] For each audio segment to be matched, if the similarity between the MFCC features of the audio segment to be matched and the MFCC features of the target audio segment is greater than or equal to a preset similarity threshold, the audio segment to be matched can be considered a matched audio segment that matches the target audio segment; otherwise, if the similarity between the MFCC features of the audio segment to be matched and the MFCC features of the target audio segment is greater than the preset similarity threshold, the audio segment to be matched can be considered a mismatch between the audio segment to be matched and the target audio segment.

[0077] Optionally, in practical applications, the DTW distance between two audio segments can be used to determine whether the audio segment to be matched matches the target audio segment. For example, if the DTW distance between the MFCC features of the audio segment to be matched and the MFCC features of the target audio segment is less than a preset distance threshold, the audio segment to be matched can be considered a matching audio segment that matches the target audio segment; otherwise, if the DTW distance between the MFCC features of the audio segment to be matched and the MFCC features of the target audio segment is greater than or equal to the preset distance threshold, the audio segment to be matched can be considered a mismatch. Optionally, the DTW distance between the two audio features can be normalized before being compared with the preset distance threshold to determine whether the audio segment to be matched matches the target audio segment, which can further improve the accuracy of matching.

[0078] When the features of each audio segment to be matched and the features of the target audio segment are both spectral centroid features, the implementation principle is the same as that of the MFCC features mentioned above, and will not be repeated here.

[0079] When the features of each audio segment to be matched and the features of the target audio segment both include MFCC features and spectral centroid features, the similarity calculation methods for MFCC features and spectral centroid features described above can be used to calculate the first similarity between the MFCC features of each audio segment to be matched and the MFCC features of the target audio segment, and the second similarity between the spectral centroid features of each audio segment to be matched and the spectral centroid features of the target audio segment. Then, it is detected whether there is a similarity greater than a preset similarity threshold in the first and second similarities. If there is, the audio segment to be matched is determined to be a matched audio segment that matches the target audio segment; otherwise, the audio segment to be matched does not match the target audio segment.

[0080] In this embodiment, following the above method, all matching audio segments can be accurately and comprehensively obtained from multiple audio segments to be matched, and the specific number of matching audio segments is not limited.

[0081] Optionally, in this embodiment, the matched audio segment can be represented by a start timestamp and an end timestamp in the audio of the resource, thus identifying the interval between the start timestamp and the end timestamp in the audio as the matched audio segment. Correspondingly, all matched audio segments can be a list of intervals including multiple pairs of start timestamps and end timestamps.

[0082] After identifying each matching audio segment, all matching audio segments can be marked on the audio display interface of the resource for manual review by the user. Once the user confirms that there are no problems, the system can receive the user's confirmation instruction to continue the subsequent operations of removing the target audio segment and matching audio segments from the resource.

[0083] Optionally, if the number of matched audio segments includes one or more, both the target audio segment and the matched audio segment are considered as audio segments to be deleted. If the number of audio segments to be deleted includes two or more, and there are adjacent or overlapping audio segments to be deleted, then the adjacent or overlapping audio segments to be deleted can be merged first to avoid duplicate deletion and resource breakage, thereby effectively improving resource editing efficiency.

[0084] S208. Remove the target audio segment and the matching audio segment from the resource.

[0085] Specifically, the resources in this embodiment can include two cases: the first is audio. In this case, when step S208 is implemented, the target audio segment and the matching audio segment can be directly deleted from the audio, and then all the remaining audio segments after deletion can be merged to obtain the edited audio.

[0086] The second type is audio and video, which means that the resource includes both audio and video. In this case, it is necessary to remove the target audio segment and the matching audio segment from the audio and video.

[0087] The resource editing method in this embodiment is applicable not only to audio editing but also to audio and video editing, offering high flexibility.

[0088] Specifically, removing the target audio segment and the matching audio segment from audio and video can include the following two methods with specific steps:

[0089] The first method may specifically include the following steps:

[0090] (1) Remove the target audio segment and the matching audio segment from the audio of the audio and video;

[0091] (2) Delete the target video segment corresponding to the target audio segment and the matching video segment corresponding to the matching audio segment from the audio and video video;

[0092] (3) Based on the audio of the deleted target audio segment and the matching audio segment, and the video of the deleted target video segment and the matching video segment, synthesize the processed audio and video.

[0093] In this embodiment, the method of audio and video synthesis can be found in relevant technologies, and will not be repeated here.

[0094] Since the matched audio segments are usually short, deleting both the target video segment corresponding to the audio segment and the matching video segment corresponding to the matched audio segment, as described in this embodiment, typically does not affect the overall audio and video quality. This method allows for fast, accurate, and effective audio and video editing.

[0095] The second method may specifically include the following steps:

[0096] (a) Obtain the first reference audio segment adjacent to the target audio segment and the second reference audio segment adjacent to the matching audio segment in the audio of the audio and video respectively; the duration of the first reference audio segment and the second reference audio segment is a preset duration, which is greater than the duration of the matching audio segment and also greater than the duration of the target audio segment;

[0097] (b) Using a pre-trained audio generation model, a first replacement audio segment is generated based on the audio of the first reference audio segment to replace the target audio segment;

[0098] The audio generation model can be a pre-trained neural network model. When used, the audio of a first reference audio segment is input into the model, which then generates a first replacement audio segment based on the first reference audio segment, capable of representing the target audio segment. The duration of the first replacement audio segment can be determined during the training of the audio generation model.

[0099] (c) Use the first replacement audio clip to replace the corresponding target audio segment in the audio;

[0100] (d) Using a pre-trained audio generation model, a second replacement audio segment is generated based on the audio of the second reference audio segment to replace the matching audio segment;

[0101] In this embodiment, if the target audio segment is at the beginning of the audio in the audio / video, the first reference audio segment is located after the target audio is determined; if the target audio segment is at the end of the audio in the audio / video, the first reference audio segment is located before the target audio is determined; if the target audio segment is in the middle of the audio in the audio / video, the first reference audio segment can be the audio segment before the target audio is determined, or the audio segment after the target audio is determined, or a combination of both. The acquisition of the second reference audio segment follows the same principle and will not be described further.

[0102] (e) Use the second replacement audio clip to replace the corresponding matching audio segment in the audio;

[0103] It should be noted that if there are multiple matching audio segments, each matching audio segment needs to be replaced using steps (d) and (e).

[0104] (e) Synthesize the processed audio and video based on the replaced audio and video.

[0105] Unlike the first method described above, this method does not directly delete the target audio segment and the matching audio segment. Instead, it uses a first reference audio segment adjacent to the target audio segment to regenerate a first replacement audio segment that can replace the target audio segment, and then replaces the target audio segment. Similarly, it uses a second reference audio segment adjacent to the matching audio segment to regenerate a second replacement audio segment that can replace the matching audio segment, and then replaces the target audio segment. By using this method to regenerate replacement audio segments using adjacent reference audio segments, and because the reference audio segments do not contain content from the target audio segment, it effectively ensures that the generated replacement audio segments will also not include content related to the target audio segment.

[0106] In this embodiment, a pre-trained audio generation model is used to generate a first replacement audio segment that can replace the target audio segment and a second replacement audio segment that can replace the matching audio segment, based on the audio of the first reference audio segment and the second reference audio segment, respectively. This makes the first replacement audio segment and the second replacement audio segment connect more smoothly with the adjacent audio segments in the audio.

[0107] This method also allows for accurate and effective editing of audio and video, while ensuring that the length of the audio and video remains unchanged.

[0108] The resource editing method in this embodiment can automatically edit resources, avoiding manual editing and effectively improving the accuracy and efficiency of resource editing.

[0109] The resource editing method in this embodiment extracts features of the target audio segment and features of each audio segment to be matched based on the target feature type of the target audio segment; calculates the similarity between the features of each audio segment to be matched and the features of the target audio segment; and then determines the matching audio segments among multiple audio segments to be matched that match the target audio segment. This method can comprehensively and accurately obtain all matching audio segments in the audio of the resource that match the target audio segment, thereby effectively improving the comprehensiveness and accuracy of resource editing and thus improving resource editing efficiency.

[0110] The resource editing method in this embodiment extracts features of different target audio segments and audio segments to be matched based on different target feature types of the target audio segments. This can effectively improve the accuracy of the obtained matching audio segments, thereby effectively improving the accuracy and efficiency of resource editing.

[0111] Figure 3 This is a schematic diagram based on the third embodiment of this disclosure; as shown Figure 3 As shown, this embodiment provides a resource editing device 300, including:

[0112] The segment acquisition module 301 is used to acquire the resource to be edited and the specified target audio segment; the target audio segment is the audio segment to be deleted from the resource; the resource includes audio.

[0113] Type acquisition module 302 is used to acquire the target feature type of the target audio segment;

[0114] The determining module 303 is used to determine the matching audio segment in the resource that matches the target audio segment based on the target feature type of the target audio segment;

[0115] The removal module 304 is used to remove the target audio segment and the matching audio segment from the resource.

[0116] The resource editing device 300 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0117] Figure 4 This is a schematic diagram based on the fourth embodiment of the present disclosure; as shown Figure 4 As shown, the resource editing device 400 of this embodiment, in the above-described... Figure 3 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 4 As shown, the resource editing device 400 of this embodiment includes the above-described... Figure 3The modules with the same name and function shown are: fragment acquisition module 401, type acquisition module 402, determination module 403, and removal module 404.

[0118] In this embodiment, the type acquisition module 402 is used for:

[0119] Obtain the target feature type of the input target audio segment; or

[0120] A pre-trained feature type recognition model is used to identify the feature type of the target audio segment, thereby obtaining the target feature type of the target audio segment.

[0121] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the determining module 403 includes:

[0122] The segmentation unit 4031 is used to segment the audio in the resource to obtain multiple audio segments to be matched;

[0123] Extraction unit 4032 is used to extract features of the target audio segment and features of each audio segment to be matched based on the target feature type of the target audio segment.

[0124] The calculation unit 4033 is used to calculate the similarity between the features of each of the audio segments to be matched and the features of the target audio segment;

[0125] The determining unit 4034 is used to determine the matching audio segment that matches the target audio segment among the plurality of audio segments to be matched based on the similarity between the features of each audio segment to be matched and the features of the target audio segment.

[0126] Further optionally, in one embodiment of this disclosure, the segmentation unit 4031 is used for:

[0127] The size of the sliding window is set with reference to the duration of the target audio segment, and the audio in the resource is slidably segmented based on the sliding window to generate the multiple audio segments to be matched.

[0128] Further optionally, in one embodiment of this disclosure, the extraction unit 4032 is used for:

[0129] If the target feature type of the target audio segment is the first type, extract the Mel frequency cepstral coefficient features of the target audio segment and each of the audio segments to be matched respectively;

[0130] If the target feature type of the target audio segment is the second type, extract the spectral centroid features of the target audio segment and each of the audio segments to be matched respectively;

[0131] If the target feature type of the target audio segment is uncertain, the Mel frequency cepstral coefficient features and spectral centroid features of the target audio segment and each of the audio segments to be matched are extracted respectively.

[0132] Further optionally, in one embodiment of this disclosure, the computing unit 4033 is used for:

[0133] If the feature type of the target audio segment is uncertain, for each audio segment to be matched, calculate the first similarity between the Mel frequency cepstral coefficient feature of the audio segment to be matched and the Mel frequency cepstral coefficient feature of the target audio segment, and the second similarity between the spectral centroid feature of the audio segment to be matched and the spectral centroid feature of the target audio segment.

[0134] Detect whether there is a similarity greater than a preset similarity threshold between the first similarity and the second similarity;

[0135] If it exists, the audio segment to be matched is determined to be the matching audio segment that matches the target audio segment.

[0136] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the determining module 403 further includes:

[0137] The normalization processing unit 4035 is used to normalize the pitch features of the target audio segment and each audio segment to be matched.

[0138] Further optionally, in one embodiment of this disclosure, the removal module 404 is configured to:

[0139] When the resource consists only of audio, delete the target audio segment and the matching audio segment from the audio.

[0140] When the resource is an audio-visual resource including audio and video, the target audio segment and the matching audio segment are removed from the audio-visual resource.

[0141] Further optionally, in one embodiment of this disclosure, the removal module 404 is configured to:

[0142] Remove the target audio segment and the matching audio segment from the audio of the audio / video recording;

[0143] Delete the target video segment corresponding to the target audio segment and the matching video segment corresponding to the matching audio segment from the video of the audio and video;

[0144] The resulting audio and video are synthesized based on the audio of the deleted target audio segment and the audio of the deleted matching audio segment, and the video of the deleted target video segment and the video of the deleted matching video segment.

[0145] Further optionally, in one embodiment of this disclosure, the removal module 404 is configured to:

[0146] The first reference audio segment adjacent to the target audio segment and the second reference audio segment adjacent to the matching audio segment are respectively obtained from the audio of the audio and video; the duration of the first reference audio segment and the second reference audio segment is a preset duration, which is greater than the duration of the matching audio segment and also greater than the duration of the target audio segment;

[0147] Using a pre-trained audio generation model, a first replacement audio segment is generated based on the audio of the first reference audio segment to replace the target audio segment;

[0148] The first replacement audio segment is used to replace the corresponding target audio segment in the audio;

[0149] Using a pre-trained audio generation model, a second replacement audio segment is generated based on the audio of the second reference audio segment to replace the matching audio segment;

[0150] The second replacement audio segment is used to replace the corresponding matching audio segment in the audio;

[0151] Based on the replaced audio and the video in the audio-video, a synthesized audio-video is generated.

[0152] The resource editing device 400 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0153] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.

[0154] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0155] Figure 5A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0156] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0157] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0158] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0163] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0164] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0165] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A resource editing method, comprising: Obtain the resource to be edited and the specified target audio segment; the target audio segment is the audio segment in the resource to be deleted; the resource includes audio. Obtain the target feature type of the target audio segment; Based on the target feature type of the target audio segment, determine the matching audio segment in the resource that matches the target audio segment; Remove the target audio segment and the matching audio segment from the resource.

2. The method according to claim 1, wherein, Obtaining the target feature type of the target audio segment includes: Obtain the target feature type of the input target audio segment; or A pre-trained feature type recognition model is used to identify the feature type of the target audio segment, thereby obtaining the target feature type of the target audio segment.

3. The method according to claim 1, wherein, Based on the target feature type of the target audio segment, determine the matching audio segments in the resource that match the target audio segment, including: The audio in the resource is segmented to obtain multiple audio segments to be matched; Based on the target feature type of the target audio segment, the features of the target audio segment and the features of each audio segment to be matched are extracted respectively; Calculate the similarity between the features of each of the audio segments to be matched and the features of the target audio segment; Based on the similarity between the features of each audio segment to be matched and the features of the target audio segment, the matching audio segment that matches the target audio segment among the plurality of audio segments to be matched is determined.

4. The method according to claim 3, wherein, The audio in the resource is segmented to obtain multiple audio segments to be matched, including: The size of the sliding window is set with reference to the duration of the target audio segment, and the audio in the resource is slidably segmented based on the sliding window to generate the multiple audio segments to be matched.

5. The method according to claim 3, wherein, Based on the target feature type of the target audio segment, the features of the target audio segment and the features of each of the audio segments to be matched are extracted, including: If the target feature type of the target audio segment is the first type, extract the Mel frequency cepstral coefficient features of the target audio segment and each of the audio segments to be matched respectively; If the target feature type of the target audio segment is the second type, extract the spectral centroid features of the target audio segment and each of the audio segments to be matched respectively; If the target feature type of the target audio segment is uncertain, the Mel frequency cepstral coefficient features and spectral centroid features of the target audio segment and each of the audio segments to be matched are extracted respectively.

6. The method according to claim 5, wherein, If the feature type of the target audio segment is uncertain, based on the similarity between the features of each audio segment to be matched and the features of the target audio segment, the matching audio segment that matches the target audio segment among the plurality of audio segments to be matched is determined, including: For each of the audio segments to be matched, calculate the first similarity between the Mel frequency cepstral coefficient features of the audio segment to be matched and the Mel frequency cepstral coefficient features of the target audio segment, and the second similarity between the spectral centroid features of the audio segment to be matched and the spectral centroid features of the target audio segment. Detect whether there is a similarity greater than a preset similarity threshold between the first similarity and the second similarity; If it exists, the audio segment to be matched is determined to be the matching audio segment that matches the target audio segment.

7. The method according to claim 3, wherein, After segmenting the audio in the resource to obtain multiple audio segments to be matched, and before extracting the features of the target audio segment and the features of each of the audio segments to be matched based on the target feature type of the target audio segment, the method further includes: The pitch features of the target audio segment and each audio segment to be matched are normalized.

8. The method according to any one of claims 1-7, wherein, Removing the target audio segment and the matching audio segment from the resource includes: When the resource consists only of audio, delete the target audio segment and the matching audio segment from the audio. When the resource is an audio-visual resource including audio and video, the target audio segment and the matching audio segment are removed from the audio-visual resource.

9. The method according to claim 8, wherein, Removing the target audio segment and the matching audio segment from the audio and video includes: Remove the target audio segment and the matching audio segment from the audio of the audio / video recording; Delete the target video segment corresponding to the target audio segment and the matching video segment corresponding to the matching audio segment from the video of the audio and video; The resulting audio and video are synthesized based on the audio of the deleted target audio segment and the audio of the deleted matching audio segment, and the video of the deleted target video segment and the video of the deleted matching video segment.

10. The method according to claim 8, wherein, Removing the target audio segment and the matching audio segment from the audio and video includes: The first reference audio segment adjacent to the target audio segment and the second reference audio segment adjacent to the matching audio segment are respectively obtained from the audio of the audio and video; the duration of the first reference audio segment and the second reference audio segment is a preset duration, which is greater than the duration of the matching audio segment and also greater than the duration of the target audio segment; Using a pre-trained audio generation model, a first replacement audio segment is generated based on the audio of the first reference audio segment to replace the target audio segment; The first replacement audio segment is used to replace the corresponding target audio segment in the audio; Using a pre-trained audio generation model, a second replacement audio segment is generated based on the audio of the second reference audio segment to replace the matching audio segment; The second replacement audio segment is used to replace the corresponding matching audio segment in the audio; Based on the replaced audio and the video in the audio-video, a synthesized audio-video is generated.

11. A resource editing device, comprising: The segment acquisition module is used to acquire the resource to be edited and the specified target audio segment; the target audio segment is the audio segment to be deleted from the resource; the resource includes audio. The type acquisition module is used to acquire the target feature type of the target audio segment; The determining module is used to determine the matching audio segments in the resource that match the target audio segment based on the target feature type of the target audio segment; The removal module is used to remove the target audio segment and the matching audio segment from the resource.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.