Audio trimming method and device

CN116612784BActive Publication Date: 2026-09-01CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310741308.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-09-01
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

传统的国际声修复方法是人工用音频处理软件对比国际声、过审版视频、mix轨音频(从过审版视频介质中提取出来的音频),通过肉眼比对、音频切分删除、视频点位复核的步骤完成修复,效率低下且价格昂贵

Benefits of technology

[0036] In a fifth aspect of the invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the audio trimming methods described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612784B_ABST
    Figure CN116612784B_ABST
Patent Text Reader

Abstract

This invention provides an audio trimming method and apparatus. The method involves sequentially matching a first sub-audio segment within a first audio segment with second sub-audio segments of the same order within a second audio segment. Based on the unmatched first and second sub-audio segments, the starting coordinates and misalignment point coordinates of the audio segment to be trimmed are determined. Then, the method sequentially matches a third sub-audio segment after the misalignment point and one or more fourth sub-audio segments after the starting coordinates. Based on the successfully matched fourth audio segments, the ending coordinates of the audio segment to be trimmed are determined. The method locates and trims the first audio segment based on the starting and ending coordinates of the audio segment to be trimmed within the first audio segment. This improves the efficiency of locating and trimming differential audio content within the audio segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to audio trimming methods and apparatus. Background Technology

[0002] The prerequisite for dubbing a film or television work is having international audio that matches the online video medium. International audio refers to all the sounds in a film or television work except for the dialogue in the language of the country of production. This mainly includes the music and action effects, and may also include other languages ​​besides the language of the country of production and the language of the dubbing country. Due to censorship, many film and television suppliers provide original international audio, which is incompatible with the approved video medium, resulting in many extra international audio clips corresponding to the deleted video segments.

[0003] Therefore, it is necessary to trim the international audio provided by film and television suppliers, cutting out the extra audio segments. This process can also be called international audio restoration. The traditional method of international audio restoration is to manually compare the international audio, the approved video, and the mix track audio (audio extracted from the approved video media) using audio processing software. The restoration is completed through steps such as visual comparison, audio segmentation and deletion, and video point verification. This method is inefficient and expensive. Summary of the Invention

[0004] The purpose of this invention is to provide an audio trimming method and apparatus to improve the efficiency of locating and trimming differential audio content. The specific technical solution is as follows:

[0005] In a first aspect of this invention, an audio trimming method is provided, the method comprising:

[0006] Obtain a first audio segment and a second audio segment; wherein, the first audio segment is a standard audio segment corresponding to the complete version of the video, and the second audio segment is a non-standard audio segment separated from the incomplete version of the video;

[0007] Audio fingerprints are extracted from the first audio segment and the second audio segment respectively, and the audio fingerprints of each audio frame in the first audio segment and the second audio segment are obtained respectively.

[0008] Using the number of first audio frames as the segment length, one or more first sub-audio segments are sequentially determined in the first audio segment, and one or more second sub-audio segments are sequentially determined in the second audio segment. The first sub-audio segments and second sub-audio segments in the same order are matched sequentially until the i-th first sub-audio segment and the i-th second sub-audio segment fail to match. The starting coordinates of the audio segment to be clipped in the first audio segment are determined based on the coordinates of the audio frames in the i-th first sub-audio segment, and the coordinates of the non-aligned points in the second audio segment are determined based on the coordinates of the audio frames in the i-th second sub-audio segment. A successful match between the first sub-audio segment and the second audio segment indicates that the similarity between the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment meets a preset condition.

[0009] Using the number of second audio frames as the segment length, a third sub-audio segment is determined in the second audio segment located after the non-alignment point coordinates, and one or more fourth sub-audio segments are determined in the first audio segment located after the starting coordinates. The third sub-audio segment and one or more fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment.

[0010] Based on the start and end coordinates of the audio segment to be cropped in the first audio segment, the first audio segment is cropped to obtain a standard audio segment that is compatible with the incomplete video.

[0011] Optionally, the matching status of the first sub-audio segment and the second sub-audio segment can be determined based on the following method:

[0012] Cross-correlation calculation is performed on the audio fingerprints corresponding to the first sub-audio segment and the second sub-audio segment to obtain a similarity sequence. It is then determined whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence.

[0013] If so, confirm that the first sub-audio segment and the second sub-audio segment have successfully matched;

[0014] If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

[0015] Optionally, the step of determining the starting coordinates of the audio segment to be trimmed in the first audio segment based on the coordinates of the audio frame in the i-th first sub-audio segment, and determining the coordinates of the non-alignment point in the second audio segment based on the position coordinates of the i-th second sub-audio segment, includes:

[0016] Using the number of the third audio frames as the shortening step, the i-th first sub-audio segment and the i-th second sub-audio segment are shortened simultaneously, and the shortened first sub-audio segment and the shortened second sub-audio segment are matched until the shortened first sub-audio segment and the shortened second sub-audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the successfully matched i-th second sub-audio segment are used as the coordinates of the non-aligned point in the second audio segment.

[0017] Optionally, the method further includes:

[0018] The first audio segment after the termination coordinate is determined as the new first audio segment, and the second audio segment after the non-aligned point coordinate is determined as the new second audio segment. The process returns to the steps of using the number of first audio frames as the segment length, sequentially determining one or more first sub-audio segments in the first audio segment, sequentially determining one or more second sub-audio segments in the second audio segment, and sequentially matching the first and second sub-audio segments in the same order. This process obtains the start and end coordinates of the audio segment to be trimmed determined for the new first audio segment. The process returns to the steps of determining the first audio segment after the current termination coordinate as the new first audio segment and the second audio segment after the current non-aligned point coordinate as the new second audio segment, until both the first and second sub-audio segments are successfully matched.

[0019] Optionally, the step of cropping the first audio segment based on the start and end coordinates of the audio segment to be cropped within the first audio segment includes:

[0020] For each identified audio segment to be trimmed, the first audio segment is trimmed based on the start and end coordinates corresponding to that audio segment.

[0021] Optionally, the complete video is the submitted version, the incomplete video is the approved version after the submitted version has been edited, the standard audio segment is the standard international audio corresponding to the complete video, and the non-standard audio segment is the international audio separated from the approved video.

[0022] In a second aspect of the invention, an audio trimming device is also provided, comprising:

[0023] The acquisition module is used to acquire a first audio segment and a second audio segment; wherein, the first audio segment is a standard audio segment corresponding to the complete version of the video, and the second audio segment is a non-standard audio segment separated from the incomplete version of the video;

[0024] The extraction module is used to extract audio fingerprints from the first audio segment and the second audio segment respectively, and obtain the audio fingerprint of each audio frame in the first audio segment and the second audio segment respectively;

[0025] The first matching module is used to determine one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment, using the number of first audio frames as the segment length. It then matches the first and second sub-audio segments in the same order until the i-th first and second sub-audio segments fail to match. Based on the coordinates of the audio frames in the i-th first sub-audio segment, it determines the starting coordinates of the audio segment to be trimmed within the first audio segment. Based on the coordinates of the audio frames in the i-th second sub-audio segment, it determines the coordinates of the non-aligned points in the second audio segment. Successful matching of the first and second sub-audio segments indicates that the similarity between the audio fingerprints corresponding to the first and second sub-audio segments meets a preset condition.

[0026] The second matching module is used to determine a third sub-audio segment in the second audio segment located after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to determine one or more fourth sub-audio segments in the first audio segment located after the start coordinates. The third sub-audio segment and one or more fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment.

[0027] The cropping module is used to crop the first audio segment based on the start and end coordinates of the audio segment to be cropped in the first audio segment, so as to obtain a standard audio segment that is compatible with the incomplete video.

[0028] Optionally, the first matching module includes:

[0029] The judgment unit is used to perform cross-correlation calculation on the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment to obtain a similarity sequence, and to determine whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence.

[0030] If so, confirm that the first sub-audio segment and the second sub-audio segment have successfully matched;

[0031] If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

[0032] In a third aspect of the present invention, an electronic device is provided, the electronic device including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0033] Memory, used to store computer programs;

[0034] A processor, when executing a program stored in memory, implements the audio trimming method according to any one of the preceding claims.

[0035] In a fourth aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements any of the audio trimming methods described above.

[0036] In a fifth aspect of the invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the audio trimming methods described above.

[0037] The audio trimming method provided in this invention involves sequentially determining one or more first sub-audio segments in a first audio segment, determining one or more second sub-audio segments in a second audio segment, and sequentially matching the first and second sub-audio segments in the same order. Based on the coordinates of audio frames in the unmatched first sub-audio segments, the starting coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the coordinates of audio frames in the unmatched second sub-audio segments, the coordinates of the misalignment points within the second audio segment are determined. A third sub-audio segment is determined within the second audio segment following the misalignment point coordinates. One or more fourth sub-audio segments are determined within the first audio segment following the starting coordinates. The third and fourth sub-audio segments are sequentially matched. Based on the coordinates of audio frames in the successfully matched fourth sub-audio segments, the ending coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the starting and ending coordinates of the audio segment to be trimmed within the first audio segment, the first audio segment is trimmed to obtain a standard audio segment adapted to the incomplete video. The audio trimming method provided in this invention locates the audio segment to be trimmed by matching sub-audio segments in the first and second audio segments. This eliminates the need for manual comparison between the first and second audio segments, enabling the location and trimming of audio content that is additional to the first audio segment compared to the second audio segment. This yields a standard audio segment that is consistent with the incomplete audio video, improving the efficiency of locating and trimming differing audio content. When applied to the restoration of international audio, this method can enhance restoration efficiency and reduce restoration costs. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0039] Figure 1 This is a flowchart illustrating the audio trimming method provided in an embodiment of the present invention;

[0040] Figure 2 This is an example diagram of a first audio segment and a second audio segment provided in an embodiment of the present invention;

[0041] Figure 3 This is an example diagram showing the starting coordinates of the audio segment to be trimmed in the first audio segment and the coordinates of the non-aligned point in the second audio segment, provided by an embodiment of the present invention.

[0042] Figure 4 This is an example diagram of the start and end coordinates of the audio segment to be trimmed, as well as the coordinates of the non-aligned point, provided in an embodiment of the present invention.

[0043] Figure 5 This is an example diagram of the cross-correlation calculation process provided in an embodiment of the present invention;

[0044] Figure 6 This is an example diagram illustrating the matching process of the first sub-audio segment and the second sub-audio segment provided in an embodiment of this disclosure;

[0045] Figure 7 This is another example diagram of the first and second audio segments provided in the embodiments of the present invention;

[0046] Figure 8 This is a schematic diagram of an audio trimming method provided in an embodiment of the present invention;

[0047] Figure 9 This is a schematic diagram of the audio trimming device provided in an embodiment of the present invention;

[0048] Figure 10 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0050] To address the inefficiency of current manual methods for audio restoration, this invention provides an audio trimming method. Figure 1 This is a flowchart illustrating the audio trimming method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method specifically includes the following steps:

[0051] Step S101: Obtain the first audio segment and the second audio segment; wherein, the first audio segment is the standard audio segment corresponding to the complete video, and the second audio segment is the non-standard audio segment separated from the incomplete video.

[0052] Step S102: Extract audio fingerprints from the first audio segment and the second audio segment respectively, and obtain the audio fingerprint of each audio frame in the first audio segment and the second audio segment respectively.

[0053] Step S103: Using the number of first audio frames as the segment length, determine one or more first sub-audio segments in the first audio segment, and determine one or more second sub-audio segments in the second audio segment. Match the first and second sub-audio segments in the same order until the i-th first sub-audio segment and the i-th second sub-audio segment fail to match. Determine the starting coordinates of the audio segment to be trimmed in the first audio segment based on the coordinates of the audio frames in the i-th first sub-audio segment, and determine the coordinates of the non-aligned point in the second audio segment based on the coordinates of the audio frames in the i-th second sub-audio segment. Successful matching of the first and second sub-audio segments indicates that the similarity between the audio fingerprints corresponding to the first and second sub-audio segments meets the preset conditions.

[0054] Step S104: Using the number of second audio frames as the segment length, determine a third sub-audio segment in the second audio segment located after the non-alignment point coordinates, and determine one or more fourth sub-audio segments in the first audio segment located after the start coordinates. Match the third sub-audio segment and one or more fourth sub-audio segments sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. Determine the termination coordinates of the audio segment to be trimmed in the first audio segment based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment.

[0055] Step S105: Based on the start and end coordinates of the audio segment to be cropped in the first audio segment, crop the first audio segment to obtain a standard audio segment that is compatible with the incomplete video.

[0056] The audio trimming method provided in this invention involves sequentially determining one or more first sub-audio segments in a first audio segment, determining one or more second sub-audio segments in a second audio segment, and sequentially matching the first and second sub-audio segments in the same order. Based on the coordinates of audio frames in the unmatched first sub-audio segments, the starting coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the coordinates of audio frames in the unmatched second sub-audio segments, the coordinates of the misalignment points within the second audio segment are determined. A third sub-audio segment is determined within the second audio segment following the misalignment point coordinates. One or more fourth sub-audio segments are determined within the first audio segment following the starting coordinates. The third and fourth sub-audio segments are sequentially matched. Based on the coordinates of audio frames in the successfully matched fourth sub-audio segments, the ending coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the starting and ending coordinates of the audio segment to be trimmed within the first audio segment, the first audio segment is trimmed to obtain a standard audio segment adapted to the incomplete video. The audio trimming method provided in this invention locates the audio segment to be trimmed by matching sub-audio segments in the first and second audio segments. This eliminates the need for manual comparison between the first and second audio segments, enabling the location and trimming of audio content that is additional to the first audio segment compared to the second audio segment. This yields a standard audio segment that is consistent with the incomplete audio video, improving the efficiency of locating and trimming differing audio content. When applied to the restoration of international audio, this method can enhance restoration efficiency and reduce restoration costs.

[0057] The steps S101-S105 described above will be explained in detail below:

[0058] In step S101, the incomplete video is missing some video content compared to the complete video. Specifically, the first audio segment is an audio segment adapted to the complete video, while the second audio segment is an audio segment separated from the incomplete video based on a separation algorithm. For example, the first audio segment could be human voice / background sound audio adapted to the complete video, while the second audio segment is human voice / background sound audio separated from the incomplete video based on a human voice / background sound separation algorithm.

[0059] Due to the limited accuracy of the separation algorithm, the second audio segment may contain some noise, and the audio quality may not meet the requirements. The first audio segment, compared to the second, contains additional audio content. This additional audio content corresponds to the extra content in the complete video compared to the incomplete video. The audio trimming method provided in this embodiment of the invention locates the position coordinates of the missing audio in the first audio segment compared to the second audio segment, and trims the first audio segment based on the location results, thereby obtaining human voice / background sound audio that is compatible with the incomplete audio and meets the audio quality requirements.

[0060] Figure 2 This is an example diagram of the first and second audio segments provided in an embodiment of the present invention, such as... Figure 2 As shown, the first audio segment contains standard audio segments m, k, and n, while the second audio segment contains non-standard audio segments m' and n'. Specifically, audio segments m' and m correspond to the same audio content, but their audio quality differs. The same applies to audio segments n' and n. Furthermore, the second audio segment lacks the audio content corresponding to audio segment k compared to the first audio segment.

[0061] In step S102, the audio fingerprint of each audio frame in the first audio segment and the second audio segment is extracted, wherein the audio fingerprint of each audio frame is specifically an array determined based on the acoustic features of that audio frame.

[0062] For methods of audio fingerprint extraction, please refer to relevant technical documents; this embodiment of the invention does not limit the scope of the invention. As an example, audio fingerprints can be extracted from the first and second audio segments using Shazam (an audio fingerprint extraction algorithm).

[0063] In step S103, using the number of first audio frames as the segment length, one or more first sub-audio segments are sequentially determined within the first audio segment, and one or more second sub-audio segments are sequentially determined within the second audio segment. Specifically, based on the audio frames in the first and second audio segments, the first and second audio segments are segmented, and each determined first or second sub-audio segment consists of the number of first audio frames. The number of first audio frames can be set based on actual needs.

[0064] As an example, if the number of first audio frames is set to 10, then the first 10 audio frames in the first audio segment are the first first sub-audio segment, the 11th to 20th audio frames are the second first sub-audio segment, and so on. The process of determining the second sub-audio segment in the second audio segment is the same.

[0065] The process of matching the first and second sub-audio segments involves sequentially matching the first and second sub-audio segments in the same order until the i-th first and second sub-audio segments fail to match. That is, the first first and second sub-audio segments are matched; if a match is found, the second first and second sub-audio segments are matched, and so on, until the first and second sub-audio segments that failed to match are identified.

[0066] Among them, the first sub-audio segment and the second sub-audio segment are successfully matched, which specifically indicates that the similarity between the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment meets the preset conditions.

[0067] When calculating the similarity between the audio fingerprints corresponding to the first and second sub-audio segments, the audio fingerprints of each audio frame in the first sub-audio segment are concatenated to obtain one array, and the audio fingerprints of each audio frame in the second sub-audio segment are concatenated to obtain another array. The similarity between these two arrays is then calculated. For specific methods on calculating the similarity between audio fingerprints, please refer to relevant technical documentation.

[0068] The preset similarity criteria used to determine whether the first and second sub-audio segments match successfully can be selected according to actual needs, and this embodiment of the invention does not limit this. For example, a similarity threshold can be preset; if the similarity between the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment is not less than the similarity threshold, then the match is considered successful.

[0069] After determining that the i-th first sub-audio segment and the i-th second sub-audio segment did not match successfully, the starting coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the i-th first sub-audio segment, and the coordinates of the non-alignment point in the second audio segment are determined based on the coordinates of the audio frames in the i-th second sub-audio segment.

[0070] In this context, a misalignment point in the second audio segment can be understood as follows: The second sub-audio segment preceding that point in the second audio segment corresponds to the audio content of the first sub-audio segment in the same order in the first audio segment, meaning a successful match is achieved. Conversely, the second sub-audio segment following that point in the second audio segment does not align with the audio content of the first sub-audio segment in the same order in the first audio segment, meaning a failed match is achieved. For example, Figure 2 The point of intersection between audio segment m' and audio segment n' in the second audio segment shown in the diagram is the non-alignment point.

[0071] Regarding the coordinates of audio frames, each audio frame in the first and second audio segments typically has a fixed duration. Therefore, the start and end coordinates of each audio frame in the first and second audio segments can be considered known. As an example, if the duration of each audio frame is 20ms, then the start and end coordinates of the first audio frame in the first audio segment are 0ms and 20ms, respectively; the start and end coordinates of the second audio frame are 20ms and 40ms, respectively; and so on.

[0072] As an example, since the duration of each audio frame is fixed, the sequence number of the audio frame can also be used to represent the time point in the audio. Therefore, the sequence number of the audio frame can also be used to represent the starting coordinates of the audio segment to be cut in the first audio segment and the coordinates of the non-aligned point in the second audio segment. For example, the starting coordinates of the audio segment to be cut in the first audio segment are the 100th frame in the first audio segment.

[0073] As an example, the starting coordinates of the audio segment to be cropped in the first audio segment and the coordinates of the non-aligned point in the second audio segment can also be represented by the time of the audio frame. For example, the starting coordinates of the audio segment to be cropped in the first audio segment are the 100th second in the first audio segment.

[0074] As mentioned earlier, the first and second audio segments each contain multiple audio frames. The starting coordinates and misalignment point coordinates of the audio segment to be trimmed are determined based on the specific audio frame's coordinates, whichever is chosen first. For example, the starting coordinates of the audio segment to be trimmed within the first audio segment can be determined based on the coordinates of the first audio frame in the i-th first sub-audio segment, and the misalignment point coordinates in the second audio segment can be determined based on the coordinates of the first audio frame in the i-th second sub-audio segment. For instance, if the number of first audio frames is set to 10, and the first first sub-audio segment and the first second sub-audio segment match successfully, but the second first sub-audio segment and the second second sub-audio segment do not match successfully, then the starting coordinates of the audio segment to be trimmed within the first audio segment are frame 11, and the misalignment point coordinates within the second audio segment are also frame 11.

[0075] As an example, the number of first audio frames can be set to 1 frame, then each first and second sub-audio segment specifically constitutes one audio frame. Therefore, each audio frame in the first audio segment can be sequentially matched with the audio frames in the same order in the second audio segment until the i-th audio frame in the first audio segment fails to match the i-th audio frame in the second audio segment. In this case, the starting coordinates of the audio segment to be trimmed in the first audio segment are the i-th audio frame, and the coordinates of the non-aligned point in the second audio segment are also the i-th audio frame.

[0076] The following will illustrate this with specific examples. Figure 3 This is an example diagram showing the starting coordinates of the audio segment to be trimmed in the first audio segment and the coordinates of the non-aligned point in the second audio segment, provided by an embodiment of the present invention. Figure 3 The shaded area represents the additional audio content in the first audio segment compared to the second audio segment, corresponding to... Figure 1 The audio segment k is shown.

[0077] like Figure 3As shown, the first sub-audio segment A in the first audio segment and the second sub-audio segment A' in the second audio segment are matched first. If the match is successful, the first sub-audio segment B in the first audio segment and the second sub-audio segment B' in the second audio segment are matched. If the match is successful, the first sub-audio segment C in the first audio segment and the second sub-audio segment C' in the second audio segment are matched. If the match fails, the starting coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the first audio segment C, i.e., x11 in the figure. The coordinates of the non-aligned point in the second audio segment are determined based on the coordinates of the audio frames in the second sub-audio segment, i.e., x1 in the figure.

[0078] In step S104, after determining the starting coordinates of the audio segment to be cut in the first audio segment, it is also necessary to determine the ending coordinates of the audio segment to be cut in the first audio segment to achieve the positioning of the audio segment to be cut.

[0079] After determining the starting coordinates of the audio segment to be trimmed in the first audio segment and the coordinates of the non-alignment point in the second audio segment in step S103, a third sub-audio segment can be determined in the second audio segment after the non-alignment point based on the number of second audio frames. That is, the audio frames after the number of second audio frames after the non-alignment point are taken as a third sub-audio segment. Similarly, one or more fourth sub-audio segments can be determined sequentially in the first audio segment after the starting coordinates. For example, if the number of second audio frames is set to 3 and the non-alignment point coordinate is the 11th frame, then frames 11-13 in the second audio segment are a third sub-audio segment. If the starting coordinate is the 11th frame, then frames 11-13 in the first audio segment are the first fourth sub-audio segment, frames 14-16 are the second fourth sub-audio segment, and so on.

[0080] Based on this, the third sub-audio segment is sequentially matched with one or more fourth sub-audio segments until a fourth sub-audio segment is found that successfully matches the third sub-audio segment. The termination coordinates of the audio segment to be trimmed within the first audio segment are then determined based on the coordinates of the audio frames within the fourth sub-audio segment. As an example, the termination coordinates of the audio segment to be trimmed within the first audio segment can be determined based on the coordinates of the first audio frame within the fourth sub-audio segment.

[0081] The number of the second audio frames can be configured according to actual needs. As an example, to improve the accuracy of the positioning results, the number of the second audio frames can be set to a value less than the number of the first audio frames; for example, the number of the second audio frames can be 3 frames or 1 frame.

[0082] As an example, the second audio frame can be 1 frame. Then, the first audio frame after the non-alignment point in the second audio segment is the third sub-audio segment, the first audio frame after the starting coordinate in the first audio segment is the first fourth sub-audio segment, the second audio frame is the second fourth sub-audio segment, and so on.

[0083] The following is combined with Figure 4 This example is explained below. Figure 4 This is an example diagram illustrating the start and end coordinates of the audio segment to be trimmed, as well as the coordinates of the non-aligned points, provided in an embodiment of the present invention. Figure 4 and Figure 2 Correspondingly, specifically, Figure 4 The shaded area represents the additional audio content in the first audio segment compared to the second audio segment, corresponding to... Figure 1 The audio segment k is shown.

[0084] Figure 4 The diagram shows the starting coordinate x11 of the audio segment to be trimmed in the first audio segment and the non-aligned point coordinate x1 in the second audio segment, determined in step S103. Using the number of second audio frames as the segment length, a third sub-audio segment a' can be determined in the second audio segment after the non-aligned point coordinate x1, and the first fourth sub-audio segment a can be determined in the first audio segment after the starting coordinate x11 of the audio segment to be trimmed. The third sub-audio segment a' and the fourth sub-audio segment a are then matched. If the match fails, the matching of the third sub-audio segment a' and the fourth sub-audio segment b continues. If the match fails, the matching of the third sub-audio segment a' and the fourth sub-audio segment c continues. If the match succeeds, the ending coordinate of the audio segment to be trimmed in the first audio segment is determined based on the coordinates of the audio frames in the fourth sub-audio segment b. For example, the ending coordinate x12 of the audio segment to be trimmed in the first audio segment can be determined based on the coordinates of the first audio frame in the fourth sub-audio segment b.

[0085] Figure 4 In the example shown, if the second audio frame is 1 frame, then the third sub-audio segment a' can be understood as... Figure 2 The first audio frame in audio segment n' of the second audio segment shown, and the fourth sub-audio segment a can be understood as the first audio frame in audio segment k of the first audio segment.

[0086] The fourth sub-audio segment c can be understood as Figure 2 The first audio frame in audio segment n of the first audio segment shown is used to determine the first audio frame. Therefore, the third sub-audio segment a' and the fourth sub-audio segment c can be matched successfully. Given that the third sub-audio segment a' and the fourth sub-audio segment c are matched successfully, the termination coordinate x12 of the audio segment to be cut in the first audio segment can be determined based on the coordinates of the audio frame in the fourth sub-audio segment c, thus achieving the positioning of the audio segment to be cut.

[0087] Similar to the matching process for the first and second sub-audio segments, a successful match between the third and fourth sub-audio segments can be considered if the similarity between the audio fingerprints corresponding to the third and fourth sub-audio segments meets a preset condition. This embodiment of the invention does not limit the preset condition. As an example, based on a pre-set similarity threshold between the third and fourth sub-audio segments, if the similarity between the audio fingerprint corresponding to the third sub-audio segment and the audio fingerprint corresponding to a certain fourth sub-audio segment is not less than the similarity threshold, then the third sub-audio segment can be considered to have successfully matched the fourth sub-audio segment.

[0088] In this embodiment of the invention, if the third sub-audio segment and the fourth sub-audio segment are sub-audio segments corresponding to different audio content, then the similarity between the audio fingerprints corresponding to the third sub-audio segment and the fourth sub-audio segment is low and does not exceed a predetermined threshold, that is, the third sub-audio segment and the fourth sub-audio segment cannot be matched successfully.

[0089] If the third and fourth sub-audio segments correspond to the same audio content and the only difference is whether the sub-audio segments are standard versions, then the similarity between the audio fingerprints corresponding to the third and fourth sub-audio segments is high and exceeds a predetermined threshold, meaning that the third and fourth sub-audio segments can be successfully matched.

[0090] In this embodiment of the invention, after obtaining the start and end coordinates of the audio segment to be trimmed in the first audio segment based on the above steps S101-S104, it is considered that the start and end coordinates are specifically the start and end coordinates of the audio segment that is additional to the first audio segment compared to the second audio segment, or the start and end coordinates of the audio segment that is different between the two. See [link to relevant documentation]. Figure 1 For example, it is assumed that the located audio segment to be trimmed is... Figure 1 The audio segment k in the text is the part that needs to be cut.

[0091] Therefore, in step S105, by cutting out the audio segment to be cut from the first audio segment, an audio segment that is compatible with the incomplete video and meets the audio quality requirements can be obtained, or in other words, a standard audio segment.

[0092] In step S105, the position range of the audio segment to be cut can be determined based on the start and end coordinates of the audio segment to be cut in the first audio segment, and the first audio segment can be cut based on the position range.

[0093] Specifically, the audio frames between the start and end coordinates in the first audio segment can be considered as the content of the audio segment to be trimmed.

[0094] As an example, when using audio frame numbers to represent coordinates, if the starting coordinate of the audio segment to be cropped in the first audio segment is x11 and the ending coordinate is x12, then the position range of the audio segment to be cropped in the first audio segment is [x11, x12). By cropping the audio frames within [x11, x12) of the first audio segment, a standard audio segment adapted to the incomplete video can be obtained. Specifically, since the ending coordinate x12 of the audio segment to be cropped is determined based on the coordinates of the audio frames in the fourth sub-audio segment that successfully matches the third sub-audio segment, and the audio frames in the fourth sub-audio segment do not belong to the content of the audio segment to be cropped when the fourth sub-audio segment and the third sub-audio segment successfully match, this endpoint coordinate is not included in the position range of the audio segment to be cropped.

[0095] For example, if the starting coordinate of the audio segment to be cropped in the first audio segment is the 100th frame and the ending coordinate is the 200th frame, then the position range of the audio segment to be cropped in the first audio segment is [100, 200), which means that the content of the 100th to 199th frames in the first audio segment needs to be cropped to obtain a standard audio segment that is compatible with the incomplete version of the video.

[0096] The audio trimming method provided in this invention involves sequentially determining one or more first sub-audio segments in a first audio segment, determining one or more second sub-audio segments in a second audio segment, and sequentially matching the first and second sub-audio segments in the same order. Based on the coordinates of audio frames in the unmatched first sub-audio segments, the starting coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the coordinates of audio frames in the unmatched second sub-audio segments, the coordinates of the misalignment points within the second audio segment are determined. A third sub-audio segment is determined within the second audio segment following the misalignment point coordinates. One or more fourth sub-audio segments are determined within the first audio segment following the starting coordinates. The third and fourth sub-audio segments are sequentially matched. Based on the coordinates of audio frames in the successfully matched fourth sub-audio segments, the ending coordinates of the audio segment to be trimmed within the first audio segment are determined. Based on the starting and ending coordinates of the audio segment to be trimmed within the first audio segment, the first audio segment is trimmed to obtain a standard audio segment adapted to the incomplete video. The audio trimming method provided in this invention locates the audio segment to be trimmed by matching sub-audio segments in the first and second audio segments. This eliminates the need for manual comparison between the first and second audio segments, enabling the location and trimming of audio content that is additional to the first audio segment compared to the second audio segment. This yields a standard audio segment that is consistent with the incomplete audio video, improving the efficiency of locating and trimming differing audio content. When applied to the restoration of international audio, this method can enhance restoration efficiency and reduce restoration costs.

[0097] In one embodiment of the present invention, the complete video is the submitted version video, the incomplete video is the approved version video after being edited from the submitted version video, the standard audio segment is the standard international audio corresponding to the complete video, and the non-standard audio segment is the international audio separated from the approved video.

[0098] In this embodiment of the invention, the standard international audio can be provided by the film and television work supplier along with the original video, i.e., the version submitted for review. By separating human voices from background sounds in the audio corresponding to the approved video medium, the separated background sounds are the non-standard audio segments.

[0099] Therefore, by processing the standard international audio corresponding to the submitted video and the international audio separated from the approved video in steps S101-S104, the location range of the international audio corresponding to the deleted video segment in the standard international audio can be determined. Based on this, the original standard international audio can be trimmed to obtain a standard international audio that matches the approved video. Specifically, the process of processing the original international audio provided by the supplier to obtain an international audio that matches the approved video is usually referred to as international audio restoration.

[0100] Therefore, this embodiment of the invention eliminates the need for manual comparison of the audio corresponding to the submitted and approved versions of the video. It can automatically locate and trim the audio segments to be trimmed included in the standard version of the approved video, obtain standard international audio that matches the approved video, and complete the international audio restoration. This can improve the efficiency of international audio restoration and reduce restoration costs.

[0101] In one embodiment of the present invention, it can be determined whether the first sub-audio segment and the second sub-audio segment are successfully matched based on the following method:

[0102] Cross-correlation calculation is performed on the audio fingerprints corresponding to the first sub-audio segment and the second sub-audio segment to obtain a similarity sequence. It is then determined whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence.

[0103] If so, confirm that the first and second sub-audio segments have matched successfully;

[0104] If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

[0105] As mentioned above, each first sub-audio segment and each second sub-audio segment may specifically include the number of first audio frames. If the number of first audio frames is N, then a cross-correlation operation is performed on the audio fingerprint corresponding to a first sub-audio segment and the audio fingerprint corresponding to a second sub-audio segment. Specifically, a similarity sequence with 2N-1 items can be obtained. If the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence, that is, the peak of the similarity sequence is the Nth item, then the first sub-audio segment and the second sub-audio segment are considered to be successfully matched. Otherwise, the first sub-audio segment and the second sub-audio segment are not successfully matched.

[0106] For details on how to perform cross-correlation operations on the audio features corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment, please refer to the relevant technical content. The following is a brief explanation of the cross-correlation operation with specific examples.

[0107] Figure 5 This is an example diagram of the cross-correlation calculation process provided in the embodiments of the present invention. For ease of description, it is specifically taken as an example that the number of the first audio frames is 3. Figure 5 As shown, the three audio frames in the first sub-audio segment are labeled 0, 1, and 2, and their corresponding audio fingerprints are denoted as x0, x1, and x2, respectively. The three audio frames in the second audio segment are labeled 3, 4, and 5, and their corresponding audio features are denoted as x3, x4, and x5, respectively. Figure 5 The audio fingerprints corresponding to the first and second sub-audio segments are cross-correlated. Specifically, this can be understood as using the first sub-audio segment as a reference, sliding the second sub-audio segment at the position of the first sub-audio segment, and calculating the similarity between the audio fingerprints corresponding to the second sub-audio segment and the audio fingerprints corresponding to the first sub-audio segment for the audio frames that slide from the second sub-audio segment to the position of the first sub-audio segment. In order to ensure that the data length of the audio features is consistent when calculating the similarity, zeros can be padded at the positions of missing data.

[0108] See Figure 5 In (a), audio frame 5 in the second sub-audio segment slides to the corresponding position of audio frame 0. At this point, the similarity between audio fingerprints (x0,x1,x2) and (x5,0,0) can be calculated. Similarly, (b)-(e) can be calculated, resulting in a total of 5 similarities. If the similarities of (a)-(e) are 0.2, 0.2, 0.9, 0.3, and 0.1 respectively, then the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence, and the first and second sub-audio segments are considered to have matched successfully.

[0109] As mentioned earlier, the first sub-audio segment in the first audio segment is a non-standard version, while the second sub-audio segment in the second audio segment is a standard version. Therefore, if matching is based solely on the similarity between the audio fingerprints corresponding to the first and second sub-audio segments, the accuracy of the matching results may be limited. In this embodiment of the invention, a similarity sequence is obtained by performing cross-correlation on the audio fingerprints corresponding to the first and second sub-audio segments. The coordinates of the similarity peak in the similarity sequence are then determined to be the center coordinates of the similarity sequence. If so, i.e., the first and second sub-audio segments are aligned, the first and second sub-audio segments are considered to have matched successfully, thus improving the accuracy of matching the first and second sub-audio segments.

[0110] In one embodiment of the present invention, the aforementioned steps of determining the starting coordinates of the audio segment to be trimmed in the first audio segment based on the coordinates of the audio frames in the i-th first sub-audio segment, and determining the coordinates of the non-aligned point in the second audio segment based on the position coordinates of the i-th second sub-audio segment, may specifically include:

[0111] Using the number of the third audio frames as the shortening step, the i-th first sub-audio segment and the i-th second sub-audio segment are shortened simultaneously, and the shortened first sub-audio segment and the shortened second sub-audio segment are matched until the shortened first sub-audio segment and the shortened second sub-audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the successfully matched i-th second sub-audio segment are used as the coordinates of the non-aligned point in the second audio segment.

[0112] The number of third audio frames can be predetermined and is less than the number of first audio frames. For example, if the number of first audio frames is 10, the number of third audio frames can be 3 or 1.

[0113] In this embodiment of the invention, the i-th first sub-audio segment and the i-th second sub-audio segment are shortened simultaneously, and the shortened first sub-audio segment and the shortened second sub-audio segment are matched until the shortened first sub-audio segment and the second sub-audio segment are successfully matched. Specifically, for the first sub-audio segment and the second sub-audio segment that are not successfully matched, the number of third audio frames of both the first sub-audio segment and the second sub-audio segment are shortened simultaneously, and the shortened first sub-audio segment and the second sub-audio segment are matched. If the match is unsuccessful, the number of third audio frames of both the first sub-audio segment and the second sub-audio segment are shortened again and matched, and so on, until the shortened first sub-audio segment and the second sub-audio segment are successfully matched.

[0114] As an example, if the third audio frame has 2 frames, then the 11th-20th audio frames in the first audio segment constitute the second first sub-audio segment, and the 11th-20th audio frames in the second audio segment constitute the second second sub-audio segment, and the second first sub-audio segment and the second second sub-audio segment do not match.

[0115] Then, the 11th to 18th audio frames in the first audio segment are taken as the shortened first sub-audio segment, and the 11th to 18th audio frames in the second audio segment are taken as the shortened second sub-audio segment. The shortened first sub-audio segment and the second sub-audio segment are matched. If the match fails, the first sub-audio segment and the second sub-audio segment are shortened again, and the 11th to 16th audio frames in the first audio segment and the 11th to 16th audio frames in the second audio segment are matched again. This process is repeated until the shortened first sub-audio segment and the second sub-audio segment are matched successfully.

[0116] After the shortened first and second sub-audio segments are successfully matched, the coordinates of the last audio frame in the i-th successfully matched first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the i-th successfully matched second sub-audio segment are used as the coordinates of the non-aligned point in the second audio segment.

[0117] As an example, if the shortened first sub-audio segment is frames 11-16, the shortened second sub-audio segment is frames 11-16, and the first and second sub-audio segments are successfully matched, then the starting coordinate of the audio segment to be trimmed in the first audio segment is frame 16, and the coordinate of the non-aligned point in the second audio segment is frame 16.

[0118] It should be noted that in this embodiment of the invention, if the starting coordinates of the audio segment to be trimmed in the first audio segment are x11, since x11 is specifically determined based on the coordinates of the audio frames in the first sub-audio segment that are successfully matched with the second sub-audio segment, the audio frames in the first sub-audio segment do not belong to the content of the audio segment to be trimmed when the second sub-audio segment and the first sub-audio segment are successfully matched. Therefore, when determining the position range of the audio segment to be trimmed in order to trim the first audio segment, the endpoint coordinates are not included in the position range of the audio segment to be trimmed. That is, if the ending coordinate of the audio segment to be trimmed in the first audio segment is x12, then the position range of the audio segment to be trimmed in the first audio segment is (x11, x12).

[0119] The following will illustrate this with specific examples. Figure 6 This is an example diagram illustrating the matching process of the first and second sub-audio segments provided in an embodiment of the present invention. Figure 6 The shaded area represents the additional audio content in the first audio segment compared to the second audio segment. Figure 6(a) shows the first sub-audio segments A, B, C, and the second sub-audio segments A', B', C'. The first sub-audio segments A and A', and the first sub-audio segments B and B' are all successfully matched. If the first sub-audio segments C and C' are not successfully matched, the first sub-audio segments C and C' are shortened simultaneously until the shortened first sub-audio segments and C' are successfully matched. (b) shows D as the shortened first sub-audio segment obtained by shortening the first sub-audio segment C, and D' as the shortened second sub-audio segment obtained by shortening the second sub-audio segment D. The shortened first sub-audio segment D and the shortened second sub-audio segment D' are successfully matched. Then, the coordinates of the last audio frame in the first sub-audio segment D are taken as the starting coordinates x11 of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the second sub-audio segment D' are taken as the non-alignment point coordinates x1 in the second audio segment.

[0120] In practical applications, when matching the first sub-audio segment and the second sub-audio segment using the number of first audio frames as the segment length, the starting position of the audio segment to be trimmed in the first audio segment has a certain probability of being located in the middle of the first sub-audio segment. If the first and second sub-audio segments fail to match, it is difficult to accurately determine the starting coordinates of the audio segment to be trimmed in the first audio segment and the coordinates of the misaligned point in the second audio segment. In this embodiment of the invention, when the first and second sub-audio segments fail to match, the unmatched first and second sub-audio segments are shortened simultaneously until the shortened first and second sub-audio segments match successfully. The coordinates of the last audio frame in the shortened first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the shortened second sub-audio segment are used as the coordinates of the misaligned point in the second audio segment, thus improving the accuracy of locating the audio segment to be trimmed.

[0121] In one embodiment of the present invention, the audio segment to be trimmed may specifically be multiple audio segments distributed within the second audio segment. Figure 7 This is another example diagram of the first and second audio segments provided in this embodiment of the invention, as shown below. Figure 7 As shown, the first audio segment includes standard audio segments m, n, o, as well as k and p, while the second audio segment only includes non-standard audio segments m', n', o'.

[0122] Therefore, in this embodiment, the audio trimming method provided by the present invention further includes:

[0123] The process involves determining the first audio segment after the termination coordinate as the new first audio segment, determining the second audio segment after the non-aligned point coordinate as the new second audio segment, returning to the steps of determining one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment sequentially, using the number of first audio frames as the segment length, and sequentially matching the first and second sub-audio segments in the same order. This process yields the start and end coordinates of the audio segment to be trimmed determined for the new first audio segment. The process then returns to the steps of determining the first audio segment after the current termination coordinate as the new first audio segment and the second audio segment after the current non-aligned point coordinate as the new second audio segment, until both the first and second sub-audio segments are successfully matched.

[0124] Specifically, after locating the start and end coordinates of an audio segment to be trimmed within the first audio segment, it is necessary to match the subsequent audio content in the first and second audio segments to locate other audio segments to be trimmed that may be included in the first audio segment. Therefore, the first audio segment after the end coordinates is determined as the new first audio segment, and the second audio segment after the non-alignment point coordinates is determined as the new second audio segment. The first sub-audio segment in the newly determined first audio segment and the second sub-audio segment in the second audio segment are then matched. The specific matching process can be found in the explanation of steps S103-S104 above. After locating the second audio segment to be trimmed within the first audio segment, the location of subsequent audio segments to be trimmed is performed in the same way, until all audio content in the first and second audio segments is matched.

[0125] The following is combined with Figure 7The embodiments of the present invention are illustrated by way of example. A1 and B1 are the original first audio segment and the second audio segment, respectively. Based on steps S103-S104, the first sub-audio segment in A1 and the second sub-audio segment in B1 are matched to determine the starting coordinates x11 and ending coordinates x12 of the first audio segment to be trimmed in the first audio segment, thus achieving the positioning of the first audio segment to be trimmed k. After positioning k, the first audio segment A2 after the ending coordinate x12 is determined as the new first audio segment, and the second audio segment B2 after the non-alignment point coordinate x1 is determined as the new second audio segment. Based on steps S103-S104, a new round of matching is performed on the first sub-audio segment in A2 and the second sub-audio segment in B2 to determine the starting coordinates x21 and ending coordinates x22 of the second audio segment to be trimmed in the first audio segment, thus achieving the positioning of the second audio segment to be trimmed p. After locating p, the first audio segment A3 after the termination coordinate x22 is determined as the new first audio segment, and the second audio segment B3 after the non-alignment point coordinate x2 is determined as the new second audio segment. A new round of matching is performed on the first sub-audio segment in A3 and the second sub-audio segment in B3. If the subsequent first audio segment includes other audio segments to be clipped, the steps for locating these audio segments to be clipped are similar.

[0126] Based on the embodiments of the present invention, the starting coordinates and ending coordinates of multiple audio segments to be trimmed in the first audio segment can be obtained. For example, the starting coordinate of the i-th audio segment to be trimmed in the first audio segment is xi1, and the ending coordinate is xi2.

[0127] In this embodiment of the invention, by determining the first audio segment after the termination coordinate as the new first audio segment and the second audio segment after the non-alignment point coordinate as the new second audio segment, it is possible to locate multiple audio segments to be trimmed included in the first audio segment, thus having better applicability.

[0128] In one embodiment of the present invention, the aforementioned step of trimming the first audio segment based on the start and end coordinates of the audio segment to be trimmed within the first audio segment specifically includes:

[0129] For each identified audio segment to be trimmed, the first audio segment is trimmed based on the start and end coordinates corresponding to that audio segment.

[0130] In this embodiment of the invention, after determining the start coordinates and end coordinates corresponding to an audio segment to be trimmed, the start coordinates and end coordinates corresponding to the audio segment to be trimmed can be recorded first. After matching all audio content in the first audio segment and the second audio segment, the first audio segment is trimmed based on the start coordinates and end coordinates corresponding to each audio segment to be trimmed, that is, the audio content between xi1 and xi2 is trimmed.

[0131] In this embodiment of the invention, by determining the first audio segment after the starting coordinate as the new first audio segment, and the second audio segment after the ending coordinate as the new second audio segment, and by cropping the first audio segment based on the starting and ending coordinates of each audio segment to be cropped, the applicability of the audio cropping method is improved.

[0132] Figure 8 This is a schematic diagram of an audio trimming method provided in an embodiment of the present invention. The following is in conjunction with... Figure 8 The audio trimming method provided in the embodiments of the present invention will be further described.

[0133] like Figure 8 As shown, the background audio of the approved video medium is first separated to obtain the separated international sound MnE_2, which is the second audio segment, and the original international sound MnE_1, which is the first audio segment, is obtained. Then, audio fingerprints are extracted from the separated international audio and the original international audio, respectively, to obtain the audio fingerprint sequence FP1 corresponding to the first audio segment and the audio fingerprint sequence FP2 corresponding to the second audio segment. The audio fingerprint FP1_sub_i corresponding to the first sub-audio segment in the first audio segment and the audio fingerprint FP2_sub_i corresponding to the second sub-audio segment in the second audio segment are matched. If the match is successful, the matching of FP1_sub_i+1 and FP2_sub_i+1 continues. If the match is unsuccessful, the matching of FP1_sub_i corresponding to the shortened first sub-audio segment and FP2_sub_i corresponding to the shortened second sub-audio segment continues until a match is successful. Based on the audio coordinates of the audio frame corresponding to FP1_sub_i+1, the starting coordinate xi1 of the audio segment to be clipped in the first audio segment is obtained. Then, the third sub-audio segment after the non-aligned point coordinate xi in the second audio segment and one or more fourth sub-audio segments after xi2 in the first audio segment are matched to obtain the ending coordinate xi2 of the audio segment to be clipped in the first audio segment. This process can also be understood as a sliding match between the first and second audio segments. After matching is completed, the original international audio can be trimmed based on the obtained xi1 and x12 to obtain the trimmed international audio MnE_out.

[0134] In this embodiment of the invention, background audio separation is performed on the approved video medium to obtain the separated international sound MnE_2. Audio fingerprints are extracted from MnE_2 and the original international sound MnE_1. Then, sliding matching is performed on the audio fingerprint FP1_sub_i corresponding to the first sub-audio segment and the audio fingerprint FP2_sub_i corresponding to the second sub-audio segment. The starting coordinates of the audio segment to be cropped are determined based on the coordinates of the audio frames in the first sub-audio segment that failed to match. Then, the third sub-audio segment after the non-aligned point coordinates in the second audio segment and one or more fourth sub-audio segments after the starting coordinates in the first audio segment are matched to obtain the ending coordinates of the audio segment to be cropped. The original international sound is cropped based on the starting and ending coordinates of multiple audio segments to be cropped to obtain the international sound of the approved video. This can improve the efficiency of international sound restoration and reduce the restoration cost.

[0135] Corresponding to the above method embodiments, this application also provides an audio trimming device, such as... Figure 9 As shown, the device includes:

[0136] The acquisition module 901 is used to acquire a first audio segment and a second audio segment; wherein, the first audio segment is the standard audio segment corresponding to the complete video, and the second audio segment is a non-standard audio segment separated from the incomplete video;

[0137] The extraction module 902 is used to extract audio fingerprints from the first audio segment and the second audio segment respectively, and obtain the audio fingerprint of each audio frame in the first audio segment and the second audio segment respectively.

[0138] The first matching module 903 is used to determine one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment, using the number of first audio frames as the segment length, and to match the first and second sub-audio segments in the same order in sequence, until the i-th first sub-audio segment and the i-th second sub-audio segment fail to match. Based on the coordinates of the audio frames in the i-th first sub-audio segment, the starting coordinates of the audio segment to be trimmed in the first audio segment are determined, and the coordinates of the non-aligned point in the second audio segment are determined based on the coordinates of the audio frames in the i-th second sub-audio segment. A successful match between the first and second sub-audio segments indicates that the similarity between the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment meets a preset condition.

[0139] The second matching module 904 is used to determine a third sub-audio segment in the second audio segment located after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to determine one or more fourth sub-audio segments in the first audio segment located after the start coordinates. The third sub-audio segment and one or more fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinate of the audio segment to be trimmed in the first audio segment is determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment.

[0140] The cropping module 905 is used to crop the first audio segment based on the start and end coordinates of the audio segment to be cropped in the first audio segment, so as to obtain a standard audio segment that is compatible with the incomplete video.

[0141] The audio trimming device provided in this embodiment of the invention determines one or more first sub-audio segments in a first audio segment, and one or more second sub-audio segments in a second audio segment in turn. It then sequentially matches the first and second sub-audio segments in the same order. Based on the coordinates of audio frames in the unmatched first sub-audio segments, it determines the starting coordinates of the audio segment to be trimmed within the first audio segment. Based on the coordinates of audio frames in the unmatched second sub-audio segments, it determines the coordinates of the misaligned points in the second audio segment. A third sub-audio segment is then determined in the second audio segment following the misaligned point coordinates. One or more fourth sub-audio segments are determined in the first audio segment following the starting coordinates. The third and fourth sub-audio segments are then matched sequentially. Based on the coordinates of audio frames in the successfully matched fourth sub-audio segments, it determines the ending coordinates of the audio segment to be trimmed within the first audio segment. Finally, based on the starting and ending coordinates of the audio segment to be trimmed within the first audio segment, the first audio segment is trimmed to obtain a standard audio segment adapted to the incomplete video. The audio trimming method provided in this invention locates the audio segment to be trimmed by matching sub-audio segments in the first and second audio segments. This eliminates the need for manual comparison between the first and second audio segments, enabling the location and trimming of additional audio content in the first audio segment compared to the second audio segment. This yields a standard audio segment similar to that incomplete audio-visual content, improving the efficiency of locating and trimming additional audio content. When applied to the restoration of international audio, this method enhances restoration efficiency and reduces restoration costs.

[0142] In one embodiment of the present invention, the first matching module 903 includes:

[0143] The judgment unit is used to perform cross-correlation calculation on the audio fingerprints corresponding to the first sub-audio segment and the second sub-audio segment to obtain a similarity sequence, and to determine whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence.

[0144] If so, confirm that the first and second sub-audio segments have matched successfully;

[0145] If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

[0146] In one embodiment of the present invention, the first matching module 903 includes:

[0147] The matching unit is used to synchronously shorten the i-th first sub-audio segment and the i-th second sub-audio segment with the number of third audio frames as the shortening step, and to match the shortened first sub-audio segment and the shortened second sub-audio segment until the shortened first sub-audio segment and the shortened second sub-audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the successfully matched i-th second sub-audio segment are used as the coordinates of the non-alignment point in the second audio segment.

[0148] In one embodiment of the present invention, the device further includes:

[0149] The return module is used to determine the first audio segment after the termination coordinate as the new first audio segment, and the second audio segment after the non-aligned point coordinate as the new second audio segment. It returns the steps of determining one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment in sequence, using the number of first audio frames as the segment length, and matching the first and second sub-audio segments in the same order in sequence. It obtains the start and end coordinates of the audio segment to be trimmed for the new first audio segment, and returns the steps of determining the first audio segment after the current termination coordinate as the new first audio segment and the second audio segment after the current non-aligned point coordinate as the new second audio segment, until both the first and second sub-audio segments are successfully matched.

[0150] In one embodiment of the present invention, the trimming module 905 is specifically used for:

[0151] For each identified audio segment to be trimmed, the first audio segment is trimmed based on the start and end coordinates corresponding to that audio segment.

[0152] In one embodiment of the present invention, the complete video is the submitted version video, the incomplete video is the approved version video after being edited from the submitted version video, the standard audio segment is the standard international audio corresponding to the complete video, and the non-standard audio segment is the international audio separated from the approved video.

[0153] This invention also provides an electronic device, such as... Figure 10As shown, it includes a processor 101, a communication interface 102, a memory 103, and a communication bus 104, wherein the processor 101, the communication interface 102, and the memory 103 communicate with each other through the communication bus 104.

[0154] Memory 103 is used to store computer programs;

[0155] When processor 101 executes a program stored in memory 103, it performs the following steps:

[0156] Obtain the first audio segment and the second audio segment; wherein, the first audio segment is the standard audio segment corresponding to the complete video, and the second audio segment is the non-standard audio segment separated from the incomplete video.

[0157] Audio fingerprints are extracted from the first audio segment and the second audio segment respectively, and the audio fingerprints of each audio frame in the first audio segment and the second audio segment are obtained respectively.

[0158] Using the number of first audio frames as the segment length, one or more first sub-audio segments are sequentially determined in the first audio segment, and one or more second sub-audio segments are sequentially determined in the second audio segment. The first and second sub-audio segments in the same order are matched sequentially until the i-th first sub-audio segment and the i-th second sub-audio segment fail to match. The starting coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the i-th first sub-audio segment, and the coordinates of the non-alignment point in the second audio segment are determined based on the coordinates of the audio frames in the i-th second sub-audio segment. A successful match between the first and second sub-audio segments indicates that the similarity between the audio fingerprints corresponding to the first and second sub-audio segments meets a preset condition.

[0159] Using the number of second audio frames as the segment length, a third sub-audio segment is determined in the second audio segment located after the non-alignment point coordinates, and one or more fourth sub-audio segments are determined in the first audio segment located after the starting coordinates. The third sub-audio segment and one or more fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment.

[0160] Based on the start and end coordinates of the audio segment to be cropped within the first audio segment, the first audio segment is cropped to obtain a standard audio segment that is compatible with the incomplete video.

[0161] The communication interface is used for communication between the aforementioned terminal and other devices.

[0162] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0163] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0164] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the audio trimming methods described in the above embodiments.

[0165] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the audio trimming methods described in the above embodiments.

[0166] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0167] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0168] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0169] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An audio trimming method, characterized in that, include: Obtain a first audio segment and a second audio segment; wherein, the first audio segment is the standard audio segment corresponding to the complete version of the video, and the second audio segment is a non-standard audio segment separated from the incomplete version of the video; the complete version of the video is the submitted version of the video, and the incomplete version of the video is the approved version of the submitted version of the video after being edited; the standard audio segment is the standard international audio corresponding to the complete version of the video, and the non-standard audio segment is the international audio separated from the approved version of the video; wherein, international audio refers to all sounds in a film or television work except for dialogue in the language of the country of production of the film or television work; Audio fingerprints are extracted from the first audio segment and the second audio segment respectively, and the audio fingerprints of each audio frame in the first audio segment and the second audio segment are obtained respectively. Using the number of first audio frames as the segment length, one or more first sub-audio segments are sequentially determined in the first audio segment, and one or more second sub-audio segments are sequentially determined in the second audio segment. The first sub-audio segments and second sub-audio segments in the same order are matched sequentially until the i-th first sub-audio segment and the i-th second sub-audio segment fail to match. The starting coordinates of the audio segment to be clipped in the first audio segment are determined based on the coordinates of the audio frames in the i-th first sub-audio segment, and the coordinates of the non-aligned points in the second audio segment are determined based on the coordinates of the audio frames in the i-th second sub-audio segment. A successful match between the first sub-audio segment and the second audio segment indicates that the similarity between the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment meets a preset condition. Using the number of second audio frames as the segment length, a third sub-audio segment is determined in the second audio segment located after the non-alignment point coordinates, and one or more fourth sub-audio segments are determined in the first audio segment located after the starting coordinates. The third sub-audio segment and one or more of the fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment. The number of second audio frames is less than the number of first audio frames. Based on the start and end coordinates of the audio segment to be cropped in the first audio segment, the first audio segment is cropped to obtain a standard audio segment that is compatible with the incomplete video. The first audio segment after the termination coordinate is determined as the new first audio segment, and the second audio segment after the non-aligned point coordinate is determined as the new second audio segment. The process returns to the steps of using the number of first audio frames as the segment length, sequentially determining one or more first sub-audio segments in the first audio segment, sequentially determining one or more second sub-audio segments in the second audio segment, and sequentially matching the first and second sub-audio segments in the same order. This process obtains the start and end coordinates of the audio segment to be trimmed determined for the new first audio segment. The process returns to the steps of determining the first audio segment after the current termination coordinate as the new first audio segment and the second audio segment after the current non-aligned point coordinate as the new second audio segment, until both the first and second sub-audio segments are successfully matched.

2. The method according to claim 1, characterized in that, The following method is used to determine whether the first sub-audio segment and the second sub-audio segment are successfully matched: Cross-correlation calculation is performed on the audio fingerprints corresponding to the first sub-audio segment and the second sub-audio segment to obtain a similarity sequence. It is then determined whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence. If so, confirm that the first sub-audio segment and the second sub-audio segment have successfully matched; If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

3. The method according to claim 1, characterized in that, The steps of determining the starting coordinates of the audio segment to be trimmed within the first audio segment based on the coordinates of the audio frames in the i-th first sub-audio segment, and determining the coordinates of the non-aligned point in the second audio segment based on the position coordinates of the i-th second sub-audio segment, include: Using the number of the third audio frames as the shortening step, the i-th first sub-audio segment and the i-th second sub-audio segment are shortened simultaneously, and the shortened first sub-audio segment and the shortened second sub-audio segment are matched until the shortened first sub-audio segment and the shortened second sub-audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first sub-audio segment are used as the starting coordinates of the audio segment to be trimmed in the first audio segment, and the coordinates of the last audio frame in the successfully matched i-th second sub-audio segment are used as the coordinates of the non-aligned point in the second audio segment.

4. The method according to claim 1, characterized in that, The step of cropping the first audio segment based on the start and end coordinates of the audio segment to be cropped within the first audio segment includes: For each identified audio segment to be trimmed, the first audio segment is trimmed based on the start and end coordinates corresponding to that audio segment.

5. An audio trimming device, characterized in that, include: The acquisition module is used to acquire a first audio segment and a second audio segment; wherein, the first audio segment is the standard audio segment corresponding to the complete version of the video, and the second audio segment is a non-standard audio segment separated from the incomplete version of the video; the complete version of the video is the submitted version of the video, and the incomplete version of the video is the approved version of the submitted version of the video after being edited; the standard audio segment is the standard international audio corresponding to the complete version of the video, and the non-standard audio segment is the international audio separated from the approved version of the video; wherein, international audio refers to all sounds in a film or television work except for dialogue in the language of the country of production of the film or television work; The extraction module is used to extract audio fingerprints from the first audio segment and the second audio segment respectively, and obtain the audio fingerprint of each audio frame in the first audio segment and the second audio segment respectively; The first matching module is used to determine one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment, using the number of first audio frames as the segment length. It then matches the first and second sub-audio segments in the same order until the i-th first and second sub-audio segments fail to match. Based on the coordinates of the audio frames in the i-th first sub-audio segment, it determines the starting coordinates of the audio segment to be trimmed within the first audio segment. Based on the coordinates of the audio frames in the i-th second sub-audio segment, it determines the coordinates of the non-aligned points in the second audio segment. Successful matching of the first and second sub-audio segments indicates that the similarity between the audio fingerprints corresponding to the first and second sub-audio segments meets a preset condition. The second matching module is used to determine a third sub-audio segment in the second audio segment located after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to determine one or more fourth sub-audio segments in the first audio segment located after the starting coordinates. The third sub-audio segment and one or more fourth sub-audio segments are matched sequentially until a fourth sub-audio segment that successfully matches the third sub-audio segment is determined. The termination coordinates of the audio segment to be trimmed in the first audio segment are determined based on the coordinates of the audio frames in the successfully matched fourth sub-audio segment. The number of second audio frames is less than the number of first audio frames. The cropping module is used to crop the first audio segment based on the start and end coordinates of the audio segment to be cropped in the first audio segment, and obtain a standard audio segment that is compatible with the incomplete video. The return module is used to determine the first audio segment after the termination coordinate as the new first audio segment, and the second audio segment after the non-alignment point coordinate as the new second audio segment. It returns to the steps of determining one or more first sub-audio segments in the first audio segment and one or more second sub-audio segments in the second audio segment sequentially, using the number of first audio frames as the segment length, and sequentially matching the first and second sub-audio segments in the same order. This process obtains the start and end coordinates of the audio segment to be trimmed determined for the new first audio segment. It then returns to the steps of determining the first audio segment after the current termination coordinate as the new first audio segment and the second audio segment after the current non-alignment point coordinate as the new second audio segment, until both the first and second sub-audio segments are successfully matched.

6. The apparatus according to claim 5, characterized in that, The first matching module includes: The judgment unit is used to perform cross-correlation calculation on the audio fingerprint corresponding to the first sub-audio segment and the audio fingerprint corresponding to the second sub-audio segment to obtain a similarity sequence, and to determine whether the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence. If so, confirm that the first sub-audio segment and the second sub-audio segment have successfully matched; If not, it indicates that the current first sub-audio segment and the current second sub-audio segment did not match successfully.

7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-4.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Data processing method, device and equipment

    CN114845164A

  • Audio difference positioning method and device

    CN116755034A

  • Differential detection apparatus, differential detection method and differential detection program

    JP2015141602A