Audio difference localization method and apparatus

CN116755034BActive Publication Date: 2026-08-21CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310741344.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-08-21
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

当这种方式需要进行肉眼比对,会耗费大量的人力和时间,效率低下且定位准确度不高

Benefits of technology

[0040]本发明实施例提供的音频差异定位方法,通过依次在完整版音频中确定一个或多个第一音频段,在非完整版音频中确定一个或多个第二音频段,并依次匹配相同次序的第一音频段和第二音频段,基于未匹配成功的第一音频段中音频帧的坐标确定差异音频段在完整版音频中的起始坐标,基于未匹配成功的第二音频段中音频帧的坐标确定非完整版音频中的非对齐点坐标,并在非对齐点坐标之后的完整版音频中确定一个第三音频段,在起始坐标之后的完整版音频中确定一个或多个第四音频段,将第三音频段和第四音频段进行依次匹配,基于匹配成功的第四音频段中音频帧的坐标确定差异音频段在完整版音频中的终止坐标,基于起始坐标和终止坐标定位差异音频段在完整版音频中的位置区间。应用本发明实施例提供的音频差异定位方法,通过对完整版音频和非完整版音频中的音频段进行匹配来定位差异音频段,无需对完整版音频和非完整版音频进行人工比对,就能够实现对非完整版音频中缺失的差异音频段进行定位,能够提升定位音频间的差异片段的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116755034B_ABST
    Figure CN116755034B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio difference positioning method and device, which sequentially match a first audio segment in a complete version audio and a second audio segment in a non-complete version audio in the same order, determine a start coordinate of a difference audio segment and a non-alignment point coordinate based on the first audio segment and the second audio segment that are not matched successfully, sequentially match a third audio segment after the non-alignment point and one or more fourth audio segments after the start coordinate, determine a termination coordinate of the difference audio segment based on the fourth audio segment that is matched successfully, and position the difference audio segment based on the start coordinate and the termination coordinate of the difference audio segment in the complete version audio, which can improve the efficiency of positioning the difference segment between audios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to an audio difference localization method and apparatus. Background Technology

[0002] Currently, there is a need to differentiate between full and incomplete audio versions, where the incomplete version lacks audio segments that differ from the full version.

[0003] To address this need, the current method primarily involves manually comparing the complete and incomplete audio versions using audio processing software to pinpoint the differences. However, this method requires visual comparison, which is time-consuming, inefficient, and lacks accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide an audio difference localization method and apparatus to improve the efficiency of locating difference segments between audio segments. The specific technical solution is as follows:

[0005] In a first aspect of this invention, an audio difference localization method is provided, the method comprising:

[0006] Obtain the complete audio and the incomplete audio; the incomplete audio is missing a difference audio segment that needs to be located compared to the complete audio.

[0007] Extract the audio features of each audio frame from the complete audio and the incomplete audio;

[0008] Using the number of first audio frames as the segment length, one or more first audio segments are sequentially determined in the complete audio version, and one or more second audio segments are sequentially determined in the incomplete audio version. The first audio segments and second audio segments in the same order are matched sequentially until the i-th first audio segment and the i-th second audio segment fail to match. The starting coordinates of the difference audio segment in the complete audio version are determined based on the coordinates of the audio frames in the i-th first audio segment, and the coordinates of the non-aligned point in the incomplete audio version are determined based on the coordinates of the audio frames in the i-th second audio segment. Wherein, a successful match between the first audio segment and the second audio segment indicates that the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment meets a preset condition.

[0009] Using the number of second audio frames as the segment length, a third audio segment is determined in the incomplete audio following the non-aligned point coordinates, and one or more fourth audio segments are sequentially determined in the complete audio following the starting coordinates. The third audio segment and one or more fourth audio segments are sequentially matched until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment.

[0010] Based on the start and end coordinates of the differing audio segment in the full audio version, the position range of the differing audio segment in the full audio version is located.

[0011] Optionally, the matching status of the first audio segment and the second audio segment can be determined based on the following method:

[0012] Cross-correlation calculation is performed on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain a cosine similarity sequence. It is then determined whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence.

[0013] If so, it is confirmed that the first audio segment and the second audio segment are successfully matched;

[0014] If not, it indicates that the current first and second audio segments have not matched successfully.

[0015] Optionally, the step of determining the starting coordinates of the differing audio segment in the complete audio version based on the position of the i-th first audio segment, and determining the coordinates of the non-aligned point in the incomplete audio version based on the position of the i-th second audio segment, includes:

[0016] Using the number of third audio frames as the shortening step, the i-th first audio segment and the i-th second audio segment are shortened simultaneously, and the shortened first audio segment and the shortened second audio segment are matched until the shortened first audio segment and the shortened second audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first audio segment are used as the starting coordinates of the difference audio segment in the complete version of the audio, and the coordinates of the last audio frame in the successfully matched i-th second audio segment are used as the coordinates of the non-aligned point in the incomplete version of the audio.

[0017] Optionally, the method further includes:

[0018] The complete audio following the termination coordinates is identified as the new complete audio, and the incomplete audio following the non-alignment point coordinates is identified as the new incomplete audio. The process returns to the steps of using the number of first audio frames as the segment length, sequentially identifying one or more first audio segments in the complete audio, sequentially identifying one or more second audio segments in the incomplete audio, and sequentially matching the first audio segments and second audio segments in the same order until both the first audio segments and the second audio segments are successfully matched.

[0019] Optionally, the complete audio is the audio corresponding to the submitted video, and the incomplete audio is the audio corresponding to the approved video.

[0020] The method further includes:

[0021] Based on the location range of the differential audio segments in the full audio version, the dubbing file obtained by dubbing the submitted video is trimmed to obtain a dubbing file adapted to the approved video version.

[0022] Optionally, the step of extracting the audio features of each audio frame in the complete audio and the incomplete audio specifically includes:

[0023] The complete audio is segmented into frames, and the segmented audio is input into a pre-trained speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio features of each audio frame in the complete audio. The speech recognition model includes a feature extraction module and a text recognition module. The feature extraction module is used to convert audio into audio features, and the text recognition module is used to recognize the audio features obtained by the feature extraction module as text.

[0024] The incomplete audio is segmented into frames, and the segmented incomplete audio is input into the speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio feature of each audio frame in the incomplete audio.

[0025] In a second aspect of the invention, an audio difference localization device is also provided, comprising:

[0026] The acquisition module is used to acquire the complete audio and the incomplete audio; the incomplete audio is missing a difference audio segment to be located relative to the complete audio.

[0027] An extraction module is used to extract the audio features of each audio frame in the complete audio and the incomplete audio.

[0028] The first matching module is used to determine one or more first audio segments in the complete audio version and one or more second audio segments in the incomplete audio version, using the number of first audio frames as the segment length. It then sequentially matches the first and second audio segments in the same order until the i-th first audio segment and the i-th second audio segment fail to match. Based on the coordinates of the audio frames in the i-th first audio segment, it determines the starting coordinates of the differing audio segment in the complete audio version and the coordinates of the non-aligned point in the incomplete audio version based on the coordinates of the audio frames in the i-th second audio segment. A successful match between the first and second audio segments indicates that the similarity between the audio features corresponding to the first and second audio segments meets a preset condition.

[0029] The second matching module is used to determine a third audio segment in the incomplete audio after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to sequentially determine one or more fourth audio segments in the complete audio after the starting coordinates. The third audio segment and one or more fourth audio segments are matched sequentially until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment.

[0030] The positioning module is used to locate the position range of the differing audio segment in the full version of the audio based on the start and end coordinates of the differing audio segment in the full version of the audio.

[0031] Optionally, the first matching module includes:

[0032] The judgment unit is used to perform cross-correlation calculation on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain a cosine similarity sequence, and to determine whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence.

[0033] If so, it is confirmed that the first audio segment and the second audio segment are successfully matched;

[0034] If not, it indicates that the current first and second audio segments have not matched successfully.

[0035] In a third aspect of the present invention, an electronic device is provided, the electronic device including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0036] Memory, used to store computer programs;

[0037] A processor, when executing a program stored in memory, implements the audio difference localization method described above.

[0038] In a fourth aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements any of the audio difference localization methods described above.

[0039] In a fifth aspect of the invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the audio difference localization methods described above.

[0040] The audio difference localization method provided in this invention involves sequentially determining one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio, and sequentially matching the first and second audio segments in the same order. The starting coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the unmatched first audio segments. The coordinates of the misalignment point in the incomplete audio are determined based on the coordinates of the audio frames in the unmatched second audio segments. A third audio segment is determined in the complete audio following the misalignment point coordinates, and one or more fourth audio segments are determined in the complete audio following the starting coordinates. The third and fourth audio segments are sequentially matched, and the ending coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segments. The position range of the difference audio segment in the complete audio is located based on the starting and ending coordinates. The audio difference localization method provided in this embodiment of the invention locates the difference audio segments by matching audio segments in the complete audio and incomplete audio. It can locate the missing difference audio segments in the incomplete audio without the need for manual comparison of the complete audio and incomplete audio, thus improving the efficiency of locating the difference segments between audio. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0042] Figure 1 This is a flowchart illustrating the audio difference localization method provided in an embodiment of the present invention;

[0043] Figure 2 This is an example diagram of a complete audio version and a non-complete audio version provided in an embodiment of the present invention;

[0044] Figure 3This is an example diagram showing the starting coordinates of the differential audio segments in the complete audio version and the coordinates of the non-aligned points in the non-complete audio version, provided by an embodiment of the present invention.

[0045] Figure 4 This is an example diagram of the start and end coordinates of the differential audio segments and the coordinates of the non-aligned points provided in the embodiments of the present invention;

[0046] Figure 5 This is an example diagram of the cross-correlation calculation process provided in an embodiment of the present invention;

[0047] Figure 6 This is an example diagram illustrating the matching process of the first audio segment and the second audio segment provided in an embodiment of the present invention;

[0048] Figure 7 This is another example diagram of the complete and incomplete audio provided in the embodiments of the present invention;

[0049] Figure 8 This is a schematic diagram of an audio difference localization method provided in an embodiment of the present invention;

[0050] Figure 9 This is a schematic diagram of the audio difference localization device provided in an embodiment of the present invention;

[0051] Figure 10 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0053] To address the low efficiency of manual audio difference localization, this invention provides an audio difference localization method. Figure 1 This is a flowchart illustrating the audio difference localization method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method specifically includes the following steps:

[0054] Step S101: Obtain the complete audio and the incomplete audio; the incomplete audio is missing the audio segment to be located compared to the complete audio.

[0055] Step S102: Extract the audio features of each audio frame from the complete audio and the incomplete audio.

[0056] Step S103: Using the number of first audio frames as the segment length, determine one or more first audio segments in the complete audio version sequentially, and determine one or more second audio segments in the incomplete audio version sequentially. Match the first and second audio segments in the same order sequentially until the i-th first audio segment and the i-th second audio segment fail to match. Determine the starting coordinates of the difference audio segment in the complete audio version based on the coordinates of the audio frames in the i-th first audio segment, and determine the coordinates of the non-aligned point in the incomplete audio version based on the coordinates of the audio frames in the i-th second audio segment. Wherein, a successful match between the first audio segment and the second audio segment indicates that the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment meets a preset condition.

[0057] Step S104: Using the number of second audio frames as the segment length, determine a third audio segment in the incomplete audio after the non-alignment point coordinates, and sequentially determine one or more fourth audio segments in the complete audio after the starting coordinates. Match the third audio segment and one or more fourth audio segments sequentially until a fourth audio segment that successfully matches the third audio segment is determined. Determine the termination coordinates of the difference audio segment in the complete audio based on the coordinates of the audio frames in the successfully matched fourth audio segment.

[0058] Step S105: Based on the start and end coordinates of the difference audio segment in the full audio, locate the position range of the difference audio segment in the full audio.

[0059] The audio difference localization method provided in this invention involves sequentially determining one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio, and sequentially matching the first and second audio segments in the same order. The starting coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the unmatched first audio segments. The coordinates of the misalignment point in the incomplete audio are determined based on the coordinates of the audio frames in the unmatched second audio segments. A third audio segment is determined in the complete audio following the misalignment point coordinates, and one or more fourth audio segments are determined in the complete audio following the starting coordinates. The third and fourth audio segments are sequentially matched, and the ending coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segments. The position range of the difference audio segment in the complete audio is located based on the starting and ending coordinates. The audio difference localization method provided in this embodiment of the invention locates the difference audio segments by matching audio segments in the complete audio and incomplete audio. It can locate the missing difference audio segments in the incomplete audio without the need for manual comparison of the complete audio and incomplete audio, thus improving the efficiency of locating the difference segments between audio.

[0060] The steps S101-S105 described above will be explained in detail below:

[0061] In step S101, some audio segments in the complete audio version have corresponding content in the incomplete audio version, but there is another part of the audio segment in the complete audio version that does not have corresponding content in the incomplete audio version. That is, the incomplete audio version is missing this part of the audio segment. In this embodiment of the invention, it is referred to as the difference audio segment.

[0062] Figure 2 This is an example diagram of a complete audio version and a non-complete audio version provided in an embodiment of the present invention, such as... Figure 2 As shown, the complete audio version contains audio segments m, k, and n, while the incomplete audio version contains audio segments m and n. Specifically, audio segment m in both the complete and incomplete audio versions corresponds to the same audio content, and audio segment n in both versions corresponds to the same audio content. Therefore, the incomplete audio version is missing audio segment k compared to the complete audio version; that is, audio segment k is the difference in audio content.

[0063] The audio difference localization method provided in this embodiment of the invention is specifically applied to scenarios where it is known that an incomplete audio version is missing compared to a complete audio version, but the specific location of the missing difference audio segment in the complete audio version is uncertain. The method locates the specific location of the difference audio segment in the complete audio version.

[0064] As an example, the complete audio version can be the audio corresponding to the submitted version of the film or television work, while the incomplete audio version is the audio corresponding to the approved version. The discrepancy audio segment corresponds to the audio corresponding to the segment that was cut during the review process. Since the time points of the cut content are not recorded during the film and television review process, the specific location of the discrepancy audio segment in the audio corresponding to the submitted version of the film or television work is uncertain.

[0065] In step S102, the audio features of each audio frame in the complete audio and incomplete audio are extracted, where the audio features of each audio frame can be understood as a feature vector.

[0066] For specific methods of extracting audio features, please refer to the relevant technical works. This embodiment of the invention does not limit the scope of the invention.

[0067] As an example, audio energy information, time domain information, or frequency domain information can be collected through filters, and the collected data can be encoded to obtain audio features.

[0068] As another example, audio features can also be extracted using deep learning neural networks.

[0069] In step S103, using the number of first audio frames as the segment length, one or more first audio segments are sequentially determined in the complete audio version, and one or more second audio segments are sequentially determined in the incomplete audio version. Specifically, based on the audio frames in the complete and incomplete audio versions, the complete and incomplete audio versions are segmented, and each determined first or second audio segment consists of the number of first audio frames. The number of first audio frames can be set based on actual needs.

[0070] As an example, if the number of the first audio frames is set to 10, then the first 10 audio frames in the full version of the audio are the first audio segment, the 11th to 20th audio frames are the second audio segment, and so on. The process of determining the second audio segment in the incomplete version of the audio is the same.

[0071] The process of matching the first and second audio segments involves sequentially matching the first and second audio segments in the same order until the i-th first audio segment and the i-th second audio segment fail to match. That is, the first first audio segment and the first second audio segment are matched; if a match is found, the second first audio segment and the second second audio segment are matched, and so on, until the first and second audio segments that failed to match are identified.

[0072] Among them, the first audio segment and the second audio segment are successfully matched, which specifically indicates that the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment meets the preset conditions.

[0073] When calculating the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment, the audio features of each audio frame in the first audio segment can be concatenated to obtain a vector, and the audio features of each audio frame in the second audio segment can be concatenated to obtain another vector. The similarity between these two vectors can then be calculated.

[0074] As an example, if the number of first audio frames is set to 3, and the audio features of the three audio frames in the first audio segment are x1, x2, and x3, and the audio features of the three audio frames in the second audio segment are x4, x5, and x6, calculating the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment can be understood as calculating the similarity between (x1, x2, x3) and (x4, x5, x6). For specific methods of calculating the similarity between feature vectors, please refer to relevant technical articles.

[0075] The preset similarity criteria used to determine whether the first and second audio segments match successfully can be selected according to actual needs, and this embodiment of the invention does not limit this. For example, a similarity threshold can be preset; if the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment is not less than the similarity threshold, then the match is considered successful.

[0076] After determining that the i-th first audio segment and the i-th second audio segment did not match successfully, the starting coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frame in the i-th first audio segment, and the coordinates of the non-alignment point in the incomplete audio are determined based on the coordinates of the audio frame in the i-th second audio segment.

[0077] In this context, misalignment points in incomplete audio can be understood as follows: if the second audio segment before the point in the incomplete audio is identical to the first audio segment in the complete audio (in the same order), then a match is successful; conversely, if the second audio segment after the point in the incomplete audio is different from the first audio segment in the complete audio (in the same order), then a match is unsuccessful. For example... Figure 2 The point of intersection between audio segment m and audio segment n in the incomplete audio version shown in the figure is the non-alignment point.

[0078] Regarding the coordinates of audio frames, each audio frame in both the full and incomplete audio versions typically has a fixed duration. Therefore, the start and end coordinates of each audio frame in both versions can be considered known. As an example, if the duration of each audio frame is 20ms, then the start and end coordinates of the first audio frame in the full audio version are 0ms and 20ms respectively, the start and end coordinates of the second audio frame are 20ms and 40ms respectively, and so on.

[0079] As an example, since the duration of each audio frame is fixed, the sequence number of the audio frame can also be used to represent the time point in the audio. Therefore, the sequence number of the audio frame can also be used to represent the starting coordinates of the difference audio segment in the full version of the audio and the coordinates of the non-aligned point in the incomplete audio segment. For example, the starting coordinates of the difference audio segment in the full version of the audio are the 100th frame in the full version of the audio.

[0080] As an example, the starting coordinates of the difference audio segment in the full audio and the non-aligned point coordinates in the non-full audio segment can also be represented by the time of the audio frame. For example, the starting coordinates of the difference audio segment in the full audio are the 100th second in the full audio.

[0081] As mentioned earlier, each first or second audio segment includes multiple audio frames. The starting coordinates and misalignment point coordinates of the differing audio segment are determined based on the coordinates of that specific audio frame, which can be chosen according to actual needs. For example, the starting coordinates of the differing audio segment in the full audio can be determined based on the coordinates of the first audio frame in the i-th first audio segment, and the misalignment point coordinates in the incomplete audio can be determined based on the coordinates of the first audio frame in the i-th second audio segment. For instance, if the number of first audio frames is set to 10, and the first first audio segment and the first second audio segment match successfully, but the second first audio segment and the second second audio segment do not match successfully, then the starting coordinates of the differing audio segment in the full audio are frame 11, and the misalignment point coordinates in the incomplete audio are also frame 11.

[0082] As an example, the number of the first audio frames can also be set to 1 frame, then each first audio segment and the second audio segment specifically constitutes one audio frame. Therefore, each audio frame in the complete audio version can be matched sequentially with the audio frames in the same order in the incomplete audio version until the i-th audio frame in the complete audio version and the i-th audio frame in the incomplete audio version fail to match. In this case, the starting coordinate of the differing audio segment in the complete audio version is the i-th audio frame, and the coordinate of the non-aligned point in the incomplete audio version is the i-th audio frame.

[0083] The following will illustrate this with specific examples. Figure 3 This is an example diagram showing the starting coordinates of the differing audio segments in the complete audio version and the coordinates of the non-aligned points in the non-complete audio version, provided by an embodiment of the present invention. Figure 3 The shaded area in the image represents the audio segment to be located. As mentioned earlier, the location of the audio segment in the incomplete audio version is unknown.

[0084] like Figure 3 As shown, the process first matches the first audio segment A in the complete audio and the second audio segment A' in the incomplete audio. If the match is successful, the process then matches the first audio segment B in the complete audio and the second audio segment B' in the incomplete audio. If the match is successful, the process continues to match the first audio segment C in the complete audio and the second audio segment C' in the incomplete audio. If the match fails, the starting coordinates of the differing audio segment in the complete audio are determined based on the coordinates of the audio frames in the first audio segment C, as shown in Figure x11. The coordinates of the non-aligned point in the incomplete audio are determined based on the coordinates of the audio frames in the second audio segment, as shown in Figure x1.

[0085] In step S104, after determining the starting coordinates of the difference audio segment in the full audio, it is also necessary to determine the ending coordinates of the difference audio segment in the full audio segment to locate the difference audio segment.

[0086] After determining the starting coordinates of the difference audio segment in the complete audio and the misalignment point coordinates in the incomplete audio in step S103, a third audio segment can be determined in the incomplete audio after the misalignment point based on the number of the second audio frames. That is, the audio frames after the misalignment point are considered as a third audio segment. Similarly, one or more fourth audio segments can be determined sequentially in the complete audio after the starting coordinates. For example, if the number of the second audio frames is set to 3 and the misalignment point coordinate is frame 11, then frames 11-13 in the incomplete audio are a third audio segment. If the starting coordinate is frame 11, then frames 11-13 in the complete audio are the first fourth audio segment, frames 14-16 are the second fourth audio segment, and so on.

[0087] Based on this, the third audio segment is sequentially matched with one or more fourth audio segments until a fourth audio segment that successfully matches the third audio segment is identified. The termination coordinates of the differing audio segment in the full audio version are then determined based on the coordinates of the audio frames within that fourth audio segment. As an example, the termination coordinates of the differing audio segment in the full audio version can be determined based on the coordinates of the first audio frame within the fourth audio segment.

[0088] The number of the second audio frames can be configured according to actual needs. As an example, to improve the accuracy of the positioning results, the number of the second audio frames can be set to a value less than the number of the first audio frames; for example, the number of the second audio frames can be 3 frames or 1 frame.

[0089] As an example, the second audio frame can be 1 frame. In the incomplete audio, the first audio frame after the non-aligned point is the third audio segment, and in the complete audio, the first audio frame after the starting coordinate is the first fourth audio segment, the second audio frame is the second fourth audio segment, and so on.

[0090] The following is combined with Figure 4 This example is illustrated. Figure 4 This is an example diagram showing the start and end coordinates of the differential audio segments and the coordinates of the non-aligned points provided in this embodiment of the invention. Figure 4 and Figure 2 Correspondingly, specifically, Figure 4 The shaded area in the image represents the audio segment with discrepancies to be located.

[0091] Figure 4The diagram shows the starting coordinate x11 of the difference audio segment determined in step S103 within the complete audio, and the non-aligned point coordinate x1 within the incomplete audio. Using the number of the second audio frames as the segment length, a third audio segment a' can be determined in the incomplete audio after the non-aligned point coordinate x1, and the first fourth audio segment a can be determined in the complete audio after the starting coordinate x11 of the difference audio segment. The third audio segment a' and the fourth audio segment a are then matched. If the match fails, the third audio segment a' and the fourth audio segment b are matched. If the match fails, the third audio segment a' and the fourth audio segment c are matched. If the match succeeds, the ending coordinate of the difference audio segment in the complete audio is determined based on the coordinates of the audio frames in the fourth audio segment b. For example, the ending coordinate x12 of the difference audio segment in the complete audio can be determined based on the coordinates of the first audio frame in the fourth audio segment b.

[0092] exist Figure 4 In the example, if the number of the second audio frames is set to 1, then the third audio segment a' can be understood as... Figure 2 The first audio frame within audio segment n in the incomplete audio version shown, and the fourth audio segment a can be understood as the first audio frame within audio segment k in the complete audio version.

[0093] The fourth audio segment c can be understood as Figure 2 The first audio frame within audio segment n in the complete audio version shown is used to determine the third audio segment a' and the fourth audio segment c. Therefore, the third audio segment a' and the fourth audio segment c can be matched successfully. Given that the third audio segment a' and the fourth audio segment c are matched successfully, the termination coordinate x12 of the difference audio segment in the complete audio version can be determined based on the coordinates of the audio frame in the fourth audio segment c, thus realizing the location of the difference audio segment.

[0094] Similar to the matching process for the first and second audio segments, a successful match between the third and fourth audio segments can be considered if the similarity between the audio features corresponding to the third and fourth audio segments meets a preset condition. This embodiment of the invention does not limit the preset condition. As an example, based on a pre-set similarity threshold between the third and fourth audio segments, if the similarity between the audio features corresponding to the third audio segment and the audio features corresponding to a certain fourth audio segment is not less than the similarity threshold, then the third audio segment can be considered to have successfully matched the fourth audio segment.

[0095] It should be understood that in this embodiment of the invention, if the third audio segment and the fourth audio segment are audio segments with different corresponding content, then the similarity between the third audio segment and the fourth audio segment is low and does not exceed a predetermined similarity threshold, that is, the third audio segment and the fourth audio segment cannot be matched successfully.

[0096] If the third and fourth audio segments have the same content, then the similarity between the third and fourth audio segments is high, exceeding the predetermined similarity threshold, meaning that the third and fourth audio segments can be successfully matched.

[0097] In step S105, the starting and ending coordinates of the difference audio segment in the full audio version can be used to determine the position range of the difference audio segment in the full audio version.

[0098] Specifically, the audio frames between the start and end coordinates in the full audio can be considered as the content of the difference audio segment.

[0099] As an example, when using audio frame numbers to represent coordinates, if the starting coordinate of the differing audio segment in the full audio is x11 and the ending coordinate is x12, then the position interval of the differing audio segment in the full audio is [x11, x12]. Specifically, since the ending coordinate x12 of the differing audio segment is determined based on the coordinates of the audio frames in the fourth audio segment that successfully matches the third audio segment, and since the audio frames in the fourth audio segment do not belong to the content of the differing audio segment when the fourth audio segment and the third audio segment successfully match, this endpoint coordinate is not included in the position interval of the differing audio segment.

[0100] For example, if the starting coordinate of the difference audio segment in the full audio is frame 100 and the ending coordinate is frame 200, then the position range of the difference audio segment in the full audio is [100, 200), which means that frames 100-199 in the full audio are the content of the difference audio segment.

[0101] The audio difference localization method provided in this invention involves sequentially determining one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio, and sequentially matching the first and second audio segments in the same order. The starting coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the unmatched first audio segments. The coordinates of the misalignment point in the incomplete audio are determined based on the coordinates of the audio frames in the unmatched second audio segments. A third audio segment is determined in the complete audio following the misalignment point coordinates, and one or more fourth audio segments are determined in the complete audio following the starting coordinates. The third and fourth audio segments are sequentially matched, and the ending coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segments. The position range of the difference audio segment in the complete audio is located based on the starting and ending coordinates. The audio difference localization method provided in this embodiment of the invention locates the difference audio segments by matching audio segments in the complete audio and incomplete audio. It can locate the missing difference audio segments in the incomplete audio without the need for manual comparison of the complete audio and incomplete audio, thus improving the efficiency of locating the difference segments between audio.

[0102] In one embodiment of the present invention, the complete audio is the audio corresponding to the submitted video, and the incomplete audio is the audio corresponding to the approved video.

[0103] Audio difference localization methods also include:

[0104] Based on the location range of the audio segments with different positioning in the full audio version, the dubbing file obtained by dubbing the submitted video is trimmed to obtain a dubbing file adapted to the approved video.

[0105] Specifically, regarding the dubbing of film and television works, the dubbing of overseas versions of film and television works usually begins after the approved version of the film and television work is ready. Since the dubbing work requires a lot of time, the release of overseas dubbed film and television works will be later than the release of the original dubbed film and television works. When the overseas dubbed film and television works are released, they cannot make full use of the promotional methods used by the original dubbed film and television works to increase viewership, such as promotional positions and external traffic.

[0106] If the overseas dubbing is done before the approved version of the film or television work is prepared, the specific location of the audio content to be cut needs to be determined in the submitted version after the approved version is ready. The corresponding content from the pre-prepared overseas dubbing is then cut to obtain an overseas dubbing that matches the approved version, allowing the film or television work to be released online. However, based on existing audio difference localization methods, manual comparison of the audio corresponding to the submitted version and the audio corresponding to the approved version is required to determine the specific location of the cut audio content in the submitted version. This process is inefficient, time-consuming, and costly, and therefore not suitable.

[0107] When applying the audio difference localization method provided in this embodiment of the invention, the pre-prepared overseas dubbing can be cut based on the position range of the located difference audio segments in the complete audio, i.e., the audio corresponding to the submitted version of the medium, to obtain an overseas dubbing that matches the approved version of the medium. Therefore, the dubbing work for the overseas version can be carried out before the approved version of the film and television work is prepared, and the matching overseas dubbing can be obtained in a timely manner after the approved version of the medium is prepared. This allows the film and television works with overseas dubbing and those with original dubbing to be launched simultaneously, and the film and television works with overseas dubbing can make full use of the promotional methods used by the film and television works with original dubbing.

[0108] In one embodiment of the present invention, it can be determined whether the first audio segment and the second audio segment are successfully matched based on the following method:

[0109] Cross-correlation calculation is performed on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain the cosine similarity sequence. It is then determined whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence.

[0110] If so, confirm that the first and second audio segments match successfully;

[0111] If not, it indicates that the current first and second audio segments have not matched successfully.

[0112] As mentioned above, each first audio segment or second audio segment may specifically include a number of first audio frames. If the number of first audio frames is N, then a cross-correlation operation is performed on the audio features corresponding to a first audio segment and the audio features corresponding to a second audio segment. Specifically, a similarity sequence with 2N-1 items can be obtained. If the coordinates of the similarity peak in the similarity sequence are the center coordinates of the similarity sequence, that is, the peak of the similarity sequence is the Nth item, then the first audio segment and the second audio segment are considered to be successfully matched; otherwise, the first audio segment and the second audio segment are not successfully matched.

[0113] As mentioned above, the audio features corresponding to the first and second audio segments are specifically feature vectors. Therefore, the similarity calculated in this embodiment of the invention can specifically be cosine similarity.

[0114] For details on how to perform cross-correlation operations on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment, please refer to the relevant technical content. The following is a brief explanation of the cross-correlation operation with specific examples.

[0115] Figure 5 This is an example diagram of the cross-correlation calculation process provided in the embodiments of the present invention. For ease of description, it is specifically taken as an example that the number of the first audio frames is 3. Figure 5 As shown, the three audio frames in the first audio segment are labeled 0, 1, and 2, with corresponding audio features x0, x1, and x2, respectively. The three audio frames in the second audio segment are labeled 3, 4, and 5, with corresponding audio features x3, x4, and x5, respectively. Figure 5 The audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment are cross-correlated. Specifically, it can be understood that, with the first audio segment as a reference, the second audio segment is slid at the position of the first audio segment. For the audio frames on the second audio segment that have slid to the position of the first audio segment, the similarity between the audio features corresponding to these audio frames and the audio features corresponding to the first audio segment is calculated. In order to ensure that the data length of the audio features is consistent when calculating the similarity, zeros can be padded at the positions of missing data.

[0116] See Figure 5 In (a), audio frame 5 in the second audio segment slides to the corresponding position of audio frame 0. At this point, the similarity between feature vectors (x0,x1,x2) and (x5,0,0) can be calculated. Similarly, in (b)-(e), a total of 5 similarity scores can be calculated. If the similarity scores of (a)-(e) are 0.2, 0.2, 0.9, 0.3, and 0.1 respectively, then the coordinates of the peak similarity score in the similarity sequence are the center coordinates of the similarity sequence, and the first and second audio segments are considered to have matched successfully.

[0117] It should be understood that if the coordinates of the similarity peak in the similarity sequence obtained by cross-correlation calculation are the center coordinates of the similarity sequence, specifically indicating that when the first and second audio segments are aligned in position, the audio features corresponding to the first and second audio segments have the highest similarity, then the first and second audio segments are considered to have matched successfully. Therefore, judging whether the first and second audio segments can match successfully based on cross-correlation calculation does not depend on the specific numerical value of the similarity, and can improve the accuracy of matching the first and second audio segments.

[0118] In one embodiment of the present invention, the aforementioned steps of determining the starting coordinates of the differing audio segment in the complete audio version based on the position of the i-th first audio segment, and determining the coordinates of the misaligned point in the incomplete audio version based on the position of the i-th second audio segment, may specifically include:

[0119] Using the number of third audio frames as the shortening step, the i-th first audio segment and the i-th second audio segment are shortened simultaneously, and the shortened first audio segment and the shortened second audio segment are matched until the shortened first audio segment and the shortened second audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first audio segment are used as the starting coordinates of the difference audio segment in the complete version of the audio, and the coordinates of the last audio frame in the successfully matched i-th second audio segment are used as the coordinates of the non-aligned point in the incomplete version of the audio.

[0120] The number of third audio frames can be predetermined and is less than the number of first audio frames. For example, if the number of first audio frames is 10, the number of third audio frames can be 3 or 1.

[0121] In this embodiment of the invention, the i-th first audio segment and the i-th second audio segment are shortened simultaneously, and the shortened first audio segment and the shortened second audio segment are matched until the shortened first audio segment and the second audio segment are successfully matched. Specifically, for the first audio segment and the second audio segment that are not successfully matched, the number of third audio frames of both the first audio segment and the second audio segment are shortened simultaneously, and the shortened first audio segment and the second audio segment are matched. If the match is unsuccessful, the number of third audio frames of both the first audio segment and the second audio segment are shortened again and matched, and so on, until the shortened first audio segment and the second audio segment are successfully matched.

[0122] As an example, if the third audio frame has 2 frames, the 11th-20th audio frames in the full version are the second first audio segment, and the 11th-20th audio frames in the incomplete version are the second second audio segment. If the second first audio segment and the second second audio segment do not match, then the 11th-18th audio frames in the full version are used as the shortened first audio segment, and the 11th-18th audio frames in the incomplete version are used as the shortened second audio segment. The shortened first audio segment and the second audio segment are matched. If the match fails, the first audio segment and the second audio segment are shortened again, and the 11th-16th audio frames in the full version and the 11th-16th audio frames in the incomplete version are matched. This process continues until the shortened first audio segment and the second audio segment match successfully.

[0123] After the shortened first and second audio segments are successfully matched, the coordinates of the last audio frame in the i-th successfully matched first audio segment are used as the starting coordinates of the difference audio segment in the full version of the audio, and the coordinates of the last audio frame in the i-th successfully matched second audio segment are used as the coordinates of the non-aligned point in the non-full version of the audio.

[0124] As an example, if the shortened first audio segment is frames 11-16, the shortened second audio segment is frames 11-16, and the first and second audio segments match successfully, then the starting coordinate of the difference audio segment in the full version of the audio is frame 16, and the coordinate of the non-aligned point in the incomplete version of the audio is frame 16.

[0125] It should be noted that in this embodiment of the invention, if the starting coordinates of the differing audio segment in the complete audio are x11, since x11 is specifically determined based on the coordinates of the audio frames in the first audio segment that successfully matches the second audio segment, the audio frames in the first audio segment do not belong to the content of the differing audio segment when the second audio segment and the first audio segment successfully match. Therefore, when locating the position range of the differing audio segment subsequently, the endpoint coordinates are not included in the position range of the differing audio segment. That is, if the ending coordinate of the differing audio segment in the complete audio is x12, then the position range of the differing audio segment in the complete audio is (x11, x12).

[0126] Specifically, since the shortened first and second audio segments match successfully, and the content to be located is the missing difference audio segment in the incomplete audio segment, the termination coordinate of the last audio frame is used as the starting coordinate of the difference audio segment or the non-alignment point coordinate in the incomplete audio. For example, after the shortened first and second audio segments match successfully, if the last audio frame in the shortened first audio segment is the 17th frame and the length of each audio frame is 20ms, then the starting coordinate of the difference audio segment in the complete audio is 34ms.

[0127] The following will illustrate this with specific examples. Figure 6 This is an example diagram illustrating the matching process of the first and second audio segments provided in an embodiment of the present invention. Figure 6 The shaded areas represent the audio segments with differences. Figure 6(a) shows the first audio segments A, B, C, and the second audio segments A', B', C'. The first audio segments A and the second audio segments A' and B and B' are both successfully matched. If the first audio segments C and the second audio segments C' are not successfully matched, the first audio segments C and C' are shortened simultaneously until the shortened first audio segments and the second audio segments are successfully matched. In (b), D is the shortened first audio segment obtained by shortening the first audio segment C, and D' is the shortened second audio segment obtained by shortening the second audio segment D. The shortened first audio segment D and the shortened second audio segment D' are successfully matched. Then, the coordinates of the last audio frame in the first audio segment D are taken as the starting coordinates x11 of the difference audio segment in the complete audio, and the coordinates of the last audio frame in the second audio segment D' are taken as the non-aligned point coordinates x1 in the incomplete audio.

[0128] In practical applications, when matching the first and second audio segments using the number of first audio frames as the segment length, the starting position of the differing audio segment in the complete audio has a certain probability of being located in the middle of the first audio segment. If the first and second audio segments fail to match, it is difficult to accurately determine the starting coordinates of the differing audio segment in the complete audio and the coordinates of the misaligned point in the incomplete audio. In this embodiment of the invention, when the first and second audio segments fail to match, the unmatched first and second audio segments are shortened simultaneously until the shortened first and second audio segments match successfully. The coordinates of the last audio frame in the shortened first audio segment are used as the starting coordinates of the differing audio segment in the complete audio, and the coordinates of the last audio frame in the shortened second audio segment are used as the coordinates of the misaligned point in the incomplete audio, thus improving the accuracy of locating the differing audio segment.

[0129] In one embodiment of the present invention, the incomplete audio version may specifically lack several features compared to the complete audio version. Figure 7 This is another example diagram of the complete and incomplete audio provided in this invention, such as... Figure 7 As shown, the incomplete audio version is missing audio segments k and p compared to the complete audio version.

[0130] Therefore, in this embodiment, the audio difference localization method provided by the present invention further includes:

[0131] The complete audio after the termination coordinate is identified as the new complete audio, and the incomplete audio after the non-alignment point coordinate is identified as the new incomplete audio. The process returns to the steps of identifying one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio, using the number of the first audio frames as the segment length, and matching the first and second audio segments in the same order, until both the first and second audio segments are successfully matched.

[0132] Specifically, after locating the start and end coordinates of a differing audio segment within the complete audio, it is necessary to match the subsequent audio content in both the complete and incomplete audio versions to locate other potentially missing differing audio segments in the incomplete audio. Therefore, the complete audio after the end coordinate is defined as the new complete audio, and the incomplete audio after the non-alignment point coordinate is defined as the new incomplete audio. The first audio segment in the newly defined complete audio and the second audio segment in the incomplete audio are then matched. The specific matching process can be found in the explanation of steps S103-S104 above. After locating the second missing differing audio segment in the incomplete audio, the location of subsequent differing audio segments is done in the same way, until all audio content in both the complete and incomplete audio versions is matched.

[0133] The following is combined with Figure 7 The following is an exemplary description of an embodiment of the present invention. A1 and B1 are the original complete audio and incomplete audio, respectively. Based on steps S103-S104, the first audio segment in A1 and the second audio segment in B1 are matched to determine the starting coordinates x11 and ending coordinates x12 of the first differing audio segment in the complete audio, thus locating the first differing audio segment k. After locating k, the complete audio A2 after the ending coordinate x12 is determined as the new complete audio, and the incomplete audio B2 after the non-alignment point coordinate x1 is determined as the new incomplete audio. Based on steps S103-S104, a new round of matching is performed on the first audio segment in A2 and the second audio segment in B2 to determine the starting coordinates x21 and ending coordinates x22 of the second differing audio segment in the complete audio, thus locating the second differing audio segment p. After locating p, the complete audio A3 after the termination coordinate x22 is identified as the new complete audio, and the incomplete audio B3 after the non-alignment point coordinate x2 is identified as the new incomplete audio. A new round of matching is performed on the first audio segment in A3 and the second audio segment in B3. If the subsequent incomplete audio is missing other differential audio segments, the steps to locate these differential audio segments are similar.

[0134] Based on the embodiments of the present invention, the starting coordinates and ending coordinates of multiple differential audio segments in the complete audio version can be obtained. For example, the starting coordinate of the i-th differential audio segment in the incomplete audio version is xi1, and the ending coordinate is xi2.

[0135] Specifically, based on xi1 and xi2, the position range of the i-th difference audio segment in the full audio version can be located.

[0136] In this embodiment of the invention, by determining the complete audio after the termination coordinate as the new complete audio and the incomplete audio after the non-alignment point coordinate as the new incomplete audio, it is possible to locate multiple differential audio segments missing in the incomplete audio, thereby improving the applicability of the audio difference localization method.

[0137] In one embodiment of the present invention, the aforementioned step S102 may specifically include:

[0138] The complete audio is segmented into frames. The segmented audio is then input into a pre-trained speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio features of each audio frame in the complete audio. The speech recognition model includes a feature extraction module and a text recognition module. The feature extraction module is used to convert audio into audio features, and the text recognition module is used to recognize the audio features obtained by the feature extraction module as text.

[0139] The incomplete audio is segmented into frames. The segmented incomplete audio is then input into a speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio feature of each audio frame in the incomplete audio. The first audio segment is input into the speech recognition model, and the output of the speech recognition model at the encoder layer is extracted to obtain the first audio feature.

[0140] Specifically, the language recognition model can refer to the encoder layer and the text conversion module can refer to the decoder layer. The encoder layer can convert the input audio into a feature vector, and the decoder layer can map the feature vector into a character sequence.

[0141] Compared to the model's final output, the character sequence obtained by the decoder layer, the feature vector obtained by the encoder layer contains much more information about the audio content. For example, the feature vector obtained by the encoder layer can not only represent the characters corresponding to the audio, but also represent non-character implicit information such as tone, intonation, and speech rate in the audio.

[0142] Furthermore, since the feature vectors obtained from the encoder layer can represent the implicit information in the audio, and since the output of the encoder layer is specifically the output of the previous layer of the decoder, the extraction of audio features does not strongly depend on the training accuracy of the speech recognition model, and can represent the audio content information at different temporal granularities.

[0143] As an example, to improve the ability and accuracy of audio features in representing implicit information, a large amount of highly expressive speech recognition corpus can be collected for model training during the training process. For instance, the model can be trained based on 8,000 hours of speech recognition corpus, which can specifically come from open-source datasets, TV series, movies, and variety shows. This not only improves the ability and accuracy of the first and second audio features in representing implicit information, but also enhances the adaptability of the audio difference localization method provided in this embodiment of the invention to different scenarios.

[0144] This invention does not specifically limit the scope of the speech recognition model. As an example, it can be an e2eASR (End-to-End automatic speech recognition) model. Furthermore, when the speech recognition model is an e2e ASR model, it can be trained using ESPNet (an open-source speech recognition training platform).

[0145] Furthermore, in this embodiment of the invention, the temporal resolution of the downsampling module in the encoder layer can also be set to a higher resolution.

[0146] Specifically, speech recognition models are typically used for speech recognition tasks, i.e., recognizing characters corresponding to audio. These tasks have low requirements for temporal resolution. For example, the temporal resolution of the downsampling module in the encoder layer may be 60ms or 80ms. However, a low temporal resolution may cause some detailed information to be lost in the obtained audio features.

[0147] Therefore, to enable the audio features output by the encoder layer to represent more detailed information, such as language content where tone, intonation, and speech rate change rapidly within a short period, the temporal resolution of the downsampling module in the encoder layer can be set to a higher resolution during model training. For example, the temporal resolution of the downsampling module in the encoder layer can be increased to the frame level, such as 10ms. Based on this, the audio features extracted by the encoder layer of the speech recognition model have higher accuracy.

[0148] In this embodiment of the invention, the complete audio is processed by frame segmentation, and the processed complete audio is input into a pre-trained speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio feature of each audio frame in the complete audio. The obtained audio features can characterize the implicit information contained in the complete audio and the incomplete audio. Based on this, the matching of the first audio segment and the second audio segment and the location of the missing difference audio segment in the incomplete audio have higher accuracy.

[0149] Figure 8 This is a schematic diagram of an audio difference localization method provided in an embodiment of the present invention. The following is in conjunction with... Figure 8 The audio difference localization method provided in the embodiments of the present invention will be further described.

[0150] like Figure 8 As shown, first, we obtain the audio Mix_1 corresponding to the submitted version of the media (i.e., the complete audio), and the audio Mix_2 corresponding to the approved version of the media (i.e., the incomplete audio). Then, we use an e2e ASR model to extract EVs from Mix_1 and Mix_2 respectively, that is, to obtain the audio features output by the encoder layer of the model, thus obtaining the audio feature sequence EV1 corresponding to the complete audio and the audio feature sequence EV2 corresponding to the incomplete audio. The process involves matching the audio feature EV1_sub_i corresponding to the first audio segment in the complete audio with the audio feature EV2_sub_i corresponding to the second audio segment in the same order in the incomplete audio. If a match is successful, the process continues to match the next audio feature EV1_sub_i+1 corresponding to the first audio segment and the next audio feature EV2_sub_i+1 corresponding to the second audio segment. If a match is unsuccessful, the first and second audio segments are shortened, and the audio features EV1_sub_i and EV2_sub_i corresponding to the shortened first and second audio segments are matched until a match is successful. This yields the starting coordinate xi1 of the difference audio segment in the complete audio. Then, the process matches the third audio segment after the non-aligned point coordinate xi in the incomplete audio with one or more fourth audio segments after xi2 in the complete audio to obtain the ending coordinate xi2 of the difference audio segment in the complete audio. This process can also be understood as a sliding match between the complete and incomplete audio. After matching is complete, you can perform sequence deletion on the overseas version of the voice acting, that is, delete the content between x11 and x12 and between xi1 and xi2 to obtain the trimmed overseas version of the voice acting.

[0151] In this embodiment of the invention, audio features of the approved and submitted audio versions are extracted using a speech recognition model. The first audio segment in the complete audio version and the second audio segment in the same order in the incomplete audio version are matched sequentially. Based on the unmatched first and second audio segments, the starting coordinates and misalignment point coordinates of the differing audio segments are determined. Then, the third audio segment after the misalignment point and one or more fourth audio segments after the starting coordinates are matched sequentially. Based on the successfully matched fourth audio segments, the ending coordinates of the differing audio segments are determined. The differing audio segments are located based on their starting and ending coordinates in the complete audio version. This allows for the identification of additional audio content in the submitted audio version, improving the efficiency of locating differing content in the audio. Furthermore, the obtained locations can be used to trim overseas dubbing, improving the dubbing efficiency of films and television dramas.

[0152] Corresponding to the above method embodiments, this application also provides an audio difference localization device, such as... Figure 9 As shown, the device includes:

[0153] Module 901 is used to acquire both the complete and incomplete audio versions; the incomplete audio version is missing audio segments that need to be located compared to the complete audio version.

[0154] Extraction module 902 is used to extract the audio features of each audio frame in the complete audio and incomplete audio.

[0155] The first matching module 903 is used to determine one or more first audio segments in the complete audio version and one or more second audio segments in the incomplete audio version, using the number of first audio frames as the segment length, and to match the first audio segments and second audio segments in the same order in sequence until the i-th first audio segment and the i-th second audio segment fail to match. Based on the coordinates of the audio frames in the i-th first audio segment, the starting coordinates of the differing audio segment in the complete audio version are determined, and based on the coordinates of the audio frames in the i-th second audio segment, the coordinates of the non-aligned point in the incomplete audio version are determined. A successful match between the first audio segment and the second audio segment indicates that the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment meets a preset condition.

[0156] The second matching module 904 is used to determine a third audio segment in the incomplete audio after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to sequentially determine one or more fourth audio segments in the complete audio after the starting coordinates. The third audio segment and one or more fourth audio segments are matched sequentially until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment.

[0157] The positioning module 905 is used to locate the position range of the difference audio segment in the full version of the audio based on the start coordinates and end coordinates of the difference audio segment in the full version of the audio.

[0158] The audio difference localization device provided in this invention determines one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio by sequentially matching the first and second audio segments in the same order. Based on the coordinates of the audio frames in the unmatched first audio segments, the starting coordinates of the difference audio segment in the complete audio are determined. Based on the coordinates of the audio frames in the unmatched second audio segments, the coordinates of the misalignment point in the incomplete audio are determined. A third audio segment is determined in the complete audio following the misalignment point coordinates. One or more fourth audio segments are determined in the complete audio following the starting coordinates. The third and fourth audio segments are matched sequentially. Based on the coordinates of the audio frames in the successfully matched fourth audio segments, the ending coordinates of the difference audio segment in the complete audio are determined. The device then locates the position range of the difference audio segment in the complete audio based on the starting and ending coordinates. The audio difference localization method provided in this embodiment of the invention locates the difference audio segments by matching audio segments in the complete audio and incomplete audio. It can locate the missing difference audio segments in the incomplete audio without the need for manual comparison of the complete audio and incomplete audio, thus improving the efficiency of locating the difference segments between audio.

[0159] In one embodiment of the present invention, the first matching module 903 includes:

[0160] The judgment unit is used to perform cross-correlation calculation on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain a cosine similarity sequence, and to determine whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence.

[0161] If so, confirm that the first and second audio segments match successfully;

[0162] If not, it indicates that the current first and second audio segments have not matched successfully.

[0163] In one embodiment of the present invention, the first matching module 903 includes:

[0164] The shortening unit is used to simultaneously shorten the i-th first audio segment and the i-th second audio segment with the number of third audio frames as the shortening step size, and to match the shortened first audio segment and the shortened second audio segment until the shortened first audio segment and the shortened second audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first audio segment are used as the starting coordinates of the difference audio segment in the complete version of the audio, and the coordinates of the last audio frame in the successfully matched i-th second audio segment are used as the coordinates of the non-aligned point in the incomplete version of the audio.

[0165] In one embodiment of the present invention, the device further includes:

[0166] The determination module is used to determine the complete audio after the termination coordinate as the new complete audio, and the incomplete audio after the non-alignment point coordinate as the new incomplete audio. It returns the steps of determining one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio in sequence, using the number of first audio frames as the segment length, and matching the first audio segments and second audio segments in the same order in sequence, until both the first audio segments and the second audio segments are successfully matched.

[0167] In one embodiment of the present invention, the complete audio is the audio corresponding to the submitted video, and the incomplete audio is the audio corresponding to the approved video.

[0168] The device also includes:

[0169] The trimming module is used to trim the dubbing file obtained by dubbing the submitted video in advance, based on the position range of the difference audio segments in the full version of the audio, to obtain a dubbing file adapted to the approved video.

[0170] In one embodiment of the present invention, the extraction module 902 is specifically used to perform frame-segmentation processing on the complete audio, input the frame-segmented complete audio into a pre-trained speech recognition model, and extract the output of the feature extraction module in the speech recognition model as the audio features of each audio frame in the complete audio; wherein, the speech recognition model includes a feature extraction module and a text recognition module, the feature extraction module is used to convert audio into audio features, and the text conversion module is used to recognize the audio features obtained by the feature extraction module as text;

[0171] The incomplete audio is segmented into frames. The segmented incomplete audio is then input into a speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio features of each audio frame in the incomplete audio.

[0172] This invention also provides an electronic device, such as... Figure 10 As shown, it includes a processor 101, a communication interface 102, a memory 103, and a communication bus 104, wherein the processor 101, the communication interface 102, and the memory 103 communicate with each other through the communication bus 104.

[0173] Memory 103 is used to store computer programs;

[0174] When processor 101 executes a program stored in memory 103, it performs the following steps:

[0175] Retrieve the full and incomplete audio versions; the incomplete audio version is missing audio segments that need to be located compared to the full audio version.

[0176] Extract the audio features of each audio frame from both the complete and incomplete audio versions;

[0177] Using the number of first audio frames as the segment length, one or more first audio segments are sequentially determined in the complete audio version, and one or more second audio segments are sequentially determined in the incomplete audio version. The first and second audio segments in the same order are matched sequentially until the i-th first audio segment and the i-th second audio segment fail to match. The starting coordinates of the differing audio segment in the complete audio version are determined based on the coordinates of the audio frames in the i-th first audio segment, and the coordinates of the non-aligned point in the incomplete audio version are determined based on the coordinates of the audio frames in the i-th second audio segment. The successful matching of the first audio segment and the second audio segment indicates that the similarity between the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment meets the preset conditions.

[0178] Using the number of second audio frames as the segment length, a third audio segment is determined in the incomplete audio after the non-aligned point coordinates, and one or more fourth audio segments are determined sequentially in the complete audio after the starting coordinates. The third audio segment and one or more fourth audio segments are matched sequentially until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment.

[0179] Based on the start and end coordinates of the differing audio segments in the full audio version, the position range of the differing audio segments in the full audio version is located.

[0180] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0181] The communication interface is used for communication between the aforementioned terminal and other devices.

[0182] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0183] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0184] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the audio difference localization methods described in the above embodiments.

[0185] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the audio difference localization methods described in the above embodiments.

[0186] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0187] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0188] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0189] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An audio difference localization method, characterized in that, include: Obtain the complete audio and the incomplete audio; the incomplete audio is missing a difference audio segment that needs to be located compared to the complete audio. Extract the audio features of each audio frame from the complete audio and the incomplete audio; Using the number of first audio frames as the segment length, one or more first audio segments are sequentially determined in the complete audio, and one or more second audio segments are sequentially determined in the incomplete audio. First and second audio segments of the same order are then matched sequentially until the i-th first audio segment and the i-th second audio segment fail to match. The starting coordinates of the differing audio segment in the complete audio are determined based on the coordinates of the audio frames in the i-th first audio segment, and the coordinates of the misaligned point in the incomplete audio are determined based on the coordinates of the audio frames in the i-th second audio segment. A successful match between the first and second audio segments indicates that the similarity between the audio features corresponding to the first and second audio segments meets a preset condition. The misaligned point coordinates are the boundary points between the second audio segments in the incomplete audio. For each misaligned point: a second audio segment in the incomplete audio preceding this point can be successfully matched with a first audio segment of the same order in the complete audio; a second audio segment in the incomplete audio following this point cannot be successfully matched with a first audio segment of the same order in the complete audio. Using the number of second audio frames as the segment length, a third audio segment is determined in the incomplete audio following the non-aligned point coordinates, and one or more fourth audio segments are sequentially determined in the complete audio following the starting coordinates. The third audio segment and one or more fourth audio segments are sequentially matched until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment. Based on the start and end coordinates of the differing audio segment in the full audio version, the position range of the differing audio segment in the full audio version is located.

2. The method according to claim 1, characterized in that, The following method is used to determine whether the first audio segment and the second audio segment match successfully: Cross-correlation calculation is performed on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain a cosine similarity sequence. It is then determined whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence. If so, it is confirmed that the first audio segment and the second audio segment are successfully matched; If not, it indicates that the current first and second audio segments have not matched successfully.

3. The method according to claim 1, characterized in that, The steps of determining the starting coordinates of the differing audio segment in the complete audio version based on the position of the i-th first audio segment, and determining the coordinates of the misaligned point in the incomplete audio version based on the position of the i-th second audio segment, include: Using the number of third audio frames as the shortening step, the i-th first audio segment and the i-th second audio segment are shortened simultaneously, and the shortened first audio segment and the shortened second audio segment are matched until the shortened first audio segment and the shortened second audio segment are successfully matched. The coordinates of the last audio frame in the successfully matched i-th first audio segment are used as the starting coordinates of the difference audio segment in the complete version of the audio, and the coordinates of the last audio frame in the successfully matched i-th second audio segment are used as the coordinates of the non-aligned point in the incomplete version of the audio.

4. The method according to claim 1, characterized in that, Also includes: The complete audio following the termination coordinates is identified as the new complete audio, and the incomplete audio following the non-alignment point coordinates is identified as the new incomplete audio. The process returns to the steps of using the number of first audio frames as the segment length, sequentially identifying one or more first audio segments in the complete audio, sequentially identifying one or more second audio segments in the incomplete audio, and sequentially matching the first audio segments and second audio segments in the same order until both the first audio segments and the second audio segments are successfully matched.

5. The method according to claim 1, characterized in that, The complete audio is the audio corresponding to the submitted video, and the incomplete audio is the audio corresponding to the approved video. The method further includes: Based on the location range of the differential audio segments in the full audio version, the dubbing file obtained by dubbing the submitted video is trimmed to obtain a dubbing file adapted to the approved video version.

6. The method according to claim 1, characterized in that, The step of extracting audio features from each audio frame in the complete audio and the incomplete audio specifically includes: The complete audio is segmented into frames, and the segmented audio is input into a pre-trained speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio features of each audio frame in the complete audio. The speech recognition model includes a feature extraction module and a text recognition module. The feature extraction module is used to convert audio into audio features, and the text recognition module is used to recognize the audio features obtained by the feature extraction module as text. The incomplete audio is segmented into frames, and the segmented incomplete audio is input into the speech recognition model. The output of the feature extraction module in the speech recognition model is extracted as the audio feature of each audio frame in the incomplete audio.

7. An audio difference localization device, characterized in that, include: The acquisition module is used to acquire the complete audio and the incomplete audio; the incomplete audio is missing a difference audio segment to be located relative to the complete audio. An extraction module is used to extract the audio features of each audio frame in the complete audio and the incomplete audio. The first matching module is used to determine one or more first audio segments in the complete audio and one or more second audio segments in the incomplete audio, using the number of first audio frames as the segment length. It then sequentially matches the first and second audio segments in the same order until the i-th first and second audio segments fail to match. Based on the coordinates of the audio frames in the i-th first audio segment, it determines the starting coordinates of the differing audio segment in the complete audio. Based on the coordinates of the audio frames in the i-th second audio segment, it determines the coordinates of the misaligned point in the incomplete audio. A successful match between the first and second audio segments indicates that the similarity between the audio features corresponding to the first and second audio segments meets a preset condition. The misaligned point coordinates are the boundary points between the second audio segments in the incomplete audio. For each misaligned point: the second audio segment before the point in the incomplete audio can be successfully matched with the first audio segment in the same order in the complete audio; the second audio segment after the point in the incomplete audio cannot be successfully matched with the first audio segment in the same order in the complete audio. The second matching module is used to determine a third audio segment in the incomplete audio after the non-alignment point coordinates, using the number of second audio frames as the segment length, and to sequentially determine one or more fourth audio segments in the complete audio after the starting coordinates. The third audio segment and one or more fourth audio segments are matched sequentially until a fourth audio segment that successfully matches the third audio segment is determined. The termination coordinates of the difference audio segment in the complete audio are determined based on the coordinates of the audio frames in the successfully matched fourth audio segment. The positioning module is used to locate the position range of the differing audio segment in the full version of the audio based on the start and end coordinates of the differing audio segment in the full version of the audio.

8. The apparatus according to claim 7, characterized in that, The first matching module includes: The judgment unit is used to perform cross-correlation calculation on the audio features corresponding to the first audio segment and the audio features corresponding to the second audio segment in the same order to obtain a cosine similarity sequence, and to determine whether the coordinates of the similarity peak in the cosine similarity sequence are the center coordinates of the cosine similarity sequence. If so, it is confirmed that the first audio segment and the second audio segment are successfully matched; If not, it indicates that the current first and second audio segments have not matched successfully.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data processing method, device and equipment

    CN114845164A

  • Audio cutting method and device

    CN116612784A

  • Differential detection apparatus, differential detection method and differential detection program

    JP2015141602A