Audio processing method and device, electronic equipment and storage medium

By selecting audio segments from preset audio files as singing material in a singing application, determining and comparing the overlap between the target audio segment and the reference audio segment, and eliminating the overlapping parts, the overlap problem in audio recording in the singing application is solved, achieving smooth audio recording and saving computing performance.

CN119920221BActive Publication Date: 2026-08-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311436315.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2026-08-25
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

In singing apps, there is a problem of repeated recording during the audio recording process, which makes the singing process time-consuming, laborious, and costly, and makes it impossible to record audio smoothly, with overlapping audio recordings.

Method used

By selecting audio segments from preset audio files as singing material, determining the attribute information of target audio segments, and directly selecting audio segments from preset audio files as singing material to continue the next singing session without waiting for audio processing after a performance, multiple target audio segments are formed. By comparing the overlap between target audio segments and reference audio segments, overlapping parts are marked and removed to achieve non-overlapping audio merging.

Benefits of technology

It achieves a smooth singing process in singing applications, reduces stuttering, saves computing power consumption, and avoids audio overlap problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920221B_ABST
    Figure CN119920221B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio processing method and device, electronic equipment and a storage medium. The method comprises: determining a plurality of target audio segments; determining target audio attribute information of each target audio segment, the target audio attribute information being used to mark available audio segments in the target audio segment, the available audio segment satisfying the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment being collected after the target audio segment, and the target audio segment and the reference audio segment being audio segments formed by using a same preset audio file as singing material; and processing the plurality of target audio segments based on the target audio attribute information of each target audio segment. The present scheme solves the problem that audio recording cannot be performed smoothly in a singing application and the recorded audio has overlap.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to data processing technology, and more particularly to an audio processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of the Internet, various singing apps have gradually entered people's field of vision and become popular. At the same time, people's requirements for the singing experience of these apps are getting higher and higher.

[0003] In related solutions, such as singing applications, it is necessary to collect audio data generated during singing, and then determine whether the collected audio data needs to be adjusted. If adjustment is needed, it will be re-recorded. This may result in repeated recording of audio data, and each time it needs to be re-recorded, making the singing process time-consuming, laborious, and costly, and causing great inconvenience to the singing recording process. Summary of the Invention

[0004] This disclosure provides an audio processing method, apparatus, electronic device, and storage medium to solve the problem of inability to smoothly record audio in singing applications and the overlap of recorded audio.

[0005] In a first aspect, embodiments of this disclosure provide an audio processing method, the method comprising:

[0006] Identify multiple target audio segments;

[0007] Determine the target audio attribute information for each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the available audio segments and the reference audio segments. The reference audio segments are collected after the target audio segments and are audio segments formed by using the same preset audio file as the singing material as the target audio segments.

[0008] Based on the target audio attribute information of each target audio segment, the plurality of target audio segments are processed.

[0009] Secondly, embodiments of this disclosure also provide an audio processing apparatus, the apparatus comprising:

[0010] The first determining module is used to determine multiple target audio segments;

[0011] The second determining module is used to determine the target audio attribute information of each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the available audio segments and the reference audio segments. The reference audio segments are collected after the target audio segments and are audio segments formed by using the same preset audio file as the singing material as the target audio segments.

[0012] An audio processing module is used to process the plurality of target audio segments based on the target audio attribute information of each target audio segment.

[0013] Thirdly, this disclosure also provides an electronic device, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the audio processing method described in any of the above embodiments.

[0017] Fourthly, this disclosure also provides a computer-readable medium storing computer instructions that cause a processor to execute the audio processing method described in any of the above embodiments.

[0018] In this embodiment, an audio segment is selected from a preset audio file in a singing application as the singing material. After one performance, the audio segment is used directly as the singing material without waiting for audio processing. The singing material is then used to continue the next performance, forming an audio clip. By selecting singing materials multiple times, multiple target audio segments are obtained. This allows for smooth singing in the singing application without waiting for audio processing to finish, reducing lag in the application to some extent. After obtaining multiple target audio segments, it can be determined whether each target audio segment has a corresponding reference audio segment. If a reference audio segment exists, the target audio segment is compared with the reference audio segment to determine the target audio attribute information of the available audio segments that do not overlap with the reference audio segment. Then, the uniformly marked target audio segments are merged without overlap according to the target audio attribute information of each target audio segment. Since no audio merging operation is performed when uniformly obtaining the target audio segments, the overlapping parts of the audio segments do not need to be processed, saving some computational performance.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0021] Figure 1 This is a schematic flowchart of an audio processing method provided in an embodiment of this disclosure;

[0022] Figure 2 This is a schematic diagram of the interface of a singing application to which this disclosure applies;

[0023] Figure 3a This is a schematic diagram showing audio segments from different acquisition times mapped onto a timeline, applicable to the embodiments of this disclosure.

[0024] Figure 3b This is a schematic diagram illustrating the merging of audio segments from different acquisition times applicable to the embodiments of this disclosure;

[0025] Figure 4 This is a schematic flowchart of another audio processing method provided in an embodiment of this disclosure;

[0026] Figure 5 This is a schematic flowchart of another audio processing method provided in the embodiments of this disclosure;

[0027] Figure 6 This is a detailed breakdown diagram illustrating the merging of audio segments from different acquisition times, as provided in the embodiments of this disclosure.

[0028] Figure 7 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this disclosure;

[0029] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the audio processing method according to an embodiment of the present disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0040] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0041] Figure 1 This is a flowchart illustrating an audio processing method provided in an embodiment of the present disclosure. This embodiment is applicable to situations where singing is performed continuously in a singing application and the audio generated from the singing is merged. The method can be executed by an audio processing device, which can be implemented in software and / or hardware and is generally integrated into any electronic device with network communication capabilities, such as a mobile terminal, a PC, or a server.

[0042] like Figure 1 As shown, the audio processing method of this disclosure embodiment may include the following processes:

[0043] S110, Identify multiple target audio segments.

[0044] S120. Determine the target audio attribute information for each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the reference audio segment and the available audio segment; the reference audio segment is collected after the target audio segment; and both the reference audio segment and the target audio segment are formed by using the same preset audio file as the singing material.

[0045] See Figure 2 The following are schematic diagrams of the "song in-between page," "singing recording page," and "preview page" in the singing application. Figure 2 The "Singing Recording Page" shown allows users to repeatedly practice singing by selecting audio segments from the same preset audio file provided by the singing application at different recording times. The audio segments selected from the preset audio file at each recording time can be the same, partially the same, or different. By singing using the selected audio segments at different recording times and recording the resulting audio segments, multiple target audio segments are obtained. Furthermore, the audio segments selected from the preset audio file at each recording time may overlap. Therefore, the newly formed audio segments created by singing using the selected audio segments as singing material will have overlapping portions, meaning that many overlapping audio segments may occur between the various target audio segments, leading to unclear audio after merging.

[0046] When singing in a karaoke app, the audio segments formed when singing the same audio segment selected from a preset audio file at different capture times will only be re-sung at the next capture time if the user is not satisfied with the newly formed audio segment from the previous capture time. This process continues until the final audio segment is satisfactory. It is evident that the usability of audio segments captured later is stronger than that of audio segments captured earlier. Therefore, the audio segments formed when singing the same audio segment are usually dominated by those captured later. At the same time, some audio segments captured earlier are not reflected in the audio segments captured later. Therefore, audio segments captured earlier cannot be directly eliminated; instead, they must be selectively eliminated and retained.

[0047] Therefore, for each target audio segment, it can be determined whether there exists a series of audio segments whose acquisition time is later than the target audio segment. If so, the series of audio segments whose acquisition time is later than the target audio segment are selected as the reference audio segments corresponding to the target audio segment; if not, it is determined that there is no corresponding reference audio segment for the target audio segment. The reference audio segment corresponding to the target audio segment is an audio segment formed by singing audio segments selected from the same preset audio file, whose acquisition time is later than the target audio segment. This allows for the use of reference audio segments with acquisition times later than the target audio segment to correct and remove overlapping, unacceptable audio segments within the target audio segment.

[0048] See Figure 3a Audio segments selected from the same preset audio file at different acquisition times are used as singing material to create new audio segments, which are sequentially named the audio segment of acquisition 1, the audio segment of acquisition 2, and the audio segment of acquisition 3. These are denoted as the multiple target audio segments that need to be merged in the above-mentioned process. Simultaneously, for multiple target audio segments, a reference audio segment can be selected from the audio segments acquired at a later time than the target audio segment.

[0049] Figure 3a The audio segments from acquisition 1, acquisition 2, and acquisition 3 shown can all be used as target audio segments in sequence. When the target audio segment is the audio segment from acquisition 1, the reference audio segments corresponding to the audio segment from acquisition 1 are, in sequence, the audio segments from acquisition 2 and acquisition 3. Figure 3a As can be seen, there are some overlapping audio segments between the audio segments of Acquisition 2 and Acquisition 1. Since the audio segments of Acquisition 2 were acquired later than those of Acquisition 1, the overlapping audio segments of the audio segments of Acquisition 1, as the target audio segments, need to be removed and discarded during the subsequent merging process.

[0050] See Figure 3a Similarly, audio segments from acquisition 3 and acquisition 1 partially overlap. Since acquisition 3 was acquired later than acquisition 1, the overlapping segments of acquisition 1, being the target audio segment, also need to be discarded during subsequent merging. See also Figure 3b When the audio segment of Acquisition 1 is used as the target audio segment, the overlapping parts of the audio segment of Acquisition 1 with the audio segments of Acquisition 2 and Acquisition 3 are discarded. The audio segment of Acquisition 1 is then divided into two audio segments corresponding to Acquisition 1 as usable audio segments.

[0051] Therefore, when a reference audio segment corresponding to the target audio segment is obtained, the audio segments in the target audio segment that do not overlap with the reference audio segment can be considered as usable audio segments based on the overlap between the target audio segment and the corresponding reference audio segment, and the positions of these usable audio segments in the target audio segment can be marked. Conversely, when a reference audio segment corresponding to the target audio segment is not obtained, the entire target audio segment can be considered as a usable audio segment, and the positions of the usable audio segments within the target audio segment can be marked. Marking the positions of the usable audio segments within the target audio segment ensures that these usable audio segments are retained during subsequent merging of multiple target audio segments.

[0052] S130. Process multiple target audio segments based on the target audio attribute information of each target audio segment.

[0053] For each of multiple target audio segments, the target audio attribute information can be used to mark the available audio segments in each target audio segment that do not overlap with the reference audio segments corresponding to the target audio segment. The marked target audio segments can be saved independently. In this way, when selecting to merge audio segments in the future, the available audio segments in the target audio segments can be directly used for audio merging, thereby eliminating duplicate audio segments in the merged audio and solving the problem of audio recording not being able to be performed smoothly and the recorded audio having overlap in the singing application.

[0054] Optionally, available audio segments that do not overlap with reference audio segments corresponding to the target audio segment are selected from each target audio segment. Then, the selected available audio segments are decoded and merged to obtain the merged audio file.

[0055] As an optional but not limited implementation, processing multiple target audio segments based on the target audio attribute information of each target audio segment may include the following steps A1-A2:

[0056] Step A1: For each target audio segment, based on the target audio attribute information of the target audio segment, cut out the usable audio segments from the target audio segment and determine the audio starting point corresponding to the usable audio segments. The usable audio segments are the audio segments selected from the target audio segments that can be used for audio merging. The audio starting point indicated by the target audio attribute information is based on the audio progress time of the usable audio segments in the preset audio file.

[0057] Step A2: According to the order of the audio starting points corresponding to the available audio segments, merge the available audio segments cut from each target audio segment.

[0058] The target audio attribute information indicates the starting point of the available audio segments that can be used for audio merging within the target audio segments. This starting point can be mapped to the corresponding audio progress time in a preset audio file. After selecting the available audio segments for merging from the target audio segments, the selected segments can be decoded. The decoded segments are then concatenated according to the chronological order of their corresponding starting points to obtain the merged audio file. The process of selecting available audio segments from the target audio segments can be implemented using the AVComposition library or a third-party audio trimming library.

[0059] The solution of this embodiment selects audio segments from a preset audio file as singing material in a singing application. After one performance, without waiting for audio processing, the selected audio segments are used directly as singing material for the next performance, thus obtaining multiple target audio segments. This allows for smooth singing in the singing application without waiting for audio processing to complete, reducing lag to some extent. After obtaining multiple target audio segments, reference audio segments can be determined for each target audio segment. By comparing the target audio segments with the reference audio segments, target audio attribute information can be determined to mark usable audio segments in the target audio segments that do not overlap with the reference audio segments. Then, according to the target audio attribute information of each target audio segment, the uniformly marked target audio segments are merged without overlap. Since no audio merging operation is performed when uniformly obtaining target audio segments, overlapping parts in the audio segments do not need to be processed, saving some computational performance.

[0060] Figure 4 This is a flowchart illustrating another audio processing method provided by an embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining multiple target audio segments in the foregoing embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments.

[0061] like Figure 4 As shown, the audio processing method of this disclosure embodiment may include the following processes:

[0062] S410. Determine multiple candidate audio segments. Multiple candidate audio segments are newly formed audio segments based on audio segments selected from the same preset audio file at different acquisition times as singing materials.

[0063] Optionally, when the audio segments selected from the same preset audio file at different acquisition times are the same or partially the same, and when the audio segments selected from the same preset audio file at different acquisition times are the same or partially the same, the multiple candidate audio segments are at least partially the same. The multiple candidate audio segments are audio segments formed by pronouncing the audio segments selected from the same preset audio file as singing material at different acquisition times.

[0064] See Figure 3a and Figure 3b When practicing singing using a singing app, overlapping areas may occur when multiple target audio segments are captured. This is because, firstly, the selected audio segments are repeatedly practiced and recorded. For example, if an audio segment doesn't meet the singing requirements, it needs repeated practice and recording. In this case, the audio segments selected from the same preset audio file at different capture times are identical. Secondly, the final audio needs to be shortened and spliced. If the audio segments cannot be connected, even if spliced ​​into a single audio file, incomplete audio will occur. To continue the previously captured audio segments, the ending portion of the previous audio segment is repeated for practice. Therefore, the audio segments selected at different capture times are partially identical.

[0065] As an optional but not limited implementation, determining multiple candidate audio segments may include the following steps B1-B2:

[0066] Step B1: In response to the audio acquisition operation, an audio segment based on preset singing material is acquired to obtain a candidate audio segment and the corresponding acquisition time is marked. The preset singing material is the singing material corresponding to the audio segment selected from the preset audio file.

[0067] See Figure 2 In the singing app's recording page, clicking the recording control initiates audio capture. At this point, audio clips from a preset audio file can be used as singing material, and the audio clips generated during the performance can be captured in real-time by the audio capture tool. Clicking the recording control again ends the audio capture operation. The audio clips captured between the start and pause of the recording operation are newly generated candidate audio clips from the preset audio file, and their capture time is marked, which can also be understood as the recording time of the candidate audio clips. The audio capture tool can be implemented using the AudioUnit library or a third-party audio capture library.

[0068] When practicing singing using a singing app, there may be situations where the selected audio segment doesn't meet the singing requirements and needs to be repeatedly practiced and recorded. Alternatively, to continue singing from a previous segment, the end of the previously sung audio segment may need to be practiced again. To address this, in response to another audio capture operation, the audio segment selected from a preset audio file is captured to obtain the next candidate audio segment, and the corresponding capture time is marked. In other words, after capturing one candidate audio segment, the next singing session can begin. Similarly, the audio segment formed during the period from the start to the end of the audio capture operation needs to be captured to obtain another candidate audio segment, and the capture time of the candidate audio segment is marked. If the audio segments selected from the same preset audio file at different capture times are the same or partially the same, then the resulting candidate audio segments will at least partially overlap.

[0069] Step B2: Determine multiple candidate audio segments corresponding to multiple audio acquisition operations, and use the same preset audio file for multiple audio acquisition operations.

[0070] It is only necessary to merge multiple audio segments and remove overlapping audio segments when the audio segments selected as singing material from different acquisition times originate from the same preset audio file. If they do not originate from the same audio file, then merging multiple audio segments and removing overlapping audio segments is not so necessary.

[0071] S420. Identify multiple target audio segments from multiple candidate audio segments.

[0072] The reference audio segment is acquired later than the target audio segment. Both the target and reference audio segments are newly formed based on audio segments selected from the same preset audio file as singing material. The reference audio segment is acquired after the target audio segment, and both the reference and target audio segments use the same preset audio file as singing material.

[0073] Optionally, all candidate audio segments from multiple candidate audio segments can be used as multiple target audio segments, with each candidate audio segment serving as a target audio segment. Alternatively, a subset of candidate audio segments can be selected from multiple candidate audio segments as multiple target audio segments, in which case each selected candidate audio segment can serve as a target audio segment.

[0074] As an optional but not limited implementation, determining multiple target audio segments from multiple candidate audio segments may include the following steps:

[0075] Step C1: Sort the candidate audio segments in the multiple candidate audio segments according to the order of their acquisition time. Then, select the candidate audio segments in sequence according to the sorting results and determine them as multiple target audio segments until all candidate audio segments or selected candidate audio segments are traversed.

[0076] S430. Determine the target audio attribute information for each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the reference audio segment and the available audio segment; the reference audio segment is collected after the target audio segment; and both the reference audio segment and the target audio segment are formed by using the same preset audio file as the singing material.

[0077] S440. Based on the target audio attribute information of each target audio segment, process the multiple target audio segments.

[0078] The solution of this embodiment selects audio segments from a preset audio file as singing material in a singing application. After one performance, the audio segments are used directly as singing material without waiting for audio processing. The singing material is then used to continue singing to form another audio segment. By selecting singing material multiple times, multiple target audio segments are obtained. This allows for smooth singing in the singing application without waiting for audio processing to finish, which can reduce lag in the singing application to a certain extent. After obtaining multiple target audio segments, it can be determined whether each target audio segment has a corresponding reference audio segment. If a reference audio segment exists, the target audio segment is compared with the reference audio segment to determine the target audio attribute information of the available audio segments that do not overlap with the reference audio segment. Then, according to the target audio attribute information of each target audio segment, the uniformly marked target audio segments are merged without overlap. Since no audio merging operation is performed when uniformly obtaining the target audio segments, the overlapping parts of the audio segments do not need to be processed, which can save some computational performance. Figure 5 This is a flowchart illustrating another audio processing method provided in this disclosure. The technical solution of this embodiment further optimizes the process of determining the target audio attribute information of each target audio segment in the foregoing embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments.

[0079] like Figure 5 As shown, the audio processing method of this disclosure embodiment may include the following processes:

[0080] S510, Identify multiple target audio segments.

[0081] S520. Determine the first audio attribute information of the target audio segment. The first audio attribute information is used to indicate the audio start point and audio duration of the target audio segment. The audio start point indicated by the first audio attribute information is based on the audio progress time representation of the target audio segment in the preset audio file.

[0082] The first audio attribute information indicates the audio start point and audio duration of the target audio segment. The audio start point indicated by the first audio attribute information can be the starting point of the target audio segment mapped to the corresponding audio progress time in a preset audio file.

[0083] S530. Determine the reference audio segment corresponding to the target audio segment, and determine the second audio attribute information of the reference audio segment. The second audio attribute information is used to indicate the audio start point and audio duration of the reference audio segment. The audio start point indicated by the second audio attribute information is based on the audio progress time representation of the reference audio segment in the preset audio file. The audio start point indicated by the second audio attribute information can refer to the starting point of the reference audio segment mapped to the corresponding audio progress time in the preset audio file.

[0084] S540. Based on the first audio attribute information and the second audio attribute information, determine the target audio attribute information of the target audio segment. The target audio attribute information is used to indicate whether there is a usable audio segment and the audio start point and audio duration of at least one usable audio segment. The usable audio segment is an audio segment selected from the target audio segment that can be used for audio merging.

[0085] The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the available audio segments and the reference audio segments; the reference audio segments are collected after the target audio segments; and both the reference audio segments and the target audio segments are formed by using the same preset audio file as the singing material.

[0086] Optionally, the target audio attribute information is used to mark the available audio segments in the target audio segment that do not overlap with the reference audio segment when the reference audio segment exists, and all of the target audio segment when the reference audio segment does not exist, as available audio segments.

[0087] The target audio attribute information indicates whether a usable audio segment exists within the target audio segment, and, if a usable audio segment exists within the target hidden segment, the starting point of the usable audio segment that can be used for audio merging. The starting point indicated by the target audio attribute information can be the starting point of the usable audio segment mapped to the corresponding audio progress time in a preset audio file.

[0088] As an optional but not limited implementation, determining the reference audio segment corresponding to the target audio segment includes the following steps C1-C2:

[0089] Step C1: Sort the target audio segments according to the order of their acquisition times.

[0090] Step C2: For each target audio segment, determine the target audio segment that is sorted after the target audio segment as the reference audio segment corresponding to the target audio segment.

[0091] As an optional but not limited implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps D1-D2:

[0092] Step D1: If, based on the first audio attribute information and the second audio attribute information, it is determined that the starting point of the target audio segment is less than the starting point of the reference audio segment and there are overlapping audio segments, then the starting point indicated by the first audio attribute information is determined as the starting point of the available audio segment indicated by the target audio attribute information.

[0093] Step D2: Subtract the audio start point indicated by the second audio attribute information from the audio start point indicated by the first audio attribute information, and determine the audio duration of the available audio segment indicated by the target audio attribute information.

[0094] See Figure 6Starting from the left, taking the audio segment of acquisition 1 as the target audio segment and the audio segment of acquisition 2 as the reference audio segment as an example, the starting point of the first audio attribute information of the target audio segment is less than the starting point of the second audio attribute information of the reference audio segment, and the sum of the starting point of the first audio attribute information and the audio duration is greater than the starting point of the second audio attribute information of the reference audio segment. It can be seen that there is an overlapping audio segment between the right part of the audio segment of acquisition 1 and the left part of the audio segment of acquisition 2. The right part of the audio segment of acquisition 1 can be segmented to the right. At this time, the starting point of the first audio attribute information can be directly determined as the starting point of the usable audio segment, and the difference between the starting point of the second audio attribute information and the starting point of the first audio attribute information can be determined as the audio duration of the usable audio segment.

[0095] As another optional but non-limiting implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps E1-E3:

[0096] Step E1: If, based on the first audio attribute information and the second audio attribute information, it is determined that the starting point of the target audio segment is greater than the starting point of the reference audio segment and there are overlapping audio segments, then the sum of the starting point indicated by the second audio attribute information and the duration indicated by the second audio attribute information is determined as the starting point of the available audio segment indicated by the target audio attribute information.

[0097] Step E2: Subtract the audio starting point indicated by the target audio attribute information from the audio starting point indicated by the first audio attribute information to determine the duration of the overlapping audio segment.

[0098] Step E3: Subtract the duration of the audio segment indicated by the first audio attribute information from the duration of the overlapping audio segment to determine the duration of the available audio segment indicated by the target audio attribute information.

[0099] See Figure 6Taking the second example from the left, with audio segment 1 as the target audio segment and audio segment 2 as the reference audio segment, the starting point indicated by the first audio attribute information of the target audio segment is greater than the starting point indicated by the second audio attribute information of the reference audio segment. Furthermore, the sum of the starting point indicated by the second audio attribute information and the audio duration is greater than the starting point indicated by the first audio attribute information of the target audio segment. Therefore, it can be concluded that there is an overlapping audio segment between audio segments 1 and 2; specifically, the left portion of audio segment 1 and the right portion of audio segment 2 overlap. In the case of overlap, the left part of the audio segment of acquisition 1 can be split to the left. At this time, the difference between the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information can be determined as the duration of the overlapping audio segment between the right part of the audio segment of acquisition 1 and the left part of the audio segment of acquisition 2. Then, the difference between the audio duration indicated by the first audio attribute information and the duration of the overlapping audio segment can be determined as the audio duration of the usable audio segment. And the sum of the audio starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information can be determined as the audio starting point of the usable audio segment.

[0100] As another optional but not limited implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps F1-F5:

[0101] Step F1: If, based on the first audio attribute information and the second audio attribute information, it is determined that the audio start point of the target audio segment is less than the audio start point of the reference audio segment, and the sum of the audio start point and the audio duration of the target audio segment is greater than the sum of the audio start point and the audio duration of the reference audio segment, then the audio start point indicated by the first audio attribute information is determined as the audio start point of the first available audio segment indicated by the target audio attribute information.

[0102] Step F2: Subtract the audio start point indicated by the second audio attribute information from the audio start point indicated by the first audio attribute information, and determine the audio duration of the first available audio segment indicated by the target audio attribute information.

[0103] Step F3: The sum of the audio start point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio start point of the second available audio segment indicated by the target audio attribute information.

[0104] Step F4: Subtract the audio start point indicated by the target audio attribute information from the audio start point indicated by the first audio attribute information to determine the duration of the overlapping audio segment.

[0105] Step F5: Subtract the duration of the first audio attribute information from the duration of the overlapping audio segment to determine the duration of the second available audio segment indicated by the target audio attribute information.

[0106] The first available audio segment and the second available audio segment are two audio segments selected from the target audio segment that can be used for audio merging.

[0107] See Figure 6 Starting from the left, the third example uses the audio segment from acquisition 1 as the target audio segment and the audio segment from acquisition 2 as the reference audio segment. The overlapping audio segment of the audio segments from acquisition 1 and acquisition 2 is located in the middle part of the audio segment from acquisition 1. At this time, the process of steps F1-F5 can be achieved by combining the right-side segmentation of steps D1-D2 and the left-side segmentation of steps E1-E3. The first usable audio segment and the second usable audio segment are sequentially segmented from the left and right sides of the audio segment from acquisition 1. The first usable audio segment and the second usable audio segment are two audio segments selected from the target audio segment that can be used for audio merging.

[0108] As an optional but not limited implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps:

[0109] If, based on the first and second audio attribute information, it is determined that the starting point of the target audio segment is greater than the starting point of the reference audio segment, and the sum of the starting point and duration of the target audio segment is less than the sum of the starting point and duration of the reference audio segment, then the target audio attribute information of the target audio segment is determined to be that there is no usable audio segment in the target audio segment.

[0110] See Figure 6 The fourth one from the left in the middle, taking the audio segment of acquisition 1 as the target audio segment and the audio segment of acquisition 2 as the reference audio segment as an example, if the audio segment of acquisition 1 is completely covered by the audio segment of acquisition 2, then the audio segment of acquisition 1 is all overlapping audio segments. At this time, the audio segment of acquisition 1 can be overwritten and deleted, and it can be considered that there are no usable audio segments in the target audio segment.

[0111] S550: Process multiple target audio segments based on the target audio attribute information of each target audio segment.

[0112] The solution of this embodiment selects audio segments from a preset audio file as singing material in a singing application. After one performance, without waiting for audio processing, the selected audio segments are used directly as singing material for the next performance, thus obtaining multiple target audio segments. This allows for smooth singing in the singing application without waiting for audio processing to complete, reducing lag to some extent. After obtaining multiple target audio segments, reference audio segments can be determined for each target audio segment. By comparing the target audio segments with the reference audio segments, target audio attribute information can be determined to mark usable audio segments in the target audio segments that do not overlap with the reference audio segments. Then, according to the target audio attribute information of each target audio segment, the uniformly marked target audio segments are merged without overlap. Since no audio merging operation is performed when uniformly obtaining target audio segments, overlapping parts in the audio segments do not need to be processed, saving some computational performance.

[0113] Figure 7 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation where singing is performed continuously in a singing application and the audio generated by the singing is merged. The audio processing device can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication function, such as a mobile terminal, a PC, or a server.

[0114] like Figure 7 As shown, the audio processing apparatus of this embodiment may include: a first determining module 710, a second determining module 720, and an audio processing module 730. Wherein:

[0115] The first determining module 710 is used to determine multiple target audio segments;

[0116] The second determining module 720 is used to determine the target audio attribute information of each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no reference audio segment that overlaps with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are formed by using the same preset audio file as the singing material.

[0117] The audio processing module 730 is used to process the plurality of target audio segments based on the target audio attribute information of each target audio segment.

[0118] Based on the above embodiments, optionally, multiple target audio segments are determined, including:

[0119] Multiple candidate audio segments are identified, which are newly formed audio segments based on audio segments selected from the same preset audio file at different acquisition times as singing material;

[0120] Identify multiple target audio segments from a pool of candidate audio segments.

[0121] Based on the above embodiments, optionally, determining multiple candidate audio segments includes:

[0122] In response to the audio acquisition operation, an audio segment based on preset singing material is acquired to obtain a candidate audio segment and the corresponding acquisition time is marked. The preset singing material is the singing material corresponding to the audio segment selected from the preset audio file.

[0123] The candidate audio segments corresponding to multiple audio acquisition operations are determined as the multiple candidate audio segments, and the preset audio file used in the multiple audio acquisition operations is the same audio file.

[0124] Optionally, based on the above embodiments, the audio segments selected from the same preset audio file at different acquisition times are the same or partially the same.

[0125] Based on the above embodiments, optionally, the target audio attribute information of each target audio segment is determined, including:

[0126] The first audio attribute information of the target audio segment is determined. The first audio attribute information is used to indicate the audio start point and audio duration of the target audio segment. The audio start point indicated by the first audio attribute information is based on the audio progress time of the target audio segment in a preset audio file.

[0127] Determine the reference audio segment corresponding to the target audio segment, and determine the second audio attribute information of the reference audio segment. The second audio attribute information is used to indicate the audio start point and audio duration of the reference audio segment. The audio start point indicated by the second audio attribute information is based on the audio progress time of the reference audio segment in a preset audio file.

[0128] Based on the first audio attribute information and the second audio attribute information, the target audio attribute information of the target audio segment is determined. The target audio attribute information is used to indicate whether there is a usable audio segment and the audio start point and audio duration of at least one usable audio segment. The usable audio segment is an audio segment selected from the target audio segments that can be used for audio merging.

[0129] Based on the above embodiments, optionally, determining the reference audio segment corresponding to the target audio segment includes:

[0130] The target audio segments are sorted according to the order of their acquisition times.

[0131] For each target audio segment, the target audio segment that is sorted after the target audio segment is determined as the reference audio segment corresponding to the target audio segment.

[0132] Based on the above embodiments, optionally, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information includes:

[0133] If, based on the first audio attribute information and the second audio attribute information, it is determined that the starting point of the target audio segment is less than the starting point of the reference audio segment and there are overlapping audio segments, then the starting point indicated by the first audio attribute information is determined as the starting point of the available audio segment indicated by the target audio attribute information.

[0134] The difference between the audio start point indicated by the second audio attribute information and the audio start point indicated by the first audio attribute information is used to determine the audio duration of the available audio segment indicated by the target audio attribute information.

[0135] Based on the above embodiments, optionally, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information includes:

[0136] If, based on the first audio attribute information and the second audio attribute information, it is determined that the starting point of the target audio segment is greater than the starting point of the reference audio segment and there are overlapping audio segments, then the sum of the starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the starting point of the available audio segment indicated by the target audio attribute information.

[0137] The difference between the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information is used to determine the duration of the overlapping audio segment.

[0138] The difference between the audio duration indicated by the first audio attribute information and the duration of the overlapping audio segment is determined as the audio duration of the available audio segment indicated by the target audio attribute information.

[0139] Based on the above embodiments, optionally, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information includes:

[0140] If, based on the first audio attribute information and the second audio attribute information, it is determined that the audio start point of the target audio segment is less than the audio start point of the reference audio segment, and the sum of the audio start point and the audio duration of the target audio segment is greater than the sum of the audio start point and the audio duration of the reference audio segment, then the audio start point indicated by the first audio attribute information is determined as the audio start point of the first available audio segment indicated by the target audio attribute information.

[0141] The difference between the audio start point indicated by the second audio attribute information and the audio start point indicated by the first audio attribute information is used to determine the audio duration of the first available audio segment indicated by the target audio attribute information.

[0142] The sum of the audio start point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio start point of the second available audio segment indicated by the target audio attribute information.

[0143] The difference between the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information is used to determine the duration of the overlapping audio segment.

[0144] The difference between the duration of the first audio attribute information indicating the audio duration and the duration of the overlapping audio segment is determined as the duration of the second available audio segment indicated by the target audio attribute information.

[0145] The first available audio segment and the second available audio segment are two audio segments selected from the target audio segment that can be used for audio merging.

[0146] Based on the above embodiments, optionally, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information includes:

[0147] If, based on the first audio attribute information and the second audio attribute information, it is determined that the starting point of the target audio segment is greater than the starting point of the reference audio segment, and the sum of the starting point and duration of the target audio segment is less than the sum of the starting point and duration of the reference audio segment, then the target audio attribute information of the target audio segment is determined to be that there are no usable audio segments in the target audio segment.

[0148] Based on the above embodiments, optionally, the plurality of target audio segments are processed based on the target audio attribute information of each target audio segment, including:

[0149] For each target audio segment, based on the target audio attribute information of the target audio segment, usable audio segments are cut out from the target audio segment and the audio starting point corresponding to the usable audio segment is determined. The usable audio segment is an audio segment selected from the target audio segment that can be used for audio merging. The audio starting point indicated by the target audio attribute information is based on the audio progress time of the usable audio segment in the preset audio file.

[0150] According to the order of the audio starting points corresponding to the available audio segments, the available audio segments cut from each target audio segment are spliced ​​and merged.

[0151] In this embodiment, by singing audio segments selected from a preset audio file in a singing application, and without waiting for audio processing after one performance, the next performance is directly performed from the selected audio segments in the preset audio file, thus obtaining multiple target audio segments. This enables smooth singing in the singing application without waiting for audio processing, which can reduce lag in the singing application to a certain extent. At the same time, after obtaining multiple target audio segments, reference audio segments corresponding to each target audio segment can be determined. By comparing the target audio segments with the reference audio segments, target audio attribute information for marking usable audio segments in the target audio segments that do not overlap with the reference audio segments can be determined. Then, according to the target audio attribute information of each target audio segment, the uniformly marked target audio segments are merged without overlap. Since no audio merging operation is performed when uniformly obtaining target audio segments, some computational performance consumption can be saved.

[0152] The audio processing apparatus provided in this disclosure can execute the audio processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0153] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0154] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 8 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 8The diagram below shows the structure of the terminal device or server 800. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0155] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An edit / output (I / O) interface 805 is also connected to the bus 804.

[0156] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0157] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the audio processing method shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined in the audio processing method of the embodiments of this disclosure.

[0158] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0159] The electronic device provided in this embodiment and the audio processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0160] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the audio processing method provided in the above embodiments.

[0161] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0162] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0163] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0164] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine a plurality of target audio segments; determine target audio attribute information for each target audio segment, the target audio attribute information being used to mark available audio segments within the target audio segments, the available audio segments satisfying the following conditions: no reference audio segment overlaps with the available audio segment, the reference audio segment being acquired after the target audio segment, and both the reference and target audio segments using the same preset audio file as the singing material; and process the plurality of target audio segments based on the target audio attribute information for each target audio segment.

[0165] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0167] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0168] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0169] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0170] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0171] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0172] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An audio processing method, characterized in that, The method includes: Identify multiple target audio segments; Determine the target audio attribute information for each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the available audio segments and the reference audio segments. The reference audio segments are collected after the target audio segments and are newly formed audio segments using the same preset audio file as the singing material as each of the target audio segments. Based on the target audio attribute information of each target audio segment, the plurality of target audio segments are processed; The processing of the plurality of target audio segments based on the target audio attribute information of each target audio segment includes: Based on the target audio attribute information of each target audio segment, the available audio segments corresponding to each target audio segment are determined, and audio merging processing is performed on each of the available audio segments to eliminate overlapping audio segments in the merged audio.

2. The method according to claim 1, characterized in that, Identify multiple target audio segments, including: Multiple candidate audio segments are identified, which are newly formed audio segments based on audio segments selected from the same preset audio file at different acquisition times as singing material; Identify multiple target audio segments from a pool of candidate audio segments.

3. The method according to claim 2, characterized in that, Multiple candidate audio segments were identified, including: In response to the audio acquisition operation, an audio segment based on preset singing material is acquired to obtain a candidate audio segment and the corresponding acquisition time is marked. The preset singing material is the singing material corresponding to the audio segment selected from the preset audio file. The candidate audio segments corresponding to multiple audio acquisition operations are determined as the multiple candidate audio segments, and the preset audio file used in the multiple audio acquisition operations is the same audio file.

4. The method according to claim 2 or 3, characterized in that, The audio segments selected from the same preset audio file at different acquisition times are the same or partially the same.

5. The method according to claim 1, characterized in that, Determine the target audio attribute information for each target audio segment, including: The first audio attribute information of the target audio segment is determined. The first audio attribute information is used to indicate the audio start point and audio duration of the target audio segment. The audio start point indicated by the first audio attribute information is based on the audio progress time of the target audio segment in a preset audio file. Determine the reference audio segment corresponding to the target audio segment, and determine the second audio attribute information of the reference audio segment. The second audio attribute information is used to indicate the audio start point and audio duration of the reference audio segment. The audio start point indicated by the second audio attribute information is based on the audio progress time of the reference audio segment in a preset audio file. Based on the first audio attribute information and the second audio attribute information, the target audio attribute information of the target audio segment is determined. The target audio attribute information is used to indicate whether there is a usable audio segment and the audio start point and audio duration of at least one usable audio segment. The usable audio segment is an audio segment selected from the target audio segments that can be used for audio merging.

6. The method according to claim 5, characterized in that, Determining the reference audio segment corresponding to the target audio segment includes: The target audio segments are sorted according to the order of their acquisition times. For each target audio segment, the target audio segment that is sorted after the target audio segment is determined as the reference audio segment corresponding to the target audio segment.

7. The method according to claim 1, characterized in that, Based on the target audio attribute information of each target audio segment, the plurality of target audio segments are processed, including: For each target audio segment, based on the target audio attribute information of the target audio segment, usable audio segments are cut out from the target audio segment and the audio starting point corresponding to the usable audio segment is determined. The usable audio segment is an audio segment selected from the target audio segment that can be used for audio merging. The audio starting point indicated by the target audio attribute information is based on the audio progress time of the usable audio segment in the preset audio file. According to the order of the audio starting points corresponding to the available audio segments, the available audio segments cut from each target audio segment are spliced ​​and merged.

8. An audio processing apparatus, characterized in that, The device includes: The first determining module is used to determine multiple target audio segments; The second determining module is used to determine the target audio attribute information of each target audio segment. The target audio attribute information is used to mark the available audio segments in the target audio segments. The available audio segments meet the following conditions: there is no reference audio segment that overlaps with the available audio segments, the reference audio segment is collected after the target audio segment, and all of the target audio segments are newly formed audio segments using the same preset audio file as the singing material. An audio processing module is used to process the plurality of target audio segments based on the target audio attribute information of each target audio segment; The audio processing module is specifically used to: determine the available audio segments corresponding to each target audio segment based on the target audio attribute information of each target audio segment, and perform audio merging processing on each of the available audio segments to eliminate overlapping audio segments in the merged audio.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the audio processing method as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the audio processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Audio replacement method, device and system and storage medium

    CN111370011A