Audio processing method and apparatus, electronic device, and storage medium

By determining and processing the attribute information of the target audio clip in the singing application, the problem of unsmooth audio recording and overlap in the application is solved, and high-quality overlapping audio merging is achieved.

WO2025092825A1PCT designated stage expired Publication Date: 2025-05-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/128533
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the singing application, audio recording cannot be smoothly performed, and the recorded audio has overlapping problems.

Method used

By determining multiple target audio clips and their target audio attribute information, non-overlapping available audio clips are marked and multiple target audio clips are processed based on these attribute information to achieve non-overlapping audio merging.

Benefits of technology

It realizes smooth audio recording in singing applications, reduces lag during recording, and eliminates duplicate parts after audio merge, improving audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128533_08052025_PF_FP_ABST
    Figure CN2024128533_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing method and apparatus, an electronic device, and a storage medium. The method comprises: determining multiple target audio clips (S110); determining target audio attribute information of each target audio clip, wherein the target audio attribute information is used for tagging an available audio clip in the target audio clip, and the available audio clip satisfies the following conditions: no overlap is present between a reference audio clip and the available audio clip, the reference audio clip is collected after the target audio clip, and the reference audio clip and the target audio clip are both audio clips formed by using the same preset audio file as singing material (S120); and on the basis of the target audio attribute information of each target audio clip, processing the multiple target audio clips (S130).
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “Audio processing method, device, electronic device and storage medium” and application number 202311436315.6 filed on October 31, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to data processing technology, and more particularly to an audio processing method, apparatus, electronic device, and storage medium. Background Art

[0003] With the continuous development of the Internet, various singing applications have gradually come into people's view and are loved by people. At the same time, the requirements for the singing experience of singing applications are getting higher and higher.

[0004] In related solutions, such as singing applications, it is necessary to collect audio data generated by the singing, and then determine whether the collected audio data needs to be adjusted. If adjustment is required, it is re-recorded. This may result in the audio data being recorded repeatedly, and each time it needs to be re-recorded, making the singing process time-consuming, labor-intensive, and costly, and causing great inconvenience during the singing recording process.

[0005] Summary of the Invention

[0006] The present disclosure provides an audio processing method, device, electronic device and storage medium to solve the problem that audio cannot be recorded smoothly in a singing application and the recorded audio overlaps.

[0007] In a first aspect, an embodiment of the present disclosure provides an audio processing method, the method comprising: determining multiple target audio segments; determining target audio attribute information of each target audio segment, the target audio attribute information being used to mark available audio segments in the target audio segments, the available audio segments satisfying the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material; processing the multiple target audio segments based on the target audio attribute information of each target audio segment.

[0008] In the second aspect, the embodiment of the present disclosure also provides an audio processing device, which includes: a first determination module for determining multiple target audio segments; a second determination module for determining target audio attribute information of each of the target audio segments, wherein the target audio attribute information is used to mark the available audio segments in the target audio segments, and the available audio segments meet the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material; an audio processing module for processing the multiple target audio segments based on the target audio attribute information of each of the target audio segments.

[0009] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the audio processing method described in any one of the above embodiments.

[0010] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable medium, wherein the computer-readable medium stores computer instructions, and the computer instructions are used to enable a processor to implement the audio processing method described in any one of the above embodiments when executed.

[0011] In a fifth aspect, a computer program product is also provided in an embodiment of the present disclosure. The computer program product is tangibly stored in a computer storage medium and includes computer executable instructions. When the computer executable instructions are executed by a device, the device executes the audio processing method described in any one of the above embodiments.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0014] FIG1 is a flow chart of an audio processing method provided by an embodiment of the present disclosure;

[0015] FIG2 is a schematic diagram of an interface of a singing application applicable to an embodiment of the present disclosure;

[0016] FIG3 a is a schematic diagram showing the mapping of audio clips with different acquisition times onto a timeline applicable to an embodiment of the present disclosure;

[0017] FIG3 b is a schematic diagram of merging audio clips collected at different times according to an embodiment of the present disclosure;

[0018] FIG4 is a flow chart of another audio processing method provided by an embodiment of the present disclosure;

[0019] FIG5 is a flow chart of another audio processing method provided by an embodiment of the present disclosure;

[0020] FIG6 is a detailed decomposition diagram of merging audio clips with different acquisition times applicable to an embodiment of the present disclosure;

[0021] FIG7 is a schematic structural diagram of an audio processing device provided by an embodiment of the present disclosure;

[0022] FIG8 is a schematic structural diagram of an electronic device for implementing an audio processing method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0024] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0025] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0027] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0028] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0029] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0030] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0031] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0032] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0033] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0034] Figure 1 is a flow chart of an audio processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to situations where continuous singing is performed in a singing application and the audio generated by the singing is merged. The method can be executed by an audio processing device, which can be implemented in the form of software and / or hardware and is generally integrated on any electronic device with network communication function, which can be a mobile terminal, PC or server, etc.

[0035] As shown in FIG1 , the audio processing method according to an embodiment of the present disclosure may include the following steps:

[0036] S110: Determine multiple target audio segments.

[0037] S120. Determine target audio attribute information of each target audio segment. The target audio attribute information is used to mark available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no overlap between the reference audio segment and the available audio segment. The reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material.

[0038] Referring to FIG2 , schematic interface diagrams of the “song middle page”, “singing recording page” and “preview page” involved in the singing application are provided in sequence. In the “singing recording page” shown in FIG2 , the singing subject can use the audio clips selected from the same preset audio file provided in the singing application at different acquisition times as singing materials to repeatedly practice singing. The audio clips selected from the preset audio file at each acquisition time may be the same, partially the same, or different. A plurality of target audio clips are obtained by singing with the selected audio clips as singing materials at different acquisition times and collecting the audio clips formed by the singing. Moreover, the audio clips selected from the preset audio file at each acquisition time may overlap, so there will be overlapping parts between the newly formed audio clips when singing according to the selected audio clips as singing materials, that is, there may be many audio clips with overlapping intervals between the target audio clips, resulting in unclear audio after the audio is merged.

[0039] When using a singing application to sing, for the audio clips formed when singing the same audio clip selected from a preset audio file at different collection times, only if the user is not satisfied with the newly formed audio clip of the previous singing, will the user choose to sing again at the next collection time to form a new audio clip until the audio clip formed by repeated singing is satisfactory. It can be seen that the availability of audio clips with a later collection time is stronger than that of audio clips with an earlier collection time. Then the audio clips formed when singing the same audio clip will usually be dominated by audio clips with a later collection time. At the same time, there are some audio clips in the audio clip with an earlier collection time that are not reflected in the audio clips after the collection time. Therefore, the audio clips with an earlier collection time cannot be directly eliminated, but they must be selectively eliminated and retained.

[0040] To this end, for each target audio segment, it can be determined whether there is a series of audio segments whose collection time is later than that of the target audio segment among the multiple target audio segments. If so, the series of audio segments whose collection time is later than that of the target segment among the multiple target audio segments is selected as the reference audio segment corresponding to the target audio segment. If not, it is determined that there is no corresponding reference audio segment for the target audio segment. In some embodiments, the reference audio segment corresponding to the target audio segment is an audio segment formed by singing an audio segment selected from the same preset audio file, whose collection time is later than that of the target audio segment, so that the reference audio segment whose collection time is later than that of the target audio segment can be used to correct and eliminate overlapping audio segments in the target audio segment that do not meet the requirements.

[0041] Referring to Figure 3a, the newly generated audio clips, which are selected from the same preset audio file at different capture times and used as singing materials, are sequentially referred to as the audio clip of capture 1, the audio clip of capture 2, and the audio clip of capture 3. These are denoted as the multiple target audio clips to be merged. Furthermore, for each of the multiple target audio clips, a reference audio clip can be selected from audio clips that were captured later than the target audio clip.

[0042] The audio segments of Collection 1, Collection 2, and Collection 3 shown in Figure 3a can all be used as target audio segments in sequence. When the target audio segment is the audio segment of Collection 1, the reference audio segments corresponding to the audio segment of Collection 1 are, in sequence, the audio segments of Collection 2 and Collection 3. As can be seen from Figure 3a, there are some overlapping audio segments between the audio segments of Collection 2 and Collection 1. Since the audio segment of Collection 2 was collected later than the audio segment of Collection 1, the audio segment of Collection 1, which is used as the target audio segment, needs to be discarded when the audio segment is subsequently merged.

[0043] Referring to Figure 3a, similarly, the audio segments of Collection 3 partially overlap with those of Collection 1. Since the audio segments of Collection 3 were collected later than those of Collection 1, the overlapping segments of Collection 1, as the target audio segments, also need to be discarded during the subsequent merging process. Referring to Figure 3b, when the audio segments of Collection 1 are used as the target audio segments, the overlapping segments of Collection 1, Collection 2, and Collection 3 are discarded, and the audio segments of Collection 1 are split into two usable audio segments corresponding to Collection 1.

[0044] To this end, when a reference audio segment corresponding to a target audio segment is obtained, based on the overlap between the target audio segment and the corresponding reference audio segment, the portion of the target audio segment that does not overlap with the reference audio segment can be identified as an available audio segment, and the positions of the available audio segments in the target audio segment can be marked. Furthermore, when a reference audio segment corresponding to the target audio segment is not obtained, the entire target audio segment can be directly identified as an available audio segment, and the positions of the available audio segments in the target audio segment can be marked. By marking the positions of the available audio segments in the target audio segment, it is possible to retain the available audio segments when merging multiple target audio segments.

[0045] S130: Process multiple target audio segments based on the target audio attribute information of each target audio segment.

[0046] For each of the multiple target audio segments, the target audio attribute information can be used to mark the available audio segments in each target audio segment that do not overlap with the reference audio segment corresponding to the target audio segment, and the marked target audio segments can be saved independently. This allows the user to directly use the available audio segments in the target audio segments for audio merging when subsequently selecting to merge the target audio segments, eliminating duplicate audio segments in the merged audio and resolving the issue of audio recording not being able to be performed smoothly and with overlapping recorded audio in singing applications.

[0047] Optionally, available audio segments that do not overlap with reference audio segments corresponding to the target audio segment are selected from each target audio segment, and then the selected available audio segments are decoded and then merged to obtain a merged audio file.

[0048] As an optional but non-limiting implementation, processing multiple target audio segments based on target audio attribute information of each target audio segment may include the following steps A1-A2:

[0049] Step A1: For each target audio segment, based on the target audio attribute information of the target audio segment, crop an available audio segment from the target audio segment and determine the audio starting point corresponding to the available audio segment. The available audio segment is an audio segment selected from the target audio segment that can be used for audio merging. The audio starting point indicated by the target audio attribute information is based on the audio progress time of the available audio segment in the preset audio file.

[0050] Step A2: merging the available audio segments cut from the target audio segments in accordance with the order of the audio starting points corresponding to the available audio segments.

[0051] The target audio attribute information indicates the audio starting point of the available audio segment in the target audio segment that can be used for audio merging. The audio starting point indicated by the target audio attribute information can refer to the starting point of the available audio segment mapped to the audio progress time corresponding to the preset audio file. After the available audio segments that need to be merged are cut out from each target audio segment, the selected available audio segments can be decoded, and the corresponding decoded available audio segments can be spliced ​​together in the time sequence of the audio starting points corresponding to each available audio segment to obtain a merged audio file. Among them, the AVComposition library or a third-party audio cutting library can be used to implement the cutting of available audio segments from the target audio segment.

[0052] The solution of the embodiment of the present disclosure is to obtain multiple target audio segments by selecting audio segments from a preset audio file in a singing application as singing materials and then continuing the next singing without waiting for audio processing after one singing is completed. In this way, it is possible to perform smooth singing in the singing application without waiting for the audio processing to be completed before continuing the singing, which can reduce the jamming in the singing application to a certain extent. At the same time, after obtaining multiple target audio segments, the reference audio segment corresponding to each target audio segment can be determined. By comparing the target audio segment with the reference audio segment, the target audio attribute information for marking the available audio segments in the target audio segment that do not overlap with the reference audio segment can be determined, and then the target audio segments that are uniformly marked can be merged without overlapping audio according to the target audio attribute information of each target audio segment. Since no audio merging operation is performed when the target audio segments are uniformly obtained, the overlapping parts in the audio segments do not need to be processed, which can save some computing performance consumption.

[0053] Figure 4 is a flow chart of another audio processing method provided in an embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining multiple target audio segments in the aforementioned embodiment on the basis of the aforementioned embodiment. This embodiment can be combined with various optional solutions in one or more of the aforementioned embodiments.

[0054] As shown in FIG4 , the audio processing method according to an embodiment of the present disclosure may include the following steps:

[0055] S410: Determine a plurality of candidate audio segments, where the plurality of candidate audio segments are newly formed audio segments based on audio segments selected from the same preset audio file at different acquisition times as singing materials.

[0056] Optionally, when the audio segments selected from the same preset audio file at different acquisition times are identical or partially identical, and the audio segments selected from the same preset audio file at different acquisition times are identical or partially identical, the multiple candidate audio segments are at least partially identical. The multiple candidate audio segments are audio segments formed by pronouncing the audio segments selected from the same preset audio file at different acquisition times as singing materials.

[0057] Referring to Figures 3a and 3b, when practicing singing in a singing application, the reason why multiple target audio clips are collected and overlapped is that, on the one hand, the selected audio clips will be practiced repeatedly and audio will be collected as singing materials. For example, if the audio clip does not meet the singing requirements and needs to be practiced repeatedly and recorded, the audio clips selected from the same preset audio file at different collection times are the same. On the other hand, it is because the audio formed by the singing needs to be shortened and spliced ​​together. If the front and back audio clips of the singing cannot be connected, then even if they are spliced ​​into one audio file, the audio will be incomplete in the middle. In order to continue the audio clip obtained from the previous singing, the singing practice will be repeated on the end of the audio clip of the previous singing. The audio clips corresponding to the two collection moments are partially the same. At this time, the audio clips selected from the same preset audio file at different collection times are partially the same.

[0058] As an optional but non-limiting implementation, determining multiple candidate audio segments may include the following steps B1-B2:

[0059] Step B1, in response to the audio collection operation, collect the audio segment formed based on the preset singing material to obtain a candidate audio segment and mark the corresponding collection time, wherein the preset singing material is the singing material corresponding to the audio segment selected from the preset audio file.

[0060] Referring to Figure 2, in the singing application, the singing recording page starts the audio collection operation when the singing recording control in the singing recording page is triggered and clicked. At this time, the singing can be performed according to the audio clip provided in the preset audio file as the singing material, and the audio clip formed by the singing can be collected in real time through the audio collector. The audio collection operation ends when the singing recording control is triggered and clicked again. The audio clip collected between the start of the audio collection operation and the pause of the audio collection operation is the audio clip selected from the preset audio file as the candidate audio clip newly formed as the singing material. At the same time, the collection time of the candidate audio clip is marked, which can also be understood as the singing recording time of the candidate audio clip. In some embodiments, the audio collector can be implemented using the AudioUnit library or a third-party audio collection library.

[0061] When practicing singing in a singing application, there may be situations where the same selected audio clip does not meet the singing requirements and needs to be repeatedly practiced and recorded, or the end of the audio clip of the previous singing needs to be repeated for singing practice in order to continue the previous singing. For this reason, in response to another audio collection operation, the audio clip selected from the preset audio file is collected as the audio clip formed by the singing material to obtain the next candidate audio clip and mark the corresponding collection time. That is to say, after completing the collection of a candidate audio clip for singing, the next singing can be continued. It is also necessary to collect the audio clip formed by singing from the start of the audio collection operation to the end of the audio collection operation, obtain another candidate audio clip and mark the collection time corresponding to the candidate audio clip. When the audio clips selected from the same preset audio file at different collection times are the same or partially the same, then the obtained candidate audio clips are at least partially overlapped.

[0062] Step B2: Determine the candidate audio segments corresponding to the multiple audio collection operations as multiple candidate audio segments, and the preset audio file used in the multiple audio collection operations is the same audio file.

[0063] Merging and removing overlapping audio segments is only necessary when the audio segments selected as singing materials at different acquisition times originate from the same preset audio file. If they do not originate from the same audio file, merging and removing overlapping audio segments is not necessary.

[0064] S420: Determine multiple target audio segments from multiple candidate audio segments.

[0065] The reference audio segment is collected later than the target audio segment, and the target audio segment and the reference audio segment are newly generated audio segments based on audio segments selected from the same preset audio file as the performance material. The reference audio segment is collected after the target audio segment and is generated from the same preset audio file as the target audio segment.

[0066] Optionally, all of the candidate audio segments may be used as the target audio segments, with each candidate audio segment serving as a target audio segment. Alternatively, some of the candidate audio segments may be selected from the candidate audio segments as the target audio segments, with each selected candidate audio segment serving as a target audio segment.

[0067] As an optional but non-limiting implementation, determining multiple target audio segments from multiple candidate audio segments may include the following steps:

[0068] Step C1: Sort the candidate audio segments included in the multiple candidate audio segments according to the chronological order of their collection time, and determine the candidate audio segments selected in sequence as the multiple target audio segments according to the sorting results of the candidate audio segments, until all the candidate audio segments in the multiple candidate audio segments are traversed or a candidate audio segment is selected.

[0069] S430. Determine target audio attribute information of each target audio segment. The target audio attribute information is used to mark available audio segments in the target audio segment. The available audio segments meet the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material.

[0070] S440: Process the multiple target audio segments based on the target audio attribute information of each target audio segment.

[0071] The solution of the disclosed embodiment selects an audio clip from a preset audio file in a singing application as singing material. After a performance is completed, the selected audio clip is directly used as singing material without waiting for audio processing. The singing material is then used to continue the next performance to form an audio clip. After multiple performances are performed by selecting singing material, multiple target audio clips are sequentially obtained. This allows for smooth singing in the singing application without waiting for audio processing to complete before continuing. This can reduce lag in the singing application to a certain extent. At the same time, after obtaining multiple target audio clips, it can be determined whether each target audio clip has a corresponding reference audio clip. If a reference audio clip exists, the target audio clip can be compared with the reference audio clip to determine the target audio attribute information used to mark the available audio clips in the target audio clip that do not overlap with the reference audio clip. Then, according to the target audio attribute information of each target audio clip, the marked target audio clips can be merged to achieve non-overlapping audio. Since no audio merging operation is performed when the target audio clips are uniformly obtained, the overlapping portions of the audio clips do not need to be processed, which saves some computing power. Figure 5 is a flow chart of another audio processing method provided by an embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining the target audio attribute information of each target audio segment in the aforementioned embodiment on the basis of the aforementioned embodiment. This embodiment can be combined with various optional solutions in one or more of the aforementioned embodiments.

[0072] As shown in FIG5 , the audio processing method according to the embodiment of the present disclosure may include the following steps:

[0073] S510: Determine multiple target audio segments.

[0074] S520. Determine first audio attribute information of the target audio segment, where the first audio attribute information is used to indicate the audio starting point and audio duration of the target audio segment. The audio starting point indicated by the first audio attribute information is based on the audio progress time representation of the target audio segment in a preset audio file.

[0075] The first audio attribute information indicates the audio starting point and audio duration of the target audio segment. The audio starting point indicated by the first audio attribute information may refer to the starting point of the target audio segment mapped to the audio progress time corresponding to the preset audio file.

[0076] S530: Determine a reference audio segment corresponding to the target audio segment, and determine second audio attribute information of the reference audio segment. The second audio attribute information is used to indicate an audio start point and audio duration of the reference audio segment. The audio start point indicated by the second audio attribute information is based on an audio progress time representation of the reference audio segment in a preset audio file. The second audio attribute information indicates the audio start point and audio duration of the reference audio segment. The audio start point indicated by the second audio attribute information may refer to the audio progress time corresponding to the mapping of the starting point of the reference audio segment to the preset audio file.

[0077] S540. Determine the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information. The target audio attribute information is used to indicate whether there is an available audio segment and the audio starting point and audio duration of at least one available audio segment. The available audio segment is an audio segment selected from the target audio segment and can be used for audio merging.

[0078] In some embodiments, the target audio attribute information is used to mark the available audio segments in the target audio segment, and the available audio segments meet the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material.

[0079] Optionally, the target audio attribute information is used to mark available audio segments in the target audio segment that do not overlap with the reference audio segment when the reference audio segment exists, and all target audio segments when the reference audio segment does not exist, as available audio segments.

[0080] The target audio attribute information may indicate whether an available audio segment exists in the target audio segment and, if an available audio segment exists in the target hidden segment, the audio start point of the available audio segment that can be used for audio merging. The audio start point indicated by the target audio attribute information may be the start point of the available audio segment mapped to the audio progress time corresponding to the preset audio file.

[0081] As an optional but non-limiting implementation, determining the reference audio segment corresponding to the target audio segment includes the following steps C1-C2:

[0082] Step C1: sort the target audio segments in the plurality of target audio segments according to the chronological order of their corresponding acquisition times.

[0083] Step C2: For each target audio segment, determine the target audio segment that is sorted after the target audio segment as the reference audio segment corresponding to the target audio segment.

[0084] As an optional but non-limiting implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps D1-D2:

[0085] Step D1: If it is determined based on the first audio attribute information and the second audio attribute information that the audio start point of the target audio segment is smaller than the audio start point of the reference audio segment and there are overlapping audio segments, then the audio start point indicated by the first audio attribute information is determined as the audio start point of the available audio segment indicated by the target audio attribute information.

[0086] Step D2: Determine the subtraction value between the audio starting point indicated by the second audio attribute information and the audio starting point indicated by the first audio attribute information as the audio duration of the available audio segment indicated by the target audio attribute information.

[0087] In a singing application, an embodiment of the present disclosure selects an audio clip from a preset audio file as singing material for a performance. After a performance is completed, the selected audio clip is directly used as singing material, without waiting for audio processing. The selected audio clip is then used to continue the next performance, forming an audio clip. By selecting singing material multiple times and performing the performance, multiple target audio clips are sequentially obtained. This allows for smooth singing in the singing application without having to wait for audio processing to complete before continuing. This can reduce lag in the singing application to a certain extent. Furthermore, after obtaining multiple target audio clips, it is possible to determine whether each target audio clip has a corresponding reference audio clip. If a reference audio clip exists, the target audio clip can be compared with the reference audio clip to determine target audio attribute information for marking available audio clips within the target audio clip that do not overlap with the reference audio clip. The marked target audio clips can then be merged according to their target audio attribute information to achieve non-overlapping audio. Since no audio merging operation is performed when the target audio clips are uniformly obtained, the overlapping portions of the audio clips do not need to be processed, saving some computing power.

[0088] Referring to the first one from the left in Figure 6, taking the audio segment of Collection 1 as the target audio segment and the audio segment of Collection 2 as the reference audio segment as an example, the audio starting point indicated by the first audio attribute information of the target audio segment is smaller than the audio starting point indicated by the second audio attribute information of the reference audio segment, and the sum of the audio starting point indicated by the first audio attribute information and the audio duration is greater than the audio starting point indicated by the second audio attribute information of the reference audio segment. From this, it can be seen that there is an overlapping audio segment between the right part of the audio segment of Collection 1 and the left part of the audio segment of Collection 2, and the right part of the audio segment of Collection 1 can be right-segmented. At this time, the audio starting point indicated by the first audio attribute information can be directly determined as the audio starting point of the available audio segment, and the subtraction value between the audio starting point indicated by the second audio attribute information and the audio starting point indicated by the first audio attribute information can be determined as the audio duration of the available audio segment.

[0089] As another optional but non-limiting implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps E1-E3:

[0090] Step E1: If it is determined based on the first audio attribute information and the second audio attribute information that the audio starting point of the target audio segment is greater than the audio starting point of the reference audio segment and there are overlapping audio segments, then the sum of the audio starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio starting point of the available audio segment indicated by the target audio attribute information.

[0091] Step E2: Determine the duration of the overlapping audio segments by subtracting the audio starting point indicated by the target audio attribute information from the audio starting point indicated by the first audio attribute information.

[0092] Step E3: Determine the audio duration of the available audio segment indicated by the target audio attribute information as the subtraction value between the audio duration indicated by the first audio attribute information and the duration of the overlapping audio segment.

[0093] Referring to the second one from the left in Figure 6, taking the audio segment of Collection 1 as the target audio segment and the audio segment of Collection 2 as the reference audio segment as an example, the audio starting point indicated by the first audio attribute information of the target audio segment is greater than the audio starting point indicated by the second audio attribute information of the reference audio segment, and the sum of the audio starting point indicated by the second audio attribute information and the audio duration is greater than the audio starting point indicated by the first audio attribute information of the target audio segment. From this, it can be seen that there are overlapping audio segments between the audio segment of Collection 1 and the audio segment of Collection 2, specifically, the left part of the audio segment of Collection 1 overlaps with the right part of the audio segment of Collection 2. The left part of the audio segment of Collection 1 can be split on the left. At this time, the subtraction value between the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information can be determined as the overlapping audio segment duration of the right part of the audio segment of Collection 1 and the left part of the audio segment of Collection 2. Then, the subtraction value of the audio duration indicated by the first audio attribute information and the duration of the overlapping audio segment is determined as the audio duration of the available audio segment, and the addition value of the audio starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio starting point of the available audio segment.

[0094] As another optional but non-limiting implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps F1-F5:

[0095] Step F1. If the audio starting point of the target audio segment is determined to be smaller than the audio starting point of the reference audio segment based on the first audio attribute information and the second audio attribute information, and the sum of the audio starting point and the audio duration of the target audio segment is greater than the sum of the audio starting point and the audio duration of the reference audio segment, then the audio starting point indicated by the first audio attribute information is determined to be the audio starting point of the first available audio segment indicated by the target audio attribute information.

[0096] Step F2: Determine the subtraction value between the audio starting point indicated by the second audio attribute information and the audio starting point indicated by the first audio attribute information as the audio duration of the first available audio segment indicated by the target audio attribute information.

[0097] Step F3: Determine the sum of the audio start point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information as the audio start point of the second available audio segment indicated by the target audio attribute information.

[0098] Step F4: Determine the duration of the overlapping audio segments by subtracting the audio starting point indicated by the target audio attribute information from the audio starting point indicated by the first audio attribute information.

[0099] Step F5: Determine the subtraction value between the audio duration indicated by the first audio attribute information and the duration of the overlapping audio segment as the audio duration of the second available audio segment indicated by the target audio attribute information.

[0100] In some embodiments, the first available audio segment and the second available audio segment are two audio segments selected from the target audio segment and capable of being used for audio merging.

[0101] Referring to the third audio segment from the left in Figure 6 , taking the audio segment from Collection 1 as the target audio segment and the audio segment from Collection 2 as the reference audio segment, the overlapping audio segment of the audio segments from Collection 1 and Collection 2 is located in the middle of the audio segment from Collection 1. In this case, the right-side segmentation of steps D1-D2 and the left-side segmentation of steps E1-E3 can be combined to implement the process from steps F1-F5. This process sequentially segments the left and right sides of the audio segment from Collection 1 into a first usable audio segment and a second usable audio segment. The first and second usable audio segments are then selected from the target audio segment and used for audio merging.

[0102] As an optional but non-limiting implementation, determining the target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information may include the following steps:

[0103] If it is determined based on the first audio attribute information and the second audio attribute information that the audio starting point of the target audio segment is greater than the audio starting point of the reference audio segment, and the sum of the audio starting point and the audio duration of the target audio segment is less than the sum of the audio starting point and the audio duration of the reference audio segment, then the target audio attribute information of the target audio segment is determined to be that there is no available audio segment in the target audio segment.

[0104] As shown in Figure 6, the fourth audio clip from the left is the target audio clip, and the audio clip from Collection 2 is the reference audio clip. If the audio clip from Collection 1 is completely overwritten by the audio clip from Collection 2, then all the audio clips from Collection 1 are overlapping audio clips. In this case, the audio clip from Collection 1 can be deleted and overwritten, indicating that there are no usable audio clips in the target audio clip.

[0105] S550: Process multiple target audio segments based on the target audio attribute information of each target audio segment.

[0106] The solution of the embodiment of the present disclosure is to obtain multiple target audio segments by selecting audio segments from a preset audio file in a singing application as singing materials and then continuing the next singing without waiting for audio processing after one singing is completed. In this way, it is possible to perform smooth singing in the singing application without waiting for the audio processing to be completed before continuing the singing, which can reduce the jamming in the singing application to a certain extent. At the same time, after obtaining multiple target audio segments, the reference audio segment corresponding to each target audio segment can be determined. By comparing the target audio segment with the reference audio segment, the target audio attribute information for marking the available audio segments in the target audio segment that do not overlap with the reference audio segment can be determined, and then the target audio segments that are uniformly marked can be merged without overlapping audio according to the target audio attribute information of each target audio segment. Since no audio merging operation is performed when the target audio segments are uniformly obtained, the overlapping parts in the audio segments do not need to be processed, which can save some computing performance consumption.

[0107] Figure 7 is a structural diagram of an audio processing device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to situations where continuous singing is performed in a singing application and the audio generated by the singing is merged. The audio processing device can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication function, which can be a mobile terminal, PC or server, etc.

[0108] As shown in FIG7 , the audio processing device of the embodiment of the present disclosure may include: a first determination module 710, a second determination module 720 and an audio processing module 730. In particular:

[0109] The first determining module 710 is configured to determine a plurality of target audio segments.

[0110] The second determination module 720 is used to determine the target audio attribute information of each target audio segment, and the target audio attribute information is used to mark the available audio segments in the target audio segment, and the available audio segments meet the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material.

[0111] The audio processing module 730 is configured to process the multiple target audio segments based on the target audio attribute information of each target audio segment.

[0112] Based on the above embodiment, optionally, determining multiple target audio segments includes:

[0113] Determining a plurality of candidate audio segments, wherein the plurality of candidate audio segments are newly formed audio segments based on audio segments selected from the same preset audio file at different acquisition times as singing materials;

[0114] A plurality of target audio segments are determined from the plurality of candidate audio segments.

[0115] Based on the above embodiment, optionally, determining a plurality of candidate audio segments includes:

[0116] In response to an audio collection operation, an audio segment formed based on a preset singing material is collected to obtain a candidate audio segment and the corresponding collection time is marked, wherein the preset singing material is the singing material corresponding to the audio segment selected from the preset audio file; the candidate audio segments corresponding to multiple audio collection operations are determined as the multiple candidate audio segments, and the preset audio file used in the multiple audio collection operations is the same audio file.

[0117] Based on the above embodiment, optionally, the audio segments selected from the same preset audio file at different acquisition times are identical or partially identical.

[0118] Based on the above embodiment, optionally, determining target audio attribute information of each target audio segment includes: determining first audio attribute information of the target audio segment, the first audio attribute information being used to indicate an audio start point and audio duration of the target audio segment, the audio start point indicated by the first audio attribute information being based on an audio progress time representation of the target audio segment in a preset audio file; determining a reference audio segment corresponding to the target audio segment, and determining second audio attribute information of the reference audio segment, the second audio attribute information being used to indicate an audio start point and audio duration of the reference audio segment, the audio start point indicated by the second audio attribute information being based on an audio progress time representation of the reference audio segment in the preset audio file; and determining target audio attribute information of the target audio segment based on the first audio attribute information and the second audio attribute information, the target audio attribute information being used to indicate whether there is an available audio segment and the audio start point and audio duration of at least one available audio segment, the available audio segment being an audio segment selected from the target audio segment and capable of being used for audio merging.

[0119] Based on the above embodiment, optionally, determining a reference audio segment corresponding to the target audio segment includes: sorting the target audio segments included in the multiple target audio segments according to the chronological order of their corresponding acquisition times; and for each target audio segment, determining a target audio segment sorted after the target audio segment as the reference audio segment corresponding to the target audio segment.

[0120] On the basis of the above embodiment, optionally, the target audio attribute information of the target audio segment is determined based on the first audio attribute information and the second audio attribute information, including: if the audio starting point of the target audio segment determined based on the first audio attribute information and the second audio attribute information is smaller than the audio starting point of the reference audio segment and there are overlapping audio segments, then the audio starting point indicated by the first audio attribute information is determined as the audio starting point of the available audio segment indicated by the target audio attribute information; and the subtraction value between the audio starting point indicated by the second audio attribute information and the audio starting point indicated by the first audio attribute information is determined as the audio duration of the available audio segment indicated by the target audio attribute information.

[0121] On the basis of the above embodiment, optionally, the target audio attribute information of the target audio segment is determined based on the first audio attribute information and the second audio attribute information, including: if the audio starting point of the target audio segment determined based on the first audio attribute information and the second audio attribute information is greater than the audio starting point of the reference audio segment and there is an overlapping audio segment, then the sum of the audio starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio starting point of the available audio segment indicated by the target audio attribute information; the subtraction value of the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information is determined as the overlapping audio segment duration; and the subtraction value of the audio duration indicated by the first audio attribute information and the overlapping audio segment duration is determined as the audio duration of the available audio segment indicated by the target audio attribute information.

[0122] On the basis of the above embodiment, optionally, the target audio attribute information of the target audio segment is determined according to the first audio attribute information and the second audio attribute information, including: if the audio starting point of the target audio segment determined according to the first audio attribute information and the second audio attribute information is smaller than the audio starting point of the reference audio segment, and the sum of the audio starting point and the audio duration of the target audio segment is greater than the sum of the audio starting point and the audio duration of the reference audio segment, then the audio starting point indicated by the first audio attribute information is determined to be the audio starting point of the first available audio segment indicated by the target audio attribute information; and the audio starting point indicated by the second audio attribute information is determined to be the audio starting point of the first available audio segment indicated by the target audio attribute information. The subtraction value of the audio starting point indicated by the target audio attribute information is determined as the audio duration of the first available audio segment indicated by the target audio attribute information; the addition value of the audio starting point indicated by the second audio attribute information and the audio duration indicated by the second audio attribute information is determined as the audio starting point of the second available audio segment indicated by the target audio attribute information; the subtraction value of the audio starting point indicated by the target audio attribute information and the audio starting point indicated by the first audio attribute information is determined as the overlapping audio segment duration; the subtraction value of the audio duration indicated by the first audio attribute information and the overlapping audio segment duration is determined as the audio duration of the second available audio segment indicated by the target audio attribute information.

[0123] In some embodiments, the first available audio segment and the second available audio segment are two audio segments selected from the target audio segment and can be used for audio merging.

[0124] On the basis of the above embodiment, optionally, the target audio attribute information of the target audio segment is determined based on the first audio attribute information and the second audio attribute information, including: if the audio starting point of the target audio segment determined based on the first audio attribute information and the second audio attribute information is greater than the audio starting point of the reference audio segment, and the sum of the audio starting point and the audio duration of the target audio segment is less than the sum of the audio starting point and the audio duration of the reference audio segment, then the target audio attribute information of the target audio segment is determined to be that there is no available audio segment in the target audio segment.

[0125] Based on the above embodiment, optionally, the multiple target audio segments are processed based on the target audio attribute information of each target audio segment, including: for each target audio segment, based on the target audio attribute information of the target audio segment, cutting out a usable audio segment from the target audio segment and determining an audio starting point corresponding to the usable audio segment, the usable audio segment being an audio segment selected from the target audio segment that can be used for audio merging, the audio starting point indicated by the target audio attribute information being represented based on an audio progress time of the usable audio segment in a preset audio file; and splicing and merging the usable audio segments cut out from the target audio segments in a chronological order of the audio starting points corresponding to the usable audio segments.

[0126] In the disclosed embodiment, multiple target audio segments are obtained in sequence by singing an audio segment selected from a preset audio file in a singing application and continuing the next singing with the audio segment selected directly from the preset audio file without waiting for audio processing after one singing is completed. In this way, smooth singing can be achieved in the singing application without waiting for audio processing, which can reduce the jamming in the singing application to a certain extent. At the same time, after obtaining multiple target audio segments, the reference audio segment corresponding to each target audio segment can be determined. By comparing the target audio segment with the reference audio segment, the target audio attribute information for marking the available audio segments in the target audio segment that do not overlap with the reference audio segment can be determined, and then the target audio segments that are uniformly marked can be merged without overlapping audio according to the target audio attribute information of each target audio segment. Since no audio merging operation is performed when the target audio segments are uniformly obtained, some computing performance consumption can be saved.

[0127] The audio processing device provided by the embodiments of the present disclosure can execute the audio processing method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0128] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0129] FIG8 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG8 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG8 ) 800 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG8 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0130] As shown in FIG8 , the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An edit / output (I / O) interface 805 is also connected to the bus 804.

[0131] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0132] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the audio processing method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the audio processing method of the embodiment of the present disclosure are performed.

[0133] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0134] The electronic device provided by the embodiment of the present disclosure and the audio processing method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0135] An embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the audio processing method provided by the above embodiment is implemented.

[0136] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0137] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0138] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0139] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: determines multiple target audio segments; determines the target audio attribute information of each of the target audio segments, and the target audio attribute information is used to mark the available audio segments in the target audio segments, and the available audio segments meet the following conditions: there is no reference audio segment overlapping with the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material; based on the target audio attribute information of each of the target audio segments, the multiple target audio segments are processed.

[0140] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0142] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0143] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0146] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0147] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An audio processing method, the method comprising: determining a plurality of target audio segments; Determine target audio attribute information of each target audio segment, wherein the target audio attribute information is used to mark available audio segments in the target audio segment, and the available audio segments meet the following conditions: there is no overlap between the reference audio segment and the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material; The multiple target audio segments are processed based on the target audio property information of each target audio segment.

2. The method according to claim 1, wherein determining a plurality of target audio segments comprises: Determine a plurality of candidate audio segments, wherein the plurality of candidate audio segments are audio segments newly formed based on audio segments selected from the same preset audio file at different acquisition times as singing materials; A plurality of target audio segments are determined from the plurality of candidate audio segments.

3. The method according to claim 2, wherein determining a plurality of candidate audio segments comprises: In response to the audio collection operation, an audio segment formed based on a preset singing material is collected to obtain a candidate audio segment and a corresponding collection time is marked, wherein the preset singing material is a singing material corresponding to the audio segment selected from a preset audio file; The candidate audio segments corresponding to the multiple audio collection operations are determined as the multiple candidate audio segments, and the preset audio files used by the multiple audio collection operations are the same audio file.

4. The method according to claim 2 or 3, wherein the audio segments selected from the same preset audio file at different acquisition times are identical or partially identical.

5. The method according to claim 1, wherein determining the target audio attribute information of each of the target audio segments comprises: Determine first audio attribute information of the target audio segment, where the first audio attribute information is used to indicate an audio start point and an audio duration of the target audio segment, and the audio start point indicated by the first audio attribute information is represented based on an audio progress time of the target audio segment in a preset audio file; Determine a reference audio segment corresponding to the target audio segment, and determine second audio attribute information of the reference audio segment, where the second audio attribute information is used to indicate an audio start point and an audio duration of the reference audio segment, and the audio start point indicated by the second audio attribute information is represented based on an audio progress time of the reference audio segment in a preset audio file; Based on the first audio attribute information and the second audio attribute information, target audio attribute information of the target audio segment is determined, and the target audio attribute information is used to indicate whether there is an available audio segment and the audio starting point and audio duration of at least one available audio segment. The available audio segment is an audio segment selected from the target audio segment and can be used for audio merging.

6. The method according to claim 5, wherein determining the reference audio segment corresponding to the target audio segment comprises: Sorting the target audio segments according to the order of the acquisition time corresponding to the target audio segments; For each of the target audio segments, a target audio segment that is ranked after the target audio segment is determined as a reference audio segment corresponding to the target audio segment.

7. The method according to claim 1, wherein the processing of the plurality of target audio segments based on the target audio attribute information of each of the target audio segments comprises: For each of the target audio segments, based on the target audio attribute information of the target audio segment, a usable audio segment is cut out from the target audio segment and an audio starting point corresponding to the usable audio segment is determined, wherein the usable audio segment is an audio segment selected from the target audio segment and can be used for audio merging, and the audio starting point indicated by the target audio attribute information is represented based on an audio progress time of the usable audio segment in a preset audio file; The available audio segments cut out from the target audio segments are spliced ​​and merged according to the order of the audio starting points corresponding to the available audio segments.

8. An audio processing device, the device comprising: A first determination module, used to determine a plurality of target audio segments; A second determination module is used to determine target audio attribute information of each target audio segment, wherein the target audio attribute information is used to mark available audio segments in the target audio segment, and the available audio segments meet the following conditions: there is no overlap between the reference audio segment and the available audio segment, the reference audio segment is collected after the target audio segment, and both the reference audio segment and the target audio segment are audio segments formed by using the same preset audio file as the singing material; The audio processing module is used to process the multiple target audio segments based on the target audio attribute information of each target audio segment.

9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the audio processing method as described in any one of claims 1 to 7.

10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the audio processing method according to any one of claims 1 to 7 when executed by a computer processor.

11. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio splicing method and apparatus, and storage medium

    CN108831424A

  • Information processing method and system, first equipment and second equipment

    CN110334240A

  • Audio replacement method, device and system and storage medium

    CN111370011A

  • Video correction method and device, readable medium and electronic equipment

    CN111935541A

  • Audio data acquisition method and device

    CN114387994A