Volume adjustment method, device, storage medium and computing device

By extracting and ratio calculations, the volume of the second recorded audio is adjusted, which solves the problem of inconsistent volume in the karaoke application and improves the overall effect and user experience of recorded audio.

CN114863953BActive Publication Date: 2025-05-27HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210427136.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-05-27
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

In the karaoke application, when users re-record audio, it is difficult for users to control the volume due to lack of pre- and rear audio preparation, resulting in inconsistent volumes of the first recorded audio and the second recorded audio, and the volume is abrupt, which reduces the overall effect and user experience of the recorded audio.

Method used

By obtaining the first recorded audio and the second recorded audio, extracting their corresponding feature sequences, and calculating the ratio to the original audio feature sequence, obtaining the volume adjustment parameters, and then adjusting the volume of the second recorded audio so that it is consistent with the volume of the first recorded audio.

Benefits of technology

It effectively avoids the abrupt volume, improves the overall effect and user experience of recording audio, so that the re-recorded audio is consistent with the first recorded audio volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863953B_ABST
    Figure CN114863953B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a volume adjustment method, apparatus, storage medium, and computing device. The method includes: obtaining a first recorded audio and a second recorded audio of a song; wherein the second recorded audio includes audio recorded additionally for a target audio segment in the first recorded audio; extracting a first feature sequence corresponding to the first recorded audio; and extracting a second feature sequence corresponding to the second recorded audio; calculating a ratio of the first feature sequence to a third feature sequence corresponding to the original singer's audio of the song to obtain a first ratio sequence; and calculating a ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence; calculating a volume adjustment parameter based on the first ratio sequence and the second ratio sequence, and adjusting the volume of the second recorded audio based on the volume adjustment parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a method and apparatus for volume adjustment, a storage medium, and a computing device. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure described in the specification. The description herein is not admitted to be prior art by including it in this section.

[0003] In some audio recording scenarios (such as karaoke applications), users can record audio (such as recording the audio of a song), and in some cases, when the user is not satisfied with a certain target audio segment in the first recorded audio that has been recorded, the user can also re-record the target audio segment separately, so as to replace the target audio segment in the first recorded audio with the re-recorded second recorded audio.

[0004] However, when re-recording, due to the lack of foreshadowing of the front and rear audio, it is very difficult for the user to control the volume, and it is easy to have the situation that the volume of the first recorded audio is inconsistent with the volume of the second recorded audio; thus, there is a phenomenon of abrupt volume in terms of hearing. The occurrence of this phenomenon will seriously reduce the overall effect of the recorded audio, thereby reducing the recording experience. Summary of the Invention

[0005] In a first aspect of the embodiments of the present disclosure, a method for volume adjustment is provided, and the method includes:

[0006] Obtain a first recorded audio and a second recorded audio of a song; wherein, the second recorded audio includes audio recorded separately for a target audio segment in the first recorded audio;

[0007] Extract a first feature sequence corresponding to the first recorded audio; and extract a second feature sequence corresponding to the second recorded audio;

[0008] Calculate the ratio of the first feature sequence to a third feature sequence corresponding to the original audio of the song to obtain a first ratio sequence; and calculate the ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence;

[0009] Calculate a volume adjustment parameter based on the first ratio sequence and the second ratio sequence, and adjust the volume of the second recorded audio based on the volume adjustment parameter.

[0010] Optionally, the first feature sequence, the second feature sequence, and the third feature sequence are obtained by using the same feature extraction method, and the feature extraction method includes:

[0011] Obtain the start and end time information of each line of lyrics in the song;

[0012] Based on the start and end time information of the lyrics, divide the audio to be processed into several audio segments, and calculate the features corresponding to each audio segment; wherein, each audio segment corresponds to one sentence of lyrics.

[0013] Based on the features corresponding to each audio segment, construct a feature sequence corresponding to the audio to be processed; wherein, when the audio to be processed is the first recorded audio, the feature is the first feature, and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature, and the feature sequence is the second feature sequence; when the audio to be processed is the original singer audio, the feature is the third feature, and the feature sequence is the third feature sequence.

[0014] Optionally, the calculating the features corresponding to each audio segment includes:

[0015] For each audio segment, perform frame division on the audio segment based on a preset frame length and the number of samples to obtain multiple audio frames of the frame length;

[0016] For each audio frame, sample the audio signals of the number of samples from the audio frame, and calculate the single-frame feature of the audio frame based on the audio signals of the number of samples;

[0017] Based on the single-frame features of each audio frame within each audio segment, construct the features corresponding to each audio segment.

[0018] Optionally, the start and end time information includes the start moment of each sentence of lyrics and the duration of each sentence of lyrics; or, the start and end time information includes the start moment and the end moment of each sentence of lyrics.

[0019] Optionally, before calculating the volume adjustment parameter based on the first ratio sequence and the second ratio sequence, it further includes:

[0020] Perform anomaly recognition on the ratios in the first ratio sequence and the second ratio sequence, and filter out the abnormal ratios in the first ratio sequence and the second ratio sequence.

[0021] Optionally, the performing anomaly recognition on the ratios in the first ratio sequence and the second ratio sequence includes:

[0022] Obtain the scoring scores of each sentence of lyrics in the first recorded audio and the scoring scores of each sentence of lyrics in the second recorded audio from the scoring system;

[0023] Determine the ratios mapped by the scoring scores lower than the preset score as abnormal ratios.

[0024] Optionally, the anomaly recognition of the ratios in the first ratio sequence and the second ratio sequence includes:

[0025] Determine the ratios outside the preset threshold range in the first ratio sequence and the second ratio sequence as abnormal ratios.

[0026] Optionally, the anomaly recognition of the ratios in the first ratio sequence and the second ratio sequence includes:

[0027] Based on the Pauta criterion, determine the outlier ratios in the first ratio sequence and the second ratio sequence as abnormal ratios; the outlier ratios refer to the ratios greater than three times the standard deviation ratio of the sequence where they are located.

[0028] Optionally, the filtering of the abnormal ratios in the first ratio sequence and the second ratio sequence includes:

[0029] Set the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

[0030] Optionally, the calculation of the volume adjustment parameter based on the first ratio sequence and the second ratio sequence includes:

[0031] Calculate the song integrity; wherein, the song integrity indicates whether the target audio segment in the first recorded audio and the second recorded audio are both sung completely;

[0032] According to the result of the song integrity, adopt the calculation method corresponding to the result, and calculate the volume adjustment parameter based on the first ratio sequence and the second ratio sequence.

[0033] Optionally, the calculation of the song integrity includes:

[0034] Obtain the first scoring sequence for each line of lyrics in the target audio segment from the scoring system, and the second scoring sequence for each line of lyrics in the second recorded audio;

[0035] Perform binarization processing on the first scoring sequence and the second scoring sequence based on a preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and the scoring scores less than the preset score to a second value;

[0036] Count the first quantity of the first value in the first scoring sequence, and the second quantity of the first value in the second scoring sequence;

[0037] If the first quantity is equal to the number of lyric sentences in the target audio segment and the second quantity is equal to the number of lyric sentences in the second recorded audio, then the result of determining the song integrity is complete; otherwise, the result of determining the song integrity is incomplete.

[0038] Optionally, according to the result of the song integrity, adopting a calculation method corresponding to the result, and calculating a volume adjustment parameter based on the first ratio sequence and the second ratio sequence includes:

[0039] If the result of the song integrity is complete, obtain a third ratio sequence corresponding to the target audio segment in the first ratio sequence;

[0040] Calculate a volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

[0041] Optionally, according to the result of the song integrity, adopting a calculation method corresponding to the result, and calculating a volume adjustment parameter based on the first ratio sequence and the second ratio sequence includes:

[0042] If the result of the song integrity is incomplete, calculate the effective mean of the first ratio sequence and the effective mean of the second ratio sequence respectively;

[0043] Calculate a volume adjustment parameter according to the first effective mean and the second effective mean.

[0044] Optionally, adjusting the volume of the second recorded audio based on the volume adjustment parameter includes:

[0045] Multiply the second recorded audio by the volume adjustment parameter.

[0046] Optionally, it further includes:

[0047] Replace the target audio segment in the first recorded audio with the volume-adjusted second recorded audio.

[0048] Optionally, the target audio segment includes an abnormal audio segment in the first recorded audio.

[0049] Optionally, the first recorded audio is a recorded first dry audio, and the second recorded audio is a recorded second dry audio; the original singer audio is a third dry audio of the original singer of the song.

[0050] Optionally, the features in the first feature sequence, the second feature sequence, and the third feature sequence include features representing the energy strength of the audio signal.

[0051] Optionally, the feature representing the energy strength of the audio signal includes the root mean square energy feature of the audio signal.

[0052] In a second aspect of the embodiments of the present disclosure, a volume adjustment device is provided. The device includes:

[0053] An acquisition unit configured to acquire a first recorded audio and a second recorded audio of a song. Wherein, the second recorded audio includes an audio separately recorded for a target audio segment in the first recorded audio;

[0054] An extraction unit configured to extract a first feature sequence corresponding to the first recorded audio; and extract a second feature sequence corresponding to the second recorded audio;

[0055] A calculation unit configured to calculate a ratio of the first feature sequence to a third feature sequence corresponding to the original singer's audio of the song to obtain a first ratio sequence; and calculate a ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence;

[0056] An adjustment unit configured to calculate a volume adjustment parameter based on the first ratio sequence and the second ratio sequence, and adjust the volume of the second recorded audio based on the volume adjustment parameter.

[0057] Optionally, the first feature sequence, the second feature sequence, and the third feature sequence are obtained by using the same feature extraction unit. The feature extraction unit includes:

[0058] An acquisition subunit configured to acquire start and end time information of each line of lyrics in the song;

[0059] A division subunit configured to divide the audio to be processed into a plurality of audio segments based on the start and end time information of the lyrics. Wherein, each audio segment corresponds to one line of lyrics;

[0060] A calculation subunit configured to calculate the feature corresponding to each audio segment;

[0061] A construction subunit configured to construct a feature sequence corresponding to the audio to be processed based on the features corresponding to each audio segment. Wherein, when the audio to be processed is the first recorded audio, the feature is the first feature and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature and the feature sequence is the second feature sequence; when the audio to be processed is the original singer's audio, the feature is the third feature and the feature sequence is the third feature sequence.

[0062] Optionally, the calculation subunit is further configured to, for each audio segment, perform framing processing on the audio segment based on a preset framing length and the number of samples, to obtain a plurality of audio frames with the framing length; for each audio frame, sample the number of audio signals from the audio frame, and calculate a single-frame feature of the audio frame based on the number of sampled audio signals; and construct a feature corresponding to each audio segment based on the single-frame features of the audio frames within each audio segment.

[0063] Optionally, the start and end time information includes the start time and the duration of each lyric sentence; or, the start and end time information includes the start time and the end time of each lyric sentence.

[0064] Optionally, before the adjustment unit, it further includes:

[0065] An identification unit that performs anomaly identification on the ratios in the first ratio sequence and the second ratio sequence;

[0066] A filtering unit that filters out the abnormal ratios in the first ratio sequence and the second ratio sequence.

[0067] Optionally, the identification unit is further configured to obtain the scoring scores of each lyric sentence in the first recorded audio and the scoring scores of each lyric sentence in the second recorded audio from a scoring system, and determine the ratios mapped by the scoring scores lower than a preset score as abnormal ratios.

[0068] Optionally, the identification unit is further configured to determine the ratios outside a preset threshold range in the first ratio sequence and the second ratio sequence as abnormal ratios.

[0069] Optionally, the identification unit is further configured to an identification subunit that, based on the Pauta criterion, determines the outlier ratios in the first ratio sequence and the second ratio sequence as abnormal ratios; the outlier ratio refers to a ratio greater than three times the standard deviation ratio of the sequence where it is located.

[0070] Optionally, the filtering unit is further configured to set the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

[0071] Optionally, the adjustment unit includes:

[0072] An integrity calculation subunit that calculates the song integrity; where the song integrity indicates whether the target audio segment in the first recorded audio and the second recorded audio are both sung completely;

[0073] A parameter calculation subunit that, according to the result of the song integrity, adopts a calculation method corresponding to the result, and calculates a volume adjustment parameter based on the first ratio sequence and the second ratio sequence.

[0074] Optionally, the integrity calculation subunit includes:

[0075] An acquisition subunit that acquires, from a scoring system, a first scoring sequence for each lyric sentence in the target audio segment and a second scoring sequence for each lyric sentence in the second recorded audio;

[0076] A processing subunit that performs binarization processing on the first scoring sequence and the second scoring sequence based on a preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and set the scoring scores less than the preset score to a second value;

[0077] A statistical subunit that statistically counts a first quantity of the first value in the first scoring sequence and a second quantity of the first value in the second scoring sequence;

[0078] A determination subunit that, if the first quantity is equal to the number of lyric sentences in the target audio segment and the second quantity is equal to the number of lyric sentences in the second recorded audio, determines that the result of the song integrity is complete; otherwise, determines that the result of the song integrity is incomplete.

[0079] Optionally, the parameter calculation subunit is further configured to, when the result of the song integrity is complete, acquire a third ratio sequence corresponding to the target audio segment in the first ratio sequence; and calculate a volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

[0080] Optionally, the parameter calculation subunit is further configured to, when the result of the song integrity is incomplete, respectively calculate an effective mean of the first ratio sequence and an effective mean of the second ratio sequence; and calculate a volume adjustment parameter according to the first effective mean and the second effective mean.

[0081] Optionally, the adjustment unit is further configured to multiply the second recorded audio by the volume adjustment parameter.

[0082] Optionally, it further includes:

[0083] A processing unit that replaces the target audio segment in the first recorded audio with the second recorded audio after volume adjustment.

[0084] Optionally, the target audio segment includes an abnormal audio segment in the first recorded audio.

[0085] Optionally, the first recorded audio is a recorded first dry audio, the second recorded audio is a recorded second dry audio; and the original singer audio is a third dry audio of the original singer of the song.

[0086] Optionally, the features in the first feature sequence, the second feature sequence, and the third feature sequence include features representing the energy strength of the audio signal.

[0087] Optionally, the features representing the energy strength of the audio signal include the root mean square energy feature of the audio signal.

[0088] In the third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, including:

[0089] When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the volume adjustment method as described in any of the preceding items.

[0090] In the fourth aspect of the embodiments of the present disclosure, a computing device is provided, including:

[0091] A processor;

[0092] A memory for storing executable instructions of the processor;

[0093] Wherein, the processor is configured to execute the executable instructions to implement the volume adjustment method as described in any of the preceding items.

[0094] According to the volume adjustment solution provided by the embodiments of the present disclosure, the original audio of the song is used to calculate the first recorded audio and the second recorded audio to determine the volume adjustment parameter required to adjust the volume of the second recorded audio to that of the first recorded audio; thus, after adjusting the volume of the second recorded audio based on this volume adjustment parameter, the volume of the first recorded audio and the second recorded audio will not produce a sudden volume change, thereby improving the overall effect of the recorded audio and contributing to improving the recording experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, wherein:

[0096] Figure 1 Schematically shows a schematic diagram of the volume adjustment system architecture provided by the present disclosure;

[0097] Figure 2 Schematically shows a schematic diagram of the recorded audio in the KTV APP provided by the present disclosure;

[0098] Figure 3 Schematically shows a schematic diagram of the re-recorded audio provided by the present disclosure;

[0099] Figure 4Schematically shows a schematic diagram of the volume adjustment method provided by the present disclosure;

[0100] Figure 5 Schematically shows a schematic diagram of the medium provided by the present disclosure;

[0101] Figure 6 Schematically shows a schematic diagram of the volume adjustment device provided by the present disclosure;

[0102] Figure 7 Schematically shows a schematic diagram of the computing device provided by the present disclosure.

[0103] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0104] Hereinafter, the principles and spirit of the present disclosure will be described with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to convey the scope of the present disclosure fully to those skilled in the art.

[0105] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an equipment, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0106] According to the embodiments of the present disclosure, a volume adjustment method, a computer-readable storage medium, a device, and a computing device are provided.

[0107] In this article, it should be understood that the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0108] Hereinafter, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be explained in detail. Summary of the Invention

[0110] As mentioned above, when the user re-records, it is difficult to control the volume size due to the lack of foreshadowing of the front and rear audio, and it is easy to have inconsistent volumes of the first recorded audio and the second recorded audio; thus, there is a phenomenon of abrupt volume in terms of hearing. The occurrence of this phenomenon will seriously reduce the overall effect of the recorded audio, thereby reducing the recording experience.

[0111] To this end, this specification aims to provide a solution that can automatically adjust the volume of the re-recorded audio, so that the volume of the second recorded audio in the re-recording is roughly equivalent to or even the same as that of the first recorded audio in the initial recording, thereby avoiding the phenomenon of abrupt volume changes.

[0112] Specifically, this specification calculates the first recorded audio and the second recorded audio using the original audio of the song to determine the volume adjustment parameter required to adjust the volume of the second recorded audio to that of the first recorded audio. In this way, after adjusting the volume of the second recorded audio based on this volume adjustment parameter, the first recorded audio and the second recorded audio will not have an abrupt volume change phenomenon, thereby improving the overall effect of the recorded audio and contributing to an improved recording experience.

[0113] After introducing the basic principles of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.

[0114] Overview of Application Scenarios

[0115] Figure 1 A schematic diagram showing an exemplary volume adjustment system architecture applicable herein is shown. Figure 1 In this, various network nodes can achieve information communication via the network, and then complete interaction and data processing. The system architecture concept diagram can include a server 105 that communicates data with one or more clients 106 via a network 112, and a database 115 that can be integrated into the server 105 or independent of the server 105.

[0116] Each network 112 can include wired or wireless telecommunications devices, and the network devices on which the clients 106 are based can exchange data through the wired or wireless telecommunications devices. For example, each network 112 can include a local area network ("LAN"), a wide area network ("WAN"), an intranet, the Internet, a mobile phone network, a virtual private network (VPN), a cellular or other mobile communication network, Bluetooth, NFC, or any combination thereof. In the discussion of exemplary embodiments, it should be understood that the terms "data" and "information" can be used interchangeably herein to refer to text, images, audio, video, or any other form of information that can exist in a computer-based environment.

[0117] Each network device on which the clients 106 are based can include a device having a communication module capable of sending and receiving data via the network 112. For example, each network device on which the clients 106 are based can include a server, a desktop computer, a laptop computer, a tablet computer, a smartphone, a handheld computer, a personal digital assistant ("PDA"), or any other wired or wireless processor-driven device. In Figure 1In the exemplary embodiments depicted, the network device on which the client 106 is based can be operated by a user.

[0118] The user can use an application such as a web browser application or a stand-alone application to view, download, upload, or otherwise access files or web pages via the network 112. The network includes a wired or wireless telecommunications system or device through which network devices (including the server 105 and the client 106) can exchange data. For example, the network 112 can include a local area network ("LAN"), a wide area network ("WAN"), an intranet, the Internet, a storage area network (SAN), a personal area network (PAN), a metropolitan area network (MAN), a wireless local area network (WLAN), a virtual private network (VPN), a cellular or other mobile communication network, Bluetooth, NFC, or any combination thereof, or any other suitable architecture or system that facilitates the communication of signals, data, and / or messages. In the discussion of the exemplary embodiments, it should be understood that the terms "data" and "information" may be used interchangeably herein to refer to text, images, audio, video, or any other form of information that may exist in a computer-based environment.

[0119] An application such as a web browser application or a stand-alone application can interact with a web server (or other server, such as a singing platform, a KTV platform, etc.) connected to the network 112 to complete the interaction.

[0120] Figure 1 In the figure, the computing device (not shown in the figure) that can be in an integrated or separate relationship with the server 105, especially in the latter case, can generally be connected through an internal network or a private network, or can also be connected through an encrypted public network. In particular, when in an integrated relationship, a connection in the form of a more efficient and faster transmission speed internal bus may be adopted. This computing device, whether in an integrated or separate relationship, can access the database 115 directly or through the server 105.

[0121] By appropriately programming the computer device, the implementation of the methods in this specification can be controlled by such instructions. In particular, when in an integrated relationship, the transactions processed by the computer device can be regarded as the processing of the server 105 without special distinction.

[0122] Taking the scenario of the KTV service as an example, the above-mentioned client can include a client installed with a KTV APP; the above-mentioned server can include a service platform corresponding to the KTV APP.

[0123] The following will be described in conjunction with Figure 2 the schematic diagram of recording audio in the KTV APP shown below.

[0124] When implemented, the user can open the KTV app on the client; and select the name of the song they want to sing from the song list. As Figure 2 shown in the song list interface 21, there are several song names displayed. When the user clicks on the control corresponding to "Song Name 3", the client responds to the click of this control and jumps from the song list interface 21 to the KTV entry interface 22.

[0125] Furthermore, the user can click on the control corresponding to "KTV", and the client responds to the click of this control and jumps from the KTV entry interface 22 to the KTV recording interface 23.

[0126] In the KTV recording interface 23, a control 24 corresponding to starting recording is displayed. After this control 24 is triggered, the user's singing voice will be collected by the activated audio receiving device, thereby obtaining the first recorded audio.

[0127] During the recording process, the KTV recording interface 23 can also display the sound wave image 25 of "recording sound wave dynamics".

[0128] Generally, after the recording duration reaches the preset duration (usually the song duration corresponding to the song name), the client can jump from the KTV recording interface 23 to the KTV upload interface 26.

[0129] In the KTV upload interface 26, several controls are displayed, such as the "Preview" control for previewing the first recorded audio, the "Rerecord" control for rerecording, and the "Upload" control for uploading the recorded singing voice information, etc.

[0130] When the user previews the first recorded audio and finds that a certain audio segment is not satisfactory, they can trigger the "Rerecord" control to rerecord, so as to enter the interface for rerecording the audio as Figure 3 shown.

[0131] As Figure 3 shown, in the interface for rerecording the audio, the progress bar of the first recorded audio can be displayed, and the user can specify a certain audio segment in this progress bar for rerecording. For example, Figure 3 in the figure, the user selects the audio segment from 15 seconds to 30 seconds for rerecording. In addition, the corresponding lyrics prompt information for the rerecorded audio can be displayed in this interface.

[0132] When the user clicks on the control corresponding to "Rerecord this audio segment", the user can rerecord this audio segment. Similarly, the user's singing voice will be collected by the activated audio receiving device, thereby obtaining the second recorded audio after rerecording.

[0133] Exemplary Method

[0134] Next, in combination with the Figure 1 application scenario shown, refer toFigure 4 A method for volume adjustment according to an exemplary embodiment of the present disclosure will be described. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0135] As Figure 4 shown, the volume adjustment method can be applied to an electronic device, and the method may include the following steps:

[0136] Step 210: Obtain a first recorded audio and a second recorded audio of the song; wherein, the second recorded audio includes audio recorded separately for a target audio segment in the first recorded audio.

[0137] In this specification, the target audio segment may include an abnormal audio segment in the first recorded audio.

[0138] Among them, the abnormal audio segment can be automatically recognized by the electronic device. For example, the electronic device scores the first recorded audio based on an existing scoring mechanism. If the score of a certain audio segment is lower than the threshold, then this audio segment can be determined as an abnormal audio segment.

[0139] The abnormal audio segment can also be selected by the user. For example, if the user is not satisfied with a certain audio segment in the first recorded audio, then this audio segment can be manually marked, so that the electronic device can determine the marked audio segment as an abnormal audio segment.

[0140] Through the above examples, the target audio segment that needs to be re-recorded can be determined from the first recorded audio, and then the target audio segment can be re-recorded separately to obtain the second recorded audio.

[0141] Step 220: Extract a first feature sequence corresponding to the first recorded audio; and extract a second feature sequence corresponding to the second recorded audio.

[0142] In this specification, the first feature sequence, the second feature sequence, and the third feature sequence in step 230 can be obtained by using the same feature extraction method, and the feature extraction method may include the following steps:

[0143] Step A1: Obtain the start and end time information of each line of lyrics in the song.

[0144] In practical applications, in addition to the song file, a song usually also includes a lyrics file. Generally, the lyrics file can include files in formats such as LRC (Lyric). In these lyrics files, the start and end time information of each line of lyrics in the song is recorded. Therefore, the start and end time information of each line of lyrics in the song can be obtained from the lyrics file related to the song.

[0145] In this specification, the start and end time information may include the start moment of each line of lyrics and the duration of each line of lyrics; alternatively, the start and end time information includes the start moment and the end moment of each line of lyrics.

[0146] Step A2: Based on the start and end time information of the lyrics, divide the audio to be processed into several audio segments, and calculate the features corresponding to each audio segment; wherein, each audio segment corresponds to one line of lyrics.

[0147] Step A3: Based on the features corresponding to each audio segment, construct the feature sequence corresponding to the audio to be processed; wherein, when the audio to be processed is the first recorded audio, the feature is the first feature and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature and the feature sequence is the second feature sequence; when the audio to be processed is the original singer audio, the feature is the third feature and the feature sequence is the third feature sequence.

[0148] Taking the first recorded audio as an example, using the start and end time information of each line of lyrics in the first recorded audio, the first recorded audio can be divided into several audio segments such that each audio segment corresponds to one line of lyrics; then calculate the first feature corresponding to each audio segment.

[0149] In an exemplary embodiment, the calculating the features corresponding to each audio segment may include:

[0150] For each audio segment, based on a preset frame length for frame division and the number of samplings, perform frame division processing on the audio segment to obtain multiple audio frames with the frame length for frame division;

[0151] For each audio frame, sample the audio signal with the number of samplings from the audio frame, and based on the audio signal with the number of samplings, calculate the single-frame feature of the audio frame;

[0152] Based on the single-frame features of each audio frame within each audio segment, construct the features corresponding to each audio segment.

[0153] In this example, the frame length for frame division refers to the duration of 1 frame, and the number of samplings refers to the number of samplings within 1 frame.

[0154] Taking the first recorded audio as an example, assume that the first recorded audio is 100 seconds long. If the preset frame length for frame segmentation is 1 second, then by performing frame segmentation at 1 - second intervals, 100 audio frames, each 1 second long, can be obtained. Further, assume that the number of samples is 10. Then, for each audio frame, 10 audio signals are sampled within 1 second, and based on these 10 audio signals, the single - frame feature of the audio frame is calculated.

[0155] Among them, the single - frame feature of the audio frame can be calculated through the following formula 1:

[0156]

[0157] Among them, X(l) represents the single - frame feature of the l - th frame, sqrt is the square - root function, x(l,t) represents the sampling value of the t - th sampled audio signal in the l - th frame, L represents the total number of frames after frame segmentation, and T represents the number of samples.

[0158] After calculating the single - frame features of each audio frame within the audio segment, the feature corresponding to each audio segment can be constructed; and based on the features corresponding to each audio segment, the feature sequence corresponding to the audio to be processed is constructed as shown in the following formula 2:

[0159]

[0160] Among them, X(n,l) represents the single - frame feature of the l - th frame in the n - th line of lyrics, sqrt is the square - root function, x(n,l,t) represents the sampling value of the t - th sampled audio signal in the l - th frame in the n - th line of lyrics, N represents the total number of lines of lyrics in the song, L represents the total number of frames after frame segmentation, and T represents the number of samples.

[0161] In this specification, the features in the first feature sequence, the second feature sequence, and the third feature sequence may include features representing the energy strength of the audio signal. For example, the feature representing the energy strength of the audio signal may include the root - mean - square energy feature (Root Mean Square, RMS) of the audio signal.

[0162] In this specification, the first recorded audio is the first dry - recorded audio, and the second recorded audio is the second dry - recorded audio; the original - singer audio is the third dry - recorded audio of the original singer of the song.

[0163] Among them, dry voice, also known as naked voice, is an audio term. Generally, it refers to pure human voice that has not undergone any post - processing or processing after recording. The corresponding human voice that has undergone post - processing or processing (such as reverberation, delay, etc.) is called wet voice.

[0164] Since the dry voice is pure human voice without accompaniment and harmony, the features extracted from the dry voice audio will not be affected by the accompaniment and harmony in terms of accuracy. Therefore, the results calculated using the dry voice audio are more accurate.

[0165] Step 230: Calculate the ratio of the first feature sequence to the third feature sequence corresponding to the original audio of the song to obtain a first ratio sequence; and calculate the ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence.

[0166] Assume that the first feature sequence corresponding to the first recorded audio is X 1 (n, l), the second feature sequence corresponding to the second recorded audio is X 2 (n, l), and the third feature sequence corresponding to the original audio is X 3 (n, l);

[0167] Then the calculation of the first ratio sequence R 1 (n, l) is shown in Formula 3 below:

[0168]

[0169] where N 1 and N 2 respectively represent the lyric numbers at the start and end of the first recorded audio, that is, the first recorded audio starts recording from the N 1 th lyric and ends recording at the N 2 th lyric.

[0170] In addition, the calculation of the second ratio sequence R 2 (n, l) is shown in Formula 4 below:

[0171]

[0172] where N 3 and N 4 respectively represent the lyric numbers at the start and end of the second recorded audio, that is, the second recorded audio starts recording from the N 3 th lyric and ends recording at the N 4 th lyric.

[0173] In an exemplary embodiment, after calculating the first ratio sequence and the second ratio sequence, it may further include:

[0174] Perform anomaly recognition on the ratios in the first ratio sequence and the second ratio sequence, and filter out the abnormal ratios in the first ratio sequence and the second ratio sequence.

[0175] During the actual frequency recording process, external influences such as recording equipment, background noise, and user operations may be encountered, which may lead to abnormal audio signals in the recorded audio. These abnormal audio signals, in turn, may cause abnormal ratios to appear in the ratio sequence. Since abnormal ratios can easily affect the accuracy of the volume adjustment parameters, in order to improve the accuracy of the volume adjustment parameters, it is necessary to identify and filter out the abnormal ratios in the first ratio sequence and the second ratio sequence.

[0176] In an exemplary embodiment, after identifying the abnormal ratios, filtering the abnormal ratios in the first ratio sequence and the second ratio sequence may include:

[0177] Setting the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

[0178] Here, setting the abnormal ratios to a preset value is to retain the time information of the abnormal ratios, thereby ensuring that the accuracy of the calculation result will not be affected due to the loss of information in the time dimension when calculating the volume adjustment parameters later.

[0179] The following provides several exemplary abnormal identification methods:

[0180] In the first implementation method, identifying the abnormal ratios in the first ratio sequence and the second ratio sequence includes:

[0181] Obtaining the scoring scores of each line of lyrics in the first recorded audio and the scoring scores of each line of lyrics in the second recorded audio from the scoring system;

[0182] Determining the ratios mapped by the scoring scores lower than the preset score as abnormal ratios.

[0183] In this example, by referring to the scoring scores of each line of lyrics in the first recorded audio and the second recorded audio, and determining the ratios mapped by the scoring scores lower than the preset score as abnormal ratios, the purpose is to eliminate the audio parts in the first recorded audio and the second recorded audio that the user did not sing or sang invalidly, and improve the reliability and robustness of the ratio calculation.

[0184] The abnormal identification and filtering for the first ratio sequence are as shown in Formula 5 below:

[0185]

[0186] Where, R 11 (n, l) is the first ratio sequence after filtering abnormal ratios by the first method, and R 1 (n, l) is the first ratio sequence before filtering abnormal ratios, is the scoring score of the nth line of lyrics in the first recorded audio, and S thr is the preset score.

[0187] The anomaly recognition and filtering for the second ratio sequence are as shown in Formula 6 below:

[0188]

[0189] where R 21 (n, l) is the second ratio sequence after filtering out anomalous ratios by the first method, and R 2 (n, l) is the second ratio sequence before filtering out anomalous ratios, is the scoring value of the nth line of lyrics in the second recorded audio, and S thr is the preset score value.

[0190] In the above Formulas 5 and 6, for the ratios mapped by the scoring values that are higher than or equal to the preset score value (i.e., normal ratios), no processing is performed, and the ratios mapped by the scoring values that are lower than the preset score value (i.e., anomalous ratios) are set to 0 (i.e., the preset value).

[0191] In the second implementation method, the anomaly recognition of the ratios in the first ratio sequence and the second ratio sequence includes:

[0192] Determining the ratios outside the preset threshold range in the first ratio sequence and the second ratio sequence as anomalous ratios.

[0193] In this example, based on business experience, a preset threshold can be set, and the ratios outside the preset threshold are determined as anomalous ratios. The purpose is to eliminate the audio parts with too high or too low audio energy in the first recorded audio and the second recorded audio, and improve the reliability and robustness of ratio calculation.

[0194] The anomaly recognition and filtering for the first ratio sequence are as shown in Formula 7 below:

[0195]

[0196] where R 12 (n, l) is the first ratio sequence after filtering out anomalous ratios by the second method, and R 1 (n, l) is the first ratio sequence before filtering out anomalous ratios, and R MIN is the minimum value of the preset threshold, and R MAX is the maximum value of the preset threshold.

[0197] The anomaly recognition and filtering for the second ratio sequence are as shown in Formula 8 below:

[0198]

[0199] where R 22(n, l) is the second ratio sequence after filtering abnormal ratios by the second method, R 2 (n, l) is the second ratio sequence before filtering abnormal ratios, R MIN is the minimum value of the preset threshold, R MAX is the maximum value of the preset threshold.

[0200] In the above formulas 7 and 8, the ratio greater than R MIN and less than R MAX is the normal ratio and is not processed. The ratio less than R MIN or greater than R MAX is the abnormal ratio, and the abnormal ratio is set to 0.

[0201] In the third implementation method, the abnormal identification of the ratios in the first ratio sequence and the second ratio sequence includes:

[0202] Based on the Pauta criterion, the outlier ratios in the first ratio sequence and the second ratio sequence are determined as abnormal ratios; the outlier ratio refers to the ratio greater than three times the standard deviation ratio of the sequence where it is located.

[0203] In this example, the outlier ratios in the first ratio sequence and the second ratio sequence are removed by the Pauta criterion to improve the stability of the volume adjustment parameters.

[0204] The abnormal identification and filtering for the first ratio sequence are as shown in formula 9 below:

[0205]

[0206] where, R 13 (n, l) is the first ratio sequence after filtering abnormal ratios by the third method, R 1 (n, l) is the first ratio sequence before filtering abnormal ratios, μ 1 is the mean of the first ratio sequence, σ 1 is the variance of the first ratio sequence.

[0207] The abnormal identification and filtering for the second ratio sequence are as shown in formula 10 below:

[0208]

[0209] where, R 23 (n, l) is the second ratio sequence after filtering abnormal ratios by the third method, R 2 (n, l) is the second ratio sequence before filtering abnormal ratios, μ 2 is the mean of the second ratio sequence, σ 2 is the variance of the second ratio sequence.

[0210] σ in the above formula 9 1 and σ in formula 10 2 The weight 3 is only an example and can be flexibly adjusted according to requirements in actual applications.

[0211] It should be noted that the above abnormal recognition method is only an example. In actual applications, any other abnormal recognition method can also be adopted, and different abnormal recognition methods can also be combined to improve the abnormal recognition efficiency.

[0212] Step 240: Calculate a volume adjustment parameter based on the first ratio sequence and the second ratio sequence, and adjust the volume of the second recorded audio based on the volume adjustment parameter.

[0213] After calculating the first ratio sequence and the second ratio sequence, a volume adjustment parameter can further be calculated based on the first ratio sequence and the second ratio sequence.

[0214] In an exemplary embodiment, the calculating the volume adjustment parameter based on the first ratio sequence and the second ratio sequence may include:

[0215] Step 241: Calculate the song integrity; wherein, the song integrity indicates whether the target audio segment in the first recorded audio and the second recorded audio are both sung completely.

[0216] The calculating the song integrity may further include:

[0217] Step B1, obtain a first scoring sequence for each line of lyrics in the target audio segment from a scoring system, and a second scoring sequence for each line of lyrics in the second recorded audio;

[0218] Step B2, perform binarization processing on the first scoring sequence and the second scoring sequence based on a preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and set the scoring scores less than the preset score to a second value.

[0219] In this step, the binarization processing for the first scoring sequence can refer to the following formula 11:

[0220]

[0221] Wherein, is the first scoring sequence of the scoring system, C 1 (n) is the binarized first scoring sequence, S thr is the preset score.

[0222] The binarization processing for the second scoring sequence can refer to the following formula 12:

[0223]

[0224] Among them, is the second scoring sequence of the scoring system, C 2 (n) is the second scoring sequence after binarization, S thr is the preset score value.

[0225] In the above formulas 11 and 12, the scoring values in the first scoring sequence and the second scoring sequence that are greater than or equal to the preset score value are set to 1, and the scoring values less than the preset score value are set to 0.

[0226] Step B3, count the first quantity of the first value in the first scoring sequence and the second quantity of the first value in the second scoring sequence.

[0227] In this step, the first quantity of the first value in the first scoring sequence is counted as shown in the following formula 13:

[0228]

[0229] Among them, C 1 represents the first quantity, and the meaning expressed by formula 13 is to count the number of the first value from the N 3 th line of lyrics to the N 4 th line of lyrics in the first scoring sequence.

[0230] The second quantity of the first value in the second scoring sequence is counted as shown in the following formula 14:

[0231]

[0232] Among them, C 2 represents the second quantity, and the meaning expressed by formula 14 is to count the number of the first value from the N 3 th line of lyrics to the N 4 th line of lyrics in the first scoring sequence.

[0233] Step B4, if the first quantity is equal to the number of lines of lyrics in the target audio segment and the second quantity is equal to the number of lines of lyrics in the second recorded audio, then determine that the result of the song integrity is complete; otherwise, determine that the result of the song integrity is incomplete.

[0234] In this step, the song integrity is as shown in the following formula 15:

[0235]

[0236] Among them, Confidenc is the song integrity, Confidence = 1 means complete, Confidenc = 0 means incomplete; N 4 -N3 Indicates the number of lyric sentences in the target audio segment.

[0237] If the first quantity C 1 is equal to the number of lyric sentences N in the target audio segment 4 -N 3 , it means that the lyrics of the target audio segment in the first recorded audio are sung completely. Similarly, if the second quantity C 2 is equal to the number of lyric sentences N in the second recorded audio 4 -N 3 , it also means that the lyrics of the second recorded audio are sung completely; therefore, the result of the song integrity is also complete.

[0238] Conversely, if the first quantity C 1 is not equal to the number of lyric sentences N in the target audio segment 4 -N 3 , it means that the lyrics of the target audio segment in the first recorded audio are not sung completely; or if the second quantity C 2 is not equal to the number of lyric sentences N in the second recorded audio 4 -N 3 , it also means that the lyrics of the second recorded audio are not sung completely; therefore, the result of the song integrity is also incomplete.

[0239] Step 245: According to the result of the song integrity, adopt a calculation method corresponding to the result, and calculate the volume adjustment parameter based on the first ratio sequence and the second ratio sequence.

[0240] By determining the result of the song integrity, different calculation methods can be used to calculate the volume adjustment parameter.

[0241] In one implementation, if the result of the song integrity is complete, obtain the third ratio sequence corresponding to the target audio segment in the first ratio sequence; calculate the volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

[0242] In this example, if the result of the song integrity is complete, the volume adjustment parameter V can be calculated using the following formula 16:

[0243]

[0244] where K represents the effective number of the ratio sequence, R 1 (n, l)' is the first ratio sequence after abnormal ratio processing, and R 2 (n, l)' is the second ratio sequence after abnormal ratio processing.

[0245] The preset values in the first ratio sequence and the second ratio sequence (for example, the abnormal ratios in the aforementioned Formulas 5 to 10 are set to 0) are regarded as invalid values, and the number of the remaining non-preset values is the valid number.

[0246] Since the song integrity is complete, the valid numbers of the first ratio sequence and the second ratio sequence are the same; thus, it is unified as K in Formula 16.

[0247] In another implementation, if the result of the song integrity is incomplete, calculate the valid mean of the first ratio sequence and the valid mean of the second ratio sequence respectively; calculate the volume adjustment parameter according to the first valid mean and the second valid mean.

[0248] In this example, if the result of the song integrity is incomplete, the volume adjustment parameter V can be calculated using the following Formula 17:

[0249]

[0250] Wherein, K1 represents the valid number of the first ratio sequence, and K2 represents the valid number of the second ratio sequence.

[0251] After calculating the volume adjustment parameter V, adjusting the volume of the second recorded audio based on the volume adjustment parameter may include:

[0252] Multiply the second recorded audio by the volume adjustment parameter to obtain a second recorded audio whose volume is close to or the same as that of the first recorded audio.

[0253] Further, after obtaining the second recorded audio with the volume adjusted, the target audio segment in the first recorded audio may also be replaced with the second recorded audio with the volume adjusted.

[0254] Through the above example, since the volume of the second recorded audio after the volume adjustment is comparable to that of the first recorded audio, when the target audio segment in the first recorded audio is replaced with the second recorded audio, there will be no sudden volume phenomenon when playing the replaced first recorded audio; thus, the overall effect of the recorded audio is improved, which helps to improve the recording experience.

[0255] Exemplary Medium

[0256] After introducing the method of the exemplary embodiments of the present disclosure, next, refer to Figure 5 to describe the medium of the exemplary embodiments of the present disclosure.

[0257] In this exemplary embodiment, the above method can be implemented by a program product. For example, a portable compact disc read-only memory (CD-ROM) can be adopted, which includes program code and can run on a device, such as a personal computer. However, the program product of the present disclosure is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0258] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0259] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program used by or in conjunction with an instruction execution system, apparatus, or device.

[0260] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0261] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the C language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0262] In summary, the present disclosure can provide a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can be enabled to execute the foregoing Figure 4 volume adjustment method embodiments shown.

[0263] Exemplary Apparatus

[0264] After introducing the medium of the exemplary embodiments of the present disclosure, next, reference is made to Figure 6 describe the device of the exemplary embodiments of the present disclosure.

[0265] Figure 6 A block diagram of a volume adjustment device according to an embodiment of the present disclosure is schematically shown, corresponding to the foregoing Figure 4 method embodiments shown. The volume adjustment device may include:

[0266] An acquisition unit 610, which acquires a first recorded audio and a second recorded audio of a song; wherein, the second recorded audio includes an audio recorded additionally for a target audio segment in the first recorded audio;

[0267] An extraction unit 620, which extracts a first feature sequence corresponding to the first recorded audio; and extracts a second feature sequence corresponding to the second recorded audio;

[0268] A calculation unit 630, which calculates a ratio of the first feature sequence to a third feature sequence corresponding to the original singer audio of the song to obtain a first ratio sequence; and calculates a ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence;

[0269] An adjustment unit 640, which calculates a volume adjustment parameter based on the first ratio sequence and the second ratio sequence, and adjusts the volume of the second recorded audio based on the volume adjustment parameter.

[0270] Optionally, the first feature sequence, the second feature sequence, and the third feature sequence are obtained by using the same feature extraction unit, and the feature extraction unit includes:

[0271] An acquisition subunit 622, which acquires the start and end time information of each sentence of lyrics in the song;

[0272] A division subunit 624, which divides the audio to be processed into several audio segments based on the start and end time information of the lyrics; wherein, each audio segment corresponds to one sentence of lyrics;

[0273] A calculation subunit 626, which calculates the features corresponding to each audio segment;

[0274] A construction subunit 628 constructs a feature sequence corresponding to the audio to be processed based on the features corresponding to each audio segment. Wherein, when the audio to be processed is the first recorded audio, the feature is the first feature and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature and the feature sequence is the second feature sequence; when the audio to be processed is the original singer audio, the feature is the third feature and the feature sequence is the third feature sequence.

[0275] Optionally, the calculation subunit 626 is further configured to, for each audio segment, perform frame segmentation on the audio segment based on a preset frame segmentation length and the number of samples to obtain multiple audio frames of the frame segmentation length; for each audio frame, sample the number of audio signals from the audio frame, and calculate a single-frame feature of the audio frame based on the number of sampled audio signals; and construct a feature corresponding to each audio segment based on the single-frame features of each audio frame within each audio segment.

[0276] Optionally, the start and end time information includes the start moment and the duration of each lyric sentence; or, the start and end time information includes the start moment and the end moment of each lyric sentence.

[0277] Optionally, before the adjustment unit 640, it further includes:

[0278] An identification unit 634 performs anomaly identification on the ratios in the first ratio sequence and the second ratio sequence;

[0279] A filtering unit 636 filters out the abnormal ratios in the first ratio sequence and the second ratio sequence.

[0280] Optionally, the identification unit 634 is further configured to obtain the scoring scores of each lyric sentence in the first recorded audio and the scoring scores of each lyric sentence in the second recorded audio from a scoring system, and determine the ratios mapped by the scoring scores lower than a preset score as abnormal ratios.

[0281] Optionally, the identification unit 634 is further configured to determine the ratios outside a preset threshold range in the first ratio sequence and the second ratio sequence as abnormal ratios.

[0282] Optionally, the identification unit 634 is further configured to be an identification subunit that determines the outlier ratios in the first ratio sequence and the second ratio sequence as abnormal ratios based on the Pauta criterion; the outlier ratio refers to a ratio greater than three times the standard deviation ratio of the sequence where it is located.

[0283] Optionally, the filtering unit 636 is further configured to set the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

[0284] Optionally, the adjustment unit 640 includes:

[0285] An integrity calculation subunit that calculates the integrity of the song; wherein the integrity of the song indicates whether the target audio segment in the first recorded audio and the second recorded audio are both sung completely;

[0286] A parameter calculation subunit that, according to the result of the song integrity, adopts a calculation method corresponding to the result, and calculates a volume adjustment parameter based on the first ratio sequence and the second ratio sequence.

[0287] Optionally, the integrity calculation subunit includes:

[0288] An acquisition subunit that acquires a first scoring sequence for each lyric sentence in the target audio segment from a scoring system, and a second scoring sequence for each lyric sentence in the second recorded audio;

[0289] A processing subunit that performs binarization processing on the first scoring sequence and the second scoring sequence based on a preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and the scoring scores less than the preset score to a second value;

[0290] A statistics subunit that statistics a first quantity of the first value in the first scoring sequence and a second quantity of the first value in the second scoring sequence;

[0291] A determination subunit that, if the first quantity is equal to the number of lyric sentences in the target audio segment and the second quantity is equal to the number of lyric sentences in the second recorded audio, determines that the result of the song integrity is complete; otherwise, determines that the result of the song integrity is incomplete.

[0292] Optionally, when the result of the song integrity is complete, the parameter calculation subunit is further configured to acquire a third ratio sequence corresponding to the target audio segment in the first ratio sequence; and calculate a volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

[0293] Optionally, when the result of the song integrity is incomplete, the parameter calculation subunit is further configured to calculate an effective mean of the first ratio sequence and an effective mean of the second ratio sequence respectively; and calculate a volume adjustment parameter according to the first effective mean and the second effective mean.

[0294] Optionally, the adjustment unit 640 is further configured to multiply the second recorded audio by the volume adjustment parameter.

[0295] Optionally, the device further includes:

[0296] The processing unit 650 replaces the target audio segment in the first recorded audio with the second recorded audio after volume adjustment.

[0297] Optionally, the target audio segment includes the audio segment with anomalies in the first recorded audio.

[0298] Optionally, the first recorded audio is the first dry audio recorded, the second recorded audio is the second dry audio recorded; the original singer audio is the third dry audio of the original singer of the song.

[0299] Optionally, the features in the first feature sequence, the second feature sequence, and the third feature sequence include features representing the energy strength of the audio signal.

[0300] Optionally, the feature representing the energy strength of the audio signal includes the root mean square energy feature of the audio signal.

[0301] Exemplary Computing Device

[0302] After introducing the methods, media, and devices of the exemplary embodiments of the present disclosure, next, reference is made to Figure 7 describe the computing device of the exemplary embodiments of the present disclosure.

[0303] Figure 7 The computing device 1500 shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0304] As Figure 7 shown, the computing device 1500 is presented in the form of a general-purpose computing device. The components of the computing device 1500 may include, but are not limited to: at least one of the above-mentioned processing units 1501, at least one of the above-mentioned storage units 1502, and a bus 1503 connecting different system components (including the processing unit 1501 and the storage unit 1502).

[0305] The bus 1503 includes a data bus, a control bus, and an address bus.

[0306] The storage unit 1502 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 15021 and / or a cache memory 15022, and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 15023.

[0307] The storage unit 1502 may also include a program / utility 15025 having a set (at least one) of program modules 15024. Such program modules 15024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0308] The computing device 1500 may also communicate with one or more external devices 1504 (such as a keyboard, a pointing device, etc.).

[0309] Such communication may be carried out through an input / output (I / O) interface 1505. Also, the computing device 1500 may further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1506. As Figure 7 shown, the network adapter 1506 communicates with other modules of the computing device 1500 through a bus 1503. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 1500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0310] In summary, the present disclosure may provide a computing device, including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the foregoing Figure 4 shown volume adjustment method.

[0311] It should be noted that although several units / modules or sub-units / modules of the volume adjustment device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.

[0312] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0313] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of various aspects does not mean that the features in these aspects cannot be combined for benefits. Such division is only for the convenience of expression. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A volume adjustment method, including: obtaining a first recorded audio and a second recorded audio of a song; wherein, the second recorded audio includes audio recorded additionally for a target audio segment in the first recorded audio; extracting a first feature sequence corresponding to the first recorded audio; and, extracting a second feature sequence corresponding to the second recorded audio; calculating a ratio of the first feature sequence to a third feature sequence corresponding to the original singer's audio of the song to obtain a first ratio sequence; and, calculating a ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence; calculating the song integrity; wherein, the song integrity indicates whether both the target audio segment in the first recorded audio and the second recorded audio are sung completely; if the result of the song integrity is incomplete, respectively calculate a first effective mean of the first ratio sequence and a second effective mean of the second ratio sequence; calculate a volume adjustment parameter based on the first effective mean and the second effective mean, and adjust the volume of the second recorded audio based on the volume adjustment parameter; wherein, the first recorded audio is a first dry audio recorded, the second recorded audio is a second dry audio recorded; the original singer's audio is a third dry audio of the original singer of the song.

2. The method according to claim 1, wherein the first feature sequence, the second feature sequence and the third feature sequence are obtained by using the same feature extraction method, and the feature extraction method includes: obtaining the start and end time information of each line of lyrics in the song; based on the start and end time information of the lyrics, dividing the audio to be processed into several audio segments, and calculating the feature corresponding to each audio segment; wherein, each audio segment corresponds to one line of lyrics; constructing a feature sequence corresponding to the audio to be processed based on the features corresponding to each audio segment; wherein, when the audio to be processed is the first recorded audio, the feature is the first feature, and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature, and the feature sequence is the second feature sequence; when the audio to be processed is the original singer's audio, the feature is the third feature, and the feature sequence is the third feature sequence.

3. The method according to claim 2, wherein the calculating the feature corresponding to each audio segment includes: for each audio segment, performing frame division processing on the audio segment based on a preset frame length and the number of samples for frame division to obtain multiple audio frames of the frame length for frame division; for each audio frame, sampling the number of audio signals of the sampling number from the audio frame, and calculating the single-frame feature of the audio frame based on the number of audio signals of the sampling number; constructing the feature corresponding to each audio segment based on the single-frame features of each audio frame within each audio segment.

4. The method according to claim 2, wherein the start and end time information includes the start moment of each line of lyrics and the duration of each line of lyrics; or, the start and end time information includes the start moment and the end moment of each line of lyrics.

5. The method according to claim 1, before calculating the volume adjustment parameter based on the first ratio sequence and the second ratio sequence, further includes: Perform anomaly identification on the ratios in the first ratio sequence and the second ratio sequence, and filter out the abnormal ratios in the first ratio sequence and the second ratio sequence.

6. The method according to claim 5, wherein the performing anomaly identification on the ratios in the first ratio sequence and the second ratio sequence comprises: Obtain the scoring scores for each line of lyrics in the first recorded audio and the scoring scores for each line of lyrics in the second recorded audio from the scoring system; Determine the ratios mapped by the scoring scores lower than the preset score as abnormal ratios.

7. The method according to claim 5, wherein the performing anomaly identification on the ratios in the first ratio sequence and the second ratio sequence comprises: Determine the ratios outside the preset threshold range in the first ratio sequence and the second ratio sequence as abnormal ratios.

8. The method according to claim 5, wherein the performing anomaly identification on the ratios in the first ratio sequence and the second ratio sequence comprises: Based on the Pauta criterion, determine the outlier ratios in the first ratio sequence and the second ratio sequence as abnormal ratios; The outlier ratio refers to a ratio greater than three times the standard deviation ratio of the sequence where it is located.

9. The method according to claim 5, wherein the filtering out the abnormal ratios in the first ratio sequence and the second ratio sequence comprises: Set the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

10. The method according to claim 1, wherein the calculating the song integrity comprises: Obtain the first scoring sequence for each line of lyrics in the target audio segment and the second scoring sequence for each line of lyrics in the second recorded audio from the scoring system; Perform binarization processing on the first scoring sequence and the second scoring sequence based on the preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and set the scoring scores less than the preset score to a second value; Count the first quantity of the first value in the first scoring sequence and the second quantity of the first value in the second scoring sequence; If the first quantity is equal to the number of lyrics lines in the target audio segment and the second quantity is equal to the number of lyrics lines in the second recorded audio, determine that the result of the song integrity is complete; Otherwise, determine that the result of the song integrity is incomplete.

11. The method according to claim 1, wherein according to the result of the song integrity, adopt a calculation method corresponding to the result, and calculate the volume adjustment parameter based on the first ratio sequence and the second ratio sequence comprises: If the result of the song integrity is complete, obtain the third ratio sequence corresponding to the target audio segment in the first ratio sequence; Calculate the volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

12. The method according to claim 1, wherein the adjusting the volume of the second recorded audio based on the volume adjustment parameter comprises: Multiply the second recorded audio by the volume adjustment parameter.

13. The method according to claim 1, further comprises: Replace the target audio segment in the first recorded audio with the second recorded audio after volume adjustment.

14. The method according to claim 13, wherein the target audio segment includes the audio segment with abnormality in the first recorded audio.

15. The method according to claim 1, wherein the features in the first feature sequence, the second feature sequence and the third feature sequence include the features representing the energy strength of the audio signal.

16. The method according to claim 15, wherein the features representing the energy strength of the audio signal include the root mean square energy feature of the audio signal.

17. A volume adjustment device comprising: an acquisition unit configured to acquire a first recorded audio and a second recorded audio of a song; wherein, the second recorded audio includes the audio recorded additionally for the target audio segment in the first recorded audio; an extraction unit configured to extract a first feature sequence corresponding to the first recorded audio; and extract a second feature sequence corresponding to the second recorded audio; a calculation unit configured to calculate the ratio of the first feature sequence to a third feature sequence corresponding to the original audio of the song to obtain a first ratio sequence; and calculate the ratio of the second feature sequence to the third feature sequence to obtain a second ratio sequence; a completeness calculation subunit configured to calculate the song completeness; wherein, the song completeness indicates whether the target audio segment in the first recorded audio and the second recorded audio are both sung completely; a parameter calculation subunit configured to, when the result of the song completeness is incomplete, calculate a first effective mean of the first ratio sequence and a second effective mean of the second ratio sequence respectively; calculate a volume adjustment parameter according to the first effective mean and the second effective mean, and adjust the volume of the second recorded audio based on the volume adjustment parameter; wherein, the first recorded audio is the first dry audio recorded, the second recorded audio is the second dry audio recorded; the original audio is the third dry audio of the song's original singer.

18. The device according to claim 17, wherein the first feature sequence, the second feature sequence and the third feature sequence are obtained by using the same feature extraction unit, and the feature extraction unit comprises: an acquisition subunit configured to acquire the start and end time information of each line of lyrics in the song; a division subunit configured to divide the audio to be processed into several audio segments based on the start and end time information of the lyrics; wherein, each audio segment corresponds to one line of lyrics; a calculation subunit configured to calculate the feature corresponding to each audio segment; a construction subunit configured to construct a feature sequence corresponding to the audio to be processed based on the features corresponding to each audio segment; wherein, when the audio to be processed is the first recorded audio, the feature is the first feature and the feature sequence is the first feature sequence; when the audio to be processed is the second recorded audio, the feature is the second feature and the feature sequence is the second feature sequence; when the audio to be processed is the original audio, the feature is the third feature and the feature sequence is the third feature sequence.

19. The device according to claim 18, wherein the calculation sub-unit is further configured to, for each audio segment, perform frame division processing on the audio segment based on a preset frame division frame length and the number of samples, so as to obtain a plurality of audio frames with the frame division frame length. ; For each audio frame, sample the number of audio signals from the audio frame, and calculate the single-frame feature of the audio frame based on the number of sampled audio signals. Construct the feature corresponding to each audio segment based on the single-frame features of each audio frame within each audio segment.

20. The device according to claim 18, wherein the start and end time information includes the start time of each lyric sentence and the duration of each lyric sentence; or, the start and end time information includes the start time and end time of each lyric sentence.

21. The device according to claim 17, before the integrity calculation sub-unit, further comprises: An identification unit configured to perform anomaly identification on the ratios in the first ratio sequence and the second ratio sequence. A filtering unit configured to filter out the abnormal ratios in the first ratio sequence and the second ratio sequence.

22. The device according to claim 21, wherein the identification unit is further configured to obtain the scoring scores of each lyric sentence in the first recorded audio and the scoring scores of each lyric sentence in the second recorded audio from a scoring system, and determine the ratios mapped by the scoring scores lower than a preset score as abnormal ratios.

23. The device according to claim 21, wherein the identification unit is further configured to determine the ratios outside a preset threshold range in the first ratio sequence and the second ratio sequence as abnormal ratios.

24. The device according to claim 21, wherein the identification unit is further configured to, based on the Pauta criterion, determine the outlier ratios in the first ratio sequence and the second ratio sequence as abnormal ratios; the outlier ratio refers to a ratio greater than three times the standard deviation ratio of the sequence where it is located.

25. The device according to claim 21, wherein the filtering unit is further configured to set the abnormal ratios in the first ratio sequence and the second ratio sequence to a preset value.

26. The device according to claim 17, wherein the integrity calculation sub-unit comprises: An acquisition sub-unit configured to obtain a first scoring sequence for each lyric sentence in the target audio segment and a second scoring sequence for each lyric sentence in the second recorded audio from a scoring system. A processing sub-unit configured to perform binarization processing on the first scoring sequence and the second scoring sequence based on a preset score, so as to set the scoring scores greater than or equal to the preset score in the first scoring sequence and the second scoring sequence to a first value, and set the scoring scores less than the preset score to a second value. A statistics sub-unit configured to count a first quantity of the first value in the first scoring sequence and a second quantity of the first value in the second scoring sequence. A determination sub-unit configured to, if the first quantity is equal to the number of lyric sentences in the target audio segment and the second quantity is equal to the number of lyric sentences in the second recorded audio, determine that the result of the song integrity is complete. Conversely, determine that the result of the song integrity is incomplete.

27. The apparatus according to claim 17, wherein the parameter calculation sub-unit is further configured to, when the result of the song integrity is complete, obtain a third ratio sequence corresponding to the target audio segment in the first ratio sequence; and calculate a volume adjustment parameter based on the third ratio sequence and the second ratio sequence.

28. The apparatus according to claim 17, wherein the parameter calculation sub-unit is further configured to multiply the second recorded audio by the volume adjustment parameter.

29. The apparatus according to claim 17, further comprising: a processing unit, configured to replace the target audio segment in the first recorded audio with the second recorded audio after volume adjustment.

30. The apparatus according to claim 29, wherein the target audio segment includes an abnormal audio segment in the first recorded audio.

31. The apparatus according to claim 17, wherein the features in the first feature sequence, the second feature sequence, and the third feature sequence include features representing the energy strength of the audio signal.

32. The apparatus according to claim 31, wherein the features representing the energy strength of the audio signal include the root mean square energy feature of the audio signal.

33. A computer-readable storage medium, comprising: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the volume adjustment method according to any one of claims 1-16.

34. A computing device, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the volume adjustment method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Video correction method and system, terminal equipment and storage medium

    CN108962293A

  • Method and device for determining volume adjustment proportion information, equipment and storage medium

    CN110688082A