Audio processing method, device, computing equipment and medium

The human voice audio and accompaniment audio in the audio are automatically extracted through neural network technology, and combined with lyrics files to generate word-by-word lyrics and MIDI files, solving the problem of low audio processing efficiency in the existing technology, realizing automatic and efficient audio processing.

CN114220410BActive Publication Date: 2025-05-23HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111322201.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-05-23
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

In the prior art, the method of artificially producing accompaniment and word-by-word lyrics is complex and time-consuming, resulting in low audio processing efficiency.

Method used

By obtaining the pending audio and using neural network technology to automatically extract human voice audio and accompaniment audio, combining lyrics files to generate word-by-word lyrics files and MIDI files, the audio is automated.

Benefits of technology

It improves the processing efficiency of the to-process audio, reduces the complexity and time-consuming of manual operations, and realizes the automatic processing process of audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220410B_ABST
    Figure CN114220410B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide an audio processing method, apparatus, computing device and medium. The method automatically creates a target data record for the audio to be processed after acquiring the audio to be processed, and automatically triggers the generation process of target data such as accompaniment audio and a second lyrics file, and then automatically adds data information to the target data record after generating the target data, and the data information can reflect the information of the data generated based on the operation of the audio to be processed, so that the target data as the accompaniment material can be obtained later through the data information recorded in the target data record, so as to realize the automatic processing of the audio to be processed, thereby improving the processing efficiency of the audio to be processed.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An audio processing method, It is characterized in that The method comprises: In response to acquiring the audio to be processed, creating a target data record for the audio to be processed, wherein the target data record is used to trigger automatic generation of target data; In response to the target data record being created, target data is generated based on the audio to be processed and a first lyrics file corresponding to the audio to be processed, the first lyrics file being a lyrics file divided sentence by sentence, the target data at least including an accompaniment audio and a second lyrics file, the second lyrics file being a lyrics file divided word by word, and the generated target data is used to trigger automatic addition of data information; In response to the target data being generated, data information is added to the target data record to obtain the target data based on the data information.

2. The method according to claim 1, It is characterized in that The generating target data based on the audio to be processed and a first lyrics file corresponding to the audio to be processed includes: Acquire the vocal audio and the accompaniment audio from the audio to be processed; The second lyrics file is generated based on the vocal audio and the first lyrics file.

3. The method according to claim 2, It is characterized in that The step of obtaining the vocal audio and the accompaniment audio from the audio to be processed includes: Input the audio to be processed into a human voice extraction neural network and an accompaniment extraction neural network respectively, perform downsampling processing and a first convolution processing on the audio to be processed through the human voice extraction neural network respectively to obtain the human voice audio, and perform downsampling processing and a second convolution processing on the audio to be processed through the accompaniment extraction neural network to obtain the accompaniment audio; Among them, the network parameters used by the vocal extraction neural network for the first convolution processing are different from the network parameters used by the accompaniment extraction neural network for the second convolution processing.

4. The method according to claim 2, It is characterized in that The step of generating the second lyrics file based on the human voice audio and the first lyrics file comprises: Inputting the human voice audio into a speech recognition neural network, and outputting a first phoneme corresponding to the human voice audio and a timestamp corresponding to the first phoneme through the speech recognition neural network; Obtaining the second phoneme corresponding to each word in the first lyrics file; Based on the first phoneme and the timestamp corresponding to the first phoneme, and the second phoneme, the timestamps corresponding to the respective characters in the first lyrics file are determined to obtain the second lyrics file.

5. The method according to claim 1, It is characterized in that The target data also includes a Musical Instrument Digital Interface (MIDI) file; The generating target data based on the audio to be processed and the first lyrics file corresponding to the audio to be processed also includes: The MIDI file is generated based on the audio to be processed and the second lyrics file.

6. The method according to claim 5, It is characterized in that The step of generating the MIDI file based on the audio to be processed and the second lyrics file comprises: Inputting the audio to be processed into a melody extraction neural network, and outputting the fundamental tone of the audio to be processed through the melody extraction neural network; The MIDI file is generated based on the fundamental pitch of the audio to be processed and the second lyrics file.

7. The method according to claim 5 or 6, It is characterized in that The adding data information to the target data record to obtain the target data based on the data information includes: The data information is added to the target data record, and audio description information is generated based on the data information, so as to obtain the target data based on the audio description information.

8. The method according to claim 1, It is characterized in that The data information includes audio data information associated with the audio to be processed, and the audio data information includes at least an audio identifier of the audio to be processed and a storage location of the audio to be processed.

9. The method according to claim 1, It is characterized in that The data information also includes a data identifier of the target data and a storage location of the target data.

10. The method according to claim 7, It is characterized in that The target data record also includes status information, and the status information is used to record the processing progress of the audio to be processed.

11. The method according to claim 10, It is characterized in that The status information includes any of the following: first state information, where the first state information is used to indicate starting to process the audio to be processed; second state information, wherein the second state information is used to indicate that the accompaniment audio is being generated; third state information, the third state information is used to indicate that the second lyrics file is being generated; fourth status information, the fourth status information being used to indicate that the MIDI file is being generated; fifth state information, the fifth state information being used to indicate that the audio description information is being generated; The sixth state information is used to indicate that the accompaniment audio, the second lyrics file, the MIDI file and the audio description information have been generated.

12. An audio processing device, It is characterized in that The device comprises: A creation module, configured to create a target data record for the audio to be processed in response to acquiring the audio to be processed, wherein the target data record is used to trigger automatic generation of target data; A generating module, configured to generate target data based on the audio to be processed and a first lyrics file corresponding to the audio to be processed in response to the target data record being created, wherein the first lyrics file is a lyrics file divided sentence by sentence, and the target data at least includes an accompaniment audio and a second lyrics file, wherein the second lyrics file is a lyrics file divided word by word, and the generated target data is used to trigger automatic addition of data information; An adding module is used for adding data information to the target data record in response to the target data having been generated, so as to obtain the target data based on the data information.

13. The device according to claim 12, It is characterized in that The generating module, when used to generate target data based on the audio to be processed and the first lyrics file corresponding to the audio to be processed, comprises an acquiring unit and a generating unit; The acquisition unit is used to acquire the human voice audio and the accompaniment audio from the audio to be processed; The generating unit is used to generate the second lyrics file based on the vocal audio and the first lyrics file.

14. The device according to claim 13, It is characterized in that The acquisition unit, when used to acquire the human voice audio and the accompaniment audio from the audio to be processed, is specifically used to: Input the audio to be processed into a human voice extraction neural network and an accompaniment extraction neural network respectively, perform downsampling processing and a first convolution processing on the audio to be processed through the human voice extraction neural network respectively to obtain the human voice audio, and perform downsampling processing and a second convolution processing on the audio to be processed through the accompaniment extraction neural network to obtain the accompaniment audio; Among them, the network parameters used by the vocal extraction neural network for the first convolution processing are different from the network parameters used by the accompaniment extraction neural network for the second convolution processing.

15. The device according to claim 13, It is characterized in that The generating unit, when used to generate the second lyrics file based on the human voice audio and the first lyrics file, is specifically used to: Inputting the human voice audio into a speech recognition neural network, and outputting a first phoneme corresponding to the human voice audio and a timestamp corresponding to the first phoneme through the speech recognition neural network; Obtaining the second phoneme corresponding to each word in the first lyrics file; Based on the first phoneme and the timestamp corresponding to the first phoneme, and the second phoneme, the timestamps corresponding to the respective characters in the first lyrics file are determined to obtain the second lyrics file.

16. The device according to claim 12, It is characterized in that The target data also includes a Musical Instrument Digital Interface (MIDI) file; The generating module, when used to generate target data based on the audio to be processed and the first lyrics file corresponding to the audio to be processed, is further used to: The MIDI file is generated based on the audio to be processed and the second lyrics file.

17. The device according to claim 16, It is characterized in that The generating module, when used to generate the MIDI file based on the audio to be processed and the second lyrics file, is specifically used to: Inputting the audio to be processed into a melody extraction neural network, and outputting the fundamental tone of the audio to be processed through the melody extraction neural network; The MIDI file is generated based on the fundamental pitch of the audio to be processed and the second lyrics file.

18. The device according to claim 16 or 17, It is characterized in that The adding module, when used to add data information to the target data record to obtain the target data based on the data information, is specifically used to: The data information is added to the target data record, and audio description information is generated based on the data information, so as to obtain the target data based on the audio description information.

19. The device according to claim 12, It is characterized in that The data information includes audio data information associated with the audio to be processed, and the audio data information includes at least an audio identifier of the audio to be processed and a storage location of the audio to be processed.

20. The device according to claim 12, It is characterized in that The data information also includes a data identifier of the target data and a storage location of the target data.

21. The device according to claim 18, It is characterized in that The target data record also includes status information, and the status information is used to record the processing progress of the audio to be processed.

22. The device according to claim 21, It is characterized in that The status information includes any of the following: first state information, where the first state information is used to indicate starting to process the audio to be processed; second state information, wherein the second state information is used to indicate that the accompaniment audio is being generated; third state information, the third state information is used to indicate that the second lyrics file is being generated; fourth status information, the fourth status information being used to indicate that the MIDI file is being generated; fifth state information, the fifth state information being used to indicate that the audio description information is being generated; The sixth state information is used to indicate that the accompaniment audio, the second lyrics file, the MIDI file and the audio description information have been generated.

23. A computing device, It is characterized in that The computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the operations performed by the audio processing method according to any one of claims 1 to 11 when executing the program.

24. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a program, and the program is used by a processor to execute the operations performed by the audio processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Audio data processing method and device

    CN107103915A

  • Method and device for acquiring lyrics data

    CN108228903A

  • Accompaniment and human voice extraction method and device and word-by-word lyric generation method and device

    CN111540374A