Data labeling method and device, equipment and storage medium

The audio element annotation model trained by AI algorithms and the graphical user interface, combined with professional tools, solves the problems of low efficiency and low accuracy in manual song annotation, realizes efficient and accurate automatic annotation and real-time verification, and simplifies the annotation process.

CN114911451BActive Publication Date: 2025-10-17NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210612437.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-10-17
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

Manual song tagging is inefficient and inaccurate, lacks the assistance of artificial intelligence algorithms, and the labeled data is easily lost. The verification mechanism is imperfect, and there is a lack of professional efficiency-enhancing tools and dynamic listening feedback functions.

Method used

AI algorithms are used to train the audio element annotation model, automatically annotating preset elements of music audio, providing graphical user interface display and adjustment functions, and combining dynamic listening feedback and professional tools such as BPM meters to achieve real-time verification and data storage.

Benefits of technology

It improves annotation efficiency and accuracy, reduces labor costs, lowers the risk of data loss, provides real-time verification and dynamic listening functions, simplifies operating procedures, and improves the concentration of annotation personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114911451B_ABST
    Figure CN114911451B_ABST
Patent Text Reader

Abstract

The application provides a data labeling method and device, equipment and a storage medium, wherein the method comprises: obtaining music audio to be labeled, processing the music audio according to at least one audio element labeling model trained to obtain at least one preset music element data, wherein the at least one audio element labeling model is trained according to a plurality of music audio samples, the plurality of music audio samples are labeled with sample labeling data for corresponding preset music elements, and corresponding data is displayed at the audio track of at least one preset music element on a graphical user interface. In the present application, the preset music element data for the music audio is automatically labeled by means of an algorithm, the labeling efficiency is high, and the accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, in particular to a data labeling method and device, equipment and a storage medium. BACKGROUND

[0002] In the field of artificial intelligence, data labeling is often needed to train neural network models. Data labeling is a behavior of processing artificial intelligence learning data, and the types of data labeling include but are not limited to image labeling, speech labeling, text labeling, and video labeling.

[0003] Currently, the labeling of songs is usually performed by manually listening to the sound and manually labeling. For example, by manually listening to the sound, sentences, syllables, and phonemes in the singing are labeled, and beats and lyrics in the music score are labeled.

[0004] However, manual labeling has low labeling efficiency and low labeling accuracy. SUMMARY

[0005] Therefore, the embodiments of the present application provide a data labeling method, device, equipment and storage medium to solve the problem of low labeling efficiency and low labeling accuracy of manual labeling.

[0006] In a first aspect, the embodiments of the present application provide a data labeling method, comprising:

[0007] Obtaining music audio to be labeled;

[0008] Processing the music audio according to at least one audio element labeling model trained to obtain labeling data of at least one preset music element, wherein the at least one audio element labeling model is trained according to a plurality of music audio samples, and the plurality of music audio samples are labeled with sample labeling data corresponding to the preset music element;

[0009] Displaying the corresponding labeling data at the audio track of the at least one preset music element on a graphical user interface.

[0010] In an optional implementation, after displaying the corresponding labeling data at the audio track of the at least one preset music element on the graphical user interface, the method further comprises:

[0011] Adjusting the labeling data of the target music element in response to an adjustment operation on the labeling data of the target music element.

[0012] In an optional implementation, before adjusting the labeling data of the target music element in response to an adjustment operation on the labeling data of the target music element, the method further comprises:

[0013] In response to a playing operation on the annotation data of the target music element, playing audio of the annotation data of the target music element.

[0014] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of a beat, lyrics, an arrangement structure, and a score pitch.

[0015] In an optional implementation, the at least one preset music element includes a beat.

[0016] After displaying the corresponding annotation data at the audio track of the at least one preset music element on the graphical user interface, the method further includes:

[0017] According to the beat annotation data, determining a beat speed of the music audio.

[0018] Displaying the beat speed at a speed track on the graphical user interface.

[0019] In an optional implementation, the method further includes:

[0020] In response to a selection operation on the beat speed measurer displayed on the graphical user interface, displaying a beat touch window and playing a beat sound of the beat annotation data in the music audio, so as to input a touch operation on the beat touch window based on the beat sound.

[0021] In response to the touch operation on the beat touch window, determining a touch speed corresponding to the touch operation.

[0022] Determining the touch speed as the adjusted beat speed.

[0023] In an optional implementation, the method further includes:

[0024] Playing the beat sound of the beat annotation data.

[0025] In response to a selection operation on the beat annotation data, determining part of the annotation data from the beat annotation data.

[0026] In an optional implementation, the at least one preset music element includes lyrics.

[0027] In response to an adjustment operation on the annotation data of the target music element, adjusting the annotation data of the target music element, including:

[0028] In response to a first insertion operation on the input separator, inserting a first separator in the lyrics annotation data.

[0029] In response to an adjustment operation on first annotation data corresponding to the first separator in the lyrics annotation data, adjusting the first annotation data.

[0030] In an optional implementation, the at least one preset music element includes a musical pitch.

[0031] In response to the adjustment operation on the annotation data of the target music element, the annotation data of the target music element is adjusted, including:

[0032] In response to the second insertion operation input for the delimiter, a second delimiter is inserted in the musical pitch annotation data.

[0033] In response to the adjustment operation on the second annotation data corresponding to the second delimiter in the musical pitch annotation data, the second annotation data is adjusted.

[0034] In an optional implementation, the method further includes:

[0035] The precision of the musical pitch annotation data is quantified to the data precision corresponding to the music audio.

[0036] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of a verse, a syllable, a phoneme, a song pitch, and a sound length.

[0037] In an optional implementation, the at least one preset music element includes a verse.

[0038] In response to the adjustment operation on the annotation data of the target music element, the annotation data of the target music element is adjusted, including:

[0039] In response to the third insertion operation input for the delimiter, a third delimiter is inserted in the verse annotation data.

[0040] The selected preset lyrics are inserted into the verse annotation data.

[0041] In an optional implementation, the music audio is processed according to the at least one trained audio element annotation model to obtain at least one preset music element data, including:

[0042] The music audio is processed according to the first audio element annotation model to obtain real-time first music element annotation data.

[0043] The music audio and the first music element annotation data are processed according to the second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes the first audio element annotation model and the second audio element annotation model.

[0044] In a second aspect, the embodiments of the present application further provide a data annotation apparatus, including:

[0045] The obtaining module is configured to obtain music audio to be annotated.

[0046] a processing module configured to process the music audio according to at least one audio element labeling model obtained through training, to obtain at least one preset music element data, wherein the at least one audio element labeling model is obtained through training according to a plurality of music audio samples, and the plurality of music audio samples are labeled with sample labeling data corresponding to the preset music element;

[0047] a display module configured to display the corresponding labeling data at the audio track of the at least one preset music element on the graphical user interface.

[0048] In an optional implementation, the processing module is further configured to:

[0049] adjust the labeling data of the target music element in response to an adjustment operation on the labeling data of the target music element.

[0050] In an optional implementation, the processing module is further configured to:

[0051] play the audio of the labeling data of the target music element in response to a play operation on the labeling data of the target music element.

[0052] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of beat, lyrics, arrangement structure, and score pitch.

[0053] In an optional implementation, the at least one preset music element includes beat.

[0054] The processing module is further configured to determine a beat speed of the music audio according to the beat labeling data.

[0055] The display module is further configured to display the beat speed at a speed track on the graphical user interface.

[0056] In an optional implementation, the display module is further configured to:

[0057] display a beat touch window and play a beat sound of the beat labeling data in the music audio in response to a selection operation on a beat speed measurer displayed on the graphical user interface, so as to input a touch operation on the beat touch window based on the beat sound.

[0058] The processing module is further configured to determine a touch speed corresponding to the touch operation on the beat touch window in response to the touch operation on the beat touch window.

[0059] The touch speed is determined as the adjusted beat speed.

[0060] In an optional implementation, the processing module is further configured to:

[0061] play the beat sound of the beat labeling data.

[0062] In response to the selection operation on the beat annotation data, determine part of the annotation data from the beat annotation data.

[0063] In an optional implementation, the at least one preset music element includes lyrics, and the processing module is specifically configured to:

[0064] In response to a first insertion operation on the separator input, insert a first separator in the lyrics annotation data;

[0065] In response to an adjustment operation on first annotation data corresponding to the first separator in the lyrics annotation data, adjust the first annotation data.

[0066] In an optional implementation, the at least one preset music element includes a score pitch, and the processing module is specifically configured to:

[0067] In response to a second insertion operation on the separator input, insert a second separator in the score pitch annotation data;

[0068] In response to an adjustment operation on second annotation data corresponding to the second separator in the score pitch annotation data, adjust the second annotation data.

[0069] In an optional implementation, the processing module is further configured to:

[0070] Quantize the precision of the score pitch annotation data to the data precision corresponding to the music audio.

[0071] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of the following elements: a phrase, a syllable, a phoneme, a song pitch, and a sound length.

[0072] In an optional implementation, the at least one preset music element includes a phrase, and the processing module is specifically configured to:

[0073] In response to a third insertion operation on the separator input, insert a third separator in the phrase annotation data;

[0074] Insert the selected preset lyrics into the phrase annotation data.

[0075] In an optional implementation, the processing module is specifically configured to:

[0076] Process the music audio according to the first audio element annotation model to obtain first music element annotation data;

[0077] The music audio and the first music element annotation data are processed according to the second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes the first audio element annotation model and the second audio element annotation model.

[0078] In a third aspect, the embodiments of the present application further provide an electronic device, including a processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine readable instructions to perform the data annotation method in any one of the first aspect.

[0079] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the data annotation method in any one of the first aspect.

[0080] The data annotation method, device, equipment and storage medium provided by the present application, wherein the method comprises: obtaining music audio to be annotated, and processing the music audio according to at least one audio element annotation model trained to obtain at least one preset music element data, wherein the at least one audio element annotation model is trained according to a plurality of music audio samples, the plurality of music audio samples are annotated with sample annotation data for corresponding preset music elements, and the corresponding data is displayed at the audio track of at least one preset music element on a graphical user interface. In the present application, the preset music element data for the music audio is automatically annotated by means of algorithm, the annotation efficiency is high, and the accuracy is high.

[0081] In order to make the above objectives, characteristics and advantages of the present application more apparent and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0082] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0083] Figure 1 Flowchart of the data annotation method provided by the embodiments of the present application Figure One ;

[0084] Figure 2 Interface diagram of the music audio uploading interface provided by the embodiments of the present application Figure One ;

[0085] Figure 3 Interface diagram of a music audio uploading interface provided by an embodiment of the present application Figure Two

[0086] Figure 4 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Two

[0087] Figure 5 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Three

[0088] Figure 6 Schematic diagram of a beat touch window provided by an embodiment of the present application

[0089] Figure 7 Interface diagram of beat labeling provided by an embodiment of the present application

[0090] Figure 8 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Four

[0091] Figure 9 Interface diagram of lyric labeling provided by an embodiment of the present application

[0092] Figure 10 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Five

[0093] Figure 11 Interface diagram of pitch labeling provided by an embodiment of the present application

[0094] Figure 12 Interface diagram of arrangement structure labeling provided by an embodiment of the present application

[0095] Figure 13 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Six

[0096] Figure 14 Flow diagram of a data labeling method provided by an embodiment of the present application Figure Seven

[0097] Figure 15 Architecture diagram of data labeling provided by an embodiment of the present application

[0098] Figure 16 Interface diagram of labeled data export provided by an embodiment of the present application

[0099] Figure 17 Structure diagram of a data labeling apparatus provided by an embodiment of the present application ​​​​​​​

[0100] Figure 18 The structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0101] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0102] Before introducing the technical solutions of the present application, the professional terms involved are first explained:

[0103] Music visualization is a non-subjective interpretation and judgment of music expression, and is a presentation technology for understanding, analyzing and comparing the expressiveness and internal structure of music. After the characteristics of music such as waveform, frequency, pitch, tone, rhythm, speed, timbre, etc. are extracted, the music visualization is mapped to the corresponding visual effect.

[0104] Standard pitch: Central A 440Hz, i.e. the frequency specified in Musical Instrument Digital Interface (MIDI) number 69. The base frequency in MIDI can be calculated by the following formula:

[0105]

[0106] P represents a difference of 12 notes in MIDI number.

[0107] 12 equal temperament: the frequency difference is one time per 12 (half) notes.

[0108] Intelligent labeling: building an AI model like a human being requires a large amount of training data. Through data labeling, the model can be trained to understand specific information, so that the model can make decisions and take actions.

[0109] Audio labeling: labeling of singing, music scores, sound effects, etc. Structured audio data produced is used for big data neural network learning.

[0110] Pitch: The human perception of the frequency of a sound, typically, low-pitched sounds have a low pitch, while high-pitched sounds have a high pitch.

[0111] Cent: A logarithmic scale unit used to measure intervals.

[0112] Chord: Two or more different pitches of sound combined together.

[0113] Beat: The regular combination of strong and weak beats, specifically referring to the total length of the notes in each measure of a musical score.

[0114] Phrase: A basic structural unit with a characteristic that constitutes a piece of music.

[0115] Paragraph: The overall framework of a song, including the introduction, verse, chorus, interlude, and ending.

[0116] Beats per minute (BPM): The number of beats per minute.

[0117] Currently, the annotation of songs is generally done manually, and there is no solution using Artificial Intelligence (AI) algorithms to assist in annotation. Manual annotation has the following shortcomings:

[0118] (1) No AI algorithm assistance, existing annotation schemes rely entirely on manual annotation, without AI algorithm assistance to pre-generate annotation data, annotation data cannot add samples, and there is no mature solution for intelligent annotation.

[0119] (2) Low confidence, manual annotation has a high probability of error, which may lead to annotation results that do not conform to the true results.

[0120] (3) Data loss, there is no maintenance of historical data, no improvement of annotation data saving anti-loss mechanism, and no quick positioning of historical editing content.

[0121] (4) Incomplete verification mechanism, the existing annotation scheme does not provide real-time verification function, the annotation personnel cannot dynamically perceive the accuracy of the annotation results, and the lack of real-time feedback leads to easy rework and increased difficulty of modification of the annotation task.

[0122] (5) Lack of professional efficiency tools to assist manual annotation (such as automatic quantization, automatic adsorption, song voice assisted word filling, BPM measurement, and dynamic feedback), the tedious and lengthy annotation work leads to a longer annotation process, causing the annotation personnel to lose focus, prone to errors, and reducing confidence.

[0123] (6) There is no sound dynamic feedback function, and the annotator cannot compare the consistency of the audio and the annotation track, and cannot correct errors in time.

[0124] (7) Lack of perfect shortcut keys.

[0125] Therefore, the present application provides a data annotation method, which is a customized annotation scheme for song score, achieves the purpose of audio annotation based on AI algorithm assisted training, aims to annotate all kinds of music elements to obtain annotation data, and the annotation data can be used for big data neural network learning or used in the form of intermediate state for related links in the industry chain. Under the assistance of AI algorithm, the confidence of the annotation data is greatly improved, and the data export format is also standardized, so that the data set formed by the annotation data is more refined and has strong usability.

[0126] Specifically, the data annotation method of the present application has the following characteristics:

[0127] (1) The pre-annotation function of AI algorithm greatly reduces the labor cost, and after the creation of the audio library, the AI algorithm can generate pre-annotation data such as beat, lyrics, paragraph, chord, and sound for the audio library, with an accuracy of more than 80%, and the accuracy of pre-annotation data for chords can reach more than 90%. The annotation data can further become sample data to train AI algorithm model to form a closed loop, continuously expand new annotation database, and improve the accuracy of annotation algorithm.

[0128] (2) Based on the pre-annotation ability of AI algorithm, the annotator only needs to fine-tune to complete the annotation work, greatly improving the annotation efficiency and confidence.

[0129] (3) Maintain the history record of the editing area, provide a lossless saving mechanism for annotation data, each annotation operation can be maintained as a piece of historical data, which can be quickly positioned to the historical editing content through withdrawal and redo, and the editing area content can be automatically saved at regular intervals, greatly reducing the rework rate caused by data loss.

[0130] (4) Provide real-time verification function, the annotator can dynamically perceive the accuracy of the annotation result in real time, the annotation platform can display the annotation error in real time and prompt the correct annotation format, which facilitates the annotator to find and correct errors in time, greatly reduces the rework rate, and the quality assurance (QUALITY ASSURANCE, QA) personnel re-verify when submitting the annotation data.

[0131] (5) Provide a large number of auxiliary artificial annotation professional efficiency tools (such as automatic quantification, automatic adsorption, auxiliary word filling, BPM meter, etc.), part of the annotation work is handed over to the machine, avoids the long annotation process caused by the tedious and long annotation work, improves the concentration and efficiency of the annotation personnel, is easy to operate, reduces the training cost of the annotation personnel, and reduces the professional threshold of the employees.

[0132] (6) Provide dynamic listening function, can hear the corresponding piano sound feedback of the annotation track in real time, facilitate the annotation personnel to compare the audio file of the annotation task, also support playing the audio track, pitch track and chord track at the same time for comparison annotation, the annotation personnel can compare the consistency of the audio and the annotation track, and correct in time to ensure the accuracy of the annotation data.

[0133] (7) Provide rich shortcut mechanism, the undo and redo of the content of the editing area, the zooming and advancing and retreating of the visible area, the insertion and deletion of the annotation unit separator, the segment selection and saving are all encapsulated with shortcut keys, which is easy to operate and improves the annotation efficiency.

[0134] The data annotation method provided by the application will be described in detail below in combination with several specific embodiments.

[0135] Figure 1 The flowchart of the data annotation method provided by the embodiment of the application Figure One The execution subject of the embodiment can be a mobile phone, a notebook computer, a desktop computer or other electronic devices with data processing capability.

[0136] As shown in Figure 1 , the method can include:

[0137] S101, obtaining music audio to be annotated.

[0138] The music audio to be annotated can be any music audio saved locally by the electronic device, or any music audio in the music library accessible by the electronic device, and the embodiment does not limit this.

[0139] In one of the application scenarios of the embodiment of the application, the annotation platform for music audio annotation can be opened through the electronic device, the annotation platform is provided with an application programming interface (Application Programming Interface, API) such as WebAudio, and the annotation interface of the annotation platform can be provided with an upload control of the music audio. The annotation personnel clicks the upload control to display the music audio upload interface through WebAudio. Figure 2 The interface diagram of the music audio upload interface provided by the embodiment of the application Figure One , Figure 3Interface schematic of a music audio uploading interface provided by an embodiment of the present application Figure Two .

[0140] As shown in Figure 2 , the music audio uploading interface provides options for the source of the music audio, and an annotator can select local uploading to drag (or click to upload) the score audio of the music audio to be annotated to an audio uploading area, and drag (or click to upload) the lyrics of the music audio to be annotated to a lyrics uploading area, where the format of the score audio and the lyrics can be Zip format.

[0141] As shown in Figure 3 , the annotator can select a music library interface to upload the score audio and lyrics within an assigned range in the music library, where Figure 3 the assigned range 1 to 100 indicates that the score audio and lyrics of the first to the 100th pieces in the music library are uploaded.

[0142] After obtaining the music audio to be annotated, the music audio can also be sampled to discretize the continuous audio function into a discrete sequence using a specified sampling frequency. Specifically, by Fourier transform, a signal is decomposed into a set of sinusoidal waves of different frequencies, and the form is transformed from the time domain to the frequency domain. Taking WebAudio as an example, a visual window is selected based on WebAudio to play the music audio, and a fixed sample is selected from the music audio to be annotated for Fourier transform as the music plays, to obtain a visual audio sequence for plotting. In addition, the vocal pitch of the music audio to be annotated can also be obtained based on WebAudio. Specifically, the pitch of the song is obtained through an interface, and all points with a non-zero frequency are connected to form a line segment, and the left and right endpoints of each line segment are extended to draw a parallel line. The vocal pitch is the result of machine sampling, which can reflect the pitch frequency of each sound segment of the music audio, and serves as a reference for subsequent annotation of the score pitch and the vocal pitch.

[0143] Wherein, the display of the audio pitch can be arbitrarily adjusted by the zoom bar and the shortcut key, which is convenient for the annotator to observe. The audio supports play / pause, dynamic positioning, and specified range playback.

[0144] S102, according to the at least one audio element annotation model obtained by training, the music audio is processed to obtain the annotation data of at least one preset music element.

[0145] input the music audio to be annotated into at least one audio element annotation model, the at least one audio element annotation model being configured to process the music audio to obtain annotation data of at least one preset music element, wherein the at least one audio element annotation model is trained according to a plurality of music audio samples, and the plurality of music audio samples are annotated with sample annotation data corresponding to the at least one preset music element. That is, each audio element annotation model is configured to annotate a corresponding audio element in the music audio to obtain annotation data of the corresponding audio element.

[0146] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of beat, lyrics, composition structure, and score pitch.

[0147] The beat refers to a combination rule of strong beats and weak beats, and specifically refers to the total length of notes in each measure in a score. The lyrics refer to the lyrics of the music audio. The composition structure refers to the overall framework of the music audio. The score pitch refers to the pitch of the score of the music audio.

[0148] Taking the beat as an example, the music audio to be annotated is input into a beat annotation model, and the model outputs beat annotation data corresponding to the music audio.

[0149] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of a song sentence, a syllable, a phoneme, a song pitch, and a sound length.

[0150] The song sentence refers to a sentence constituting the music audio. The syllable refers to the smallest phonetic unit of a combination of a single vowel phoneme and a consonant phoneme. The phoneme refers to the smallest phonetic unit divided according to the natural properties of speech. The song pitch refers to the pitch of the song of the music audio. The sound length refers to the length of the song of the music audio, which is determined by the duration of the vibration of the sound source.

[0151] S103, display the corresponding annotation data at the audio track of the at least one preset music element on the graphical user interface.

[0152] The graphical user interface displays an audio track of the at least one preset music element. The at least one music audio annotation model is used to process the music audio to be annotated to obtain at least one preset music audio data. The corresponding annotation can be displayed at the audio track of the at least one preset music element on the graphical user interface. That is, the beat annotation data is displayed at the beat track, the lyrics data is displayed at the lyrics track, and other music elements are displayed in the same manner.

[0153] In the data labeling method of the embodiment, the music audio to be labeled is obtained, and the music audio is processed according to at least one audio element labeling model obtained through training to obtain at least one preset music element data, and the corresponding data is displayed at the audio track of the at least one preset music element on the graphical user interface. The preset music element data of the music audio is automatically labeled by means of an algorithm, and the labeling efficiency is high and the accuracy is high.

[0154] In an optional implementation, after the AI algorithm is used to label the music audio to be labeled, the labeling data can also be manually adjusted, which will be described below. Figure 4

[0155] Figure 4 The flowchart of the data labeling method provided by the embodiment of the application is shown in Figure Two After the corresponding labeling data is displayed at the audio track of the at least one preset music element on the graphical user interface, the method can further include: Figure 4

[0156] S201, in response to an adjustment operation of the labeling data of the target music element, adjusting the labeling data of the target music element.

[0157] After the at least one preset music element data of the music audio is labeled by means of the AI algorithm model, the labeling personnel can also fine-tune the labeling data. The labeling personnel can input an adjustment operation of the labeling data of the target music element, and accordingly, the electronic device can adjust the labeling data of the target music element in response to the adjustment operation of the labeling data of the target music element, wherein the target music element can be any one or more music elements in the at least one preset music element.

[0158] In an optional implementation, before step S201, in response to the adjustment operation of the labeling data of the target music element, adjusting the labeling data of the target music element, the method can further include:

[0159] S202, in response to a playing operation of the labeling data of the target music element, playing the audio of the labeling data of the target music element.

[0160] To realize the adjustment of the labeling data of the target music element, the audio of the labeling data of the target music element can also be played. The electronic device plays the audio of the labeling data of the target music element in response to the playing operation of the labeling data of the target music element. For example, the target music element is a beat, and the audio played is the audio of the beat labeling data, that is, the beat sound. For example, the target music element is a lyric, and the audio played is the audio of the lyric labeling data, that is, the lyric sound. In addition, the audio of the labeling data can also be played in the form of a specific musical instrument sound, such as playing the beat sound and the lyric sound in the form of a piano sound. ​​

[0161] In this way, the labeling personnel can listen to the pre-labeled data of the AI ​​algorithm model to compare the audio of the labeled data with the music audio to be labeled, and by comparing the consistency of the music audio and the audio track data, the labeling personnel can promptly correct the labeled data determined by the AI ​​algorithm model. In other words, based on the listening function, the labeling personnel can input adjustment operations for the labeled data of the target music element to adjust the labeled data of the target music element.

[0162] Taking the target music element as lyrics as an example, for the lyrics annotation data, if the end time point of the last lyrics of a certain sentence is the first time point, and the end time point of the sentence is the second time point, manual correction can be made in time to adjust the end time point of the last lyrics to the second time point to ensure the accuracy of the lyrics annotation data.

[0163] In the data annotation method of this embodiment, in response to a playback operation of the annotated data for the target music element, the audio of the annotated data of the target music element is played, and in response to an adjustment operation of the annotated data for the target music element, the annotated data of the target music element is adjusted. Based on the pre-annotation capability of the AI ​​algorithm model, the annotator can complete the annotation work by making fine adjustments on this basis, greatly improving the annotation efficiency and confidence. It also provides a dynamic listening feedback function, which allows the user to hear the feedback audio corresponding to the audio track in real time, making it convenient for the annotator to compare the annotated data with the music audio. It can also play each audio track simultaneously for comparative annotation, and timely correct errors to ensure the accuracy of the annotated data.

[0164] At least one preset music element includes: beat. In step S103, after the corresponding annotation data is displayed at the audio track of at least one preset music element on the graphical user interface, the beat speed of the music audio to be annotated can also be determined. Figure 3 Provide explanation.

[0165] Figure 5 Schematic diagram of the data annotation method provided in this application embodiment Figure Three ,like Figure 5 As shown, after displaying corresponding annotation data at the audio track of at least one preset music element on the graphical user interface, the method may further include:

[0166] S301: Determine the tempo of the music audio according to the tempo marking data.

[0167] S302: Display the tempo on the tempo track of the graphical user interface.

[0168] The at least one preset music element includes a beat, and the annotation data of the at least one preset music element includes beat annotation data, according to which a beat speed of the music audio can be determined and displayed at a speed track on the graphical user interface, and the speed track is also displayed on the graphical user interface.

[0169] The annotator can determine whether the music audio has a variable speed boundary according to the beat speed displayed by the speed track, and classify the music audio into paragraphs, which can be divided into uniform speed segments, stable variable speed segments and scordatura, where the scordatura refers to free beats with slow speed and irregular rhythm.

[0170] In an optional implementation, the method can further include:

[0171] In response to a selection operation on the beat speed measurer displayed on the graphical user interface, a beat touch window is displayed, and a beat sound of the beat annotation data in the music audio is played, so as to input a touch operation on the beat touch window based on the beat sound.

[0172] In response to the touch operation on the beat touch window, a touch speed corresponding to the touch operation is determined.

[0173] The touch speed is determined as the adjusted beat speed.

[0174] The graphical user interface displays a beat speed measurer, that is, a BPM tool, and the annotator can input a selection operation on the beat speed measurer, and the electronic device displays a beat speed touch window and plays a beat sound of the beat annotation data in the music audio in response to the selection operation on the beat speed measurer, so that the annotator inputs a touch operation on the beat touch window according to the beat speed of the beat sound, and the touch operation can be a click operation.

[0175] Correspondingly, the electronic device determines a touch speed corresponding to the touch operation on the beat touch window, and determines the touch speed as the adjusted beat speed, that is, the beat speed of the beat annotation data obtained based on the AI algorithm model can be corrected through the provided beat speed measurer.

[0176] Figure 6 The beat touch window provided by the embodiments of the present application is shown in the schematic diagram of the beat touch window as shown in Figure 6 The window provides a beat touch window, and the beat touch window has a prompt information of “listen to the beat sound and click the area”, so that the annotator clicks the beat touch window based on the beat sound heard, and the window can also display the beat speed BPM (taking 116.6 as an example) and the number of clicks determined based on the click speed of clicking the beat touch window. Figure 6 ​Figure 6 For example, 8 times), wherein the window is also provided with a reset control, by clicking the reset control, the beat touch window can be clicked again to recalculate the BPM.

[0177] In this way, by providing an auxiliary artificial annotation professional efficiency tool (BPM measurer), part of the annotation work is completely handed over to machine operation, avoiding the lengthening of the annotation process caused by tedious and lengthy annotation work, improving the concentration and efficiency of the annotation personnel, simple operation, reducing the training cost of the annotation personnel, and reducing the professional threshold of the practitioners.

[0178] In an optional implementation, the method can further include:

[0179] Playing the beat sound of the beat annotation data.

[0180] In response to a selection operation on the beat annotation data, determining part of the annotation data from the beat annotation data.

[0181] The annotation personnel can input an activation operation on the audio track with the beat annotation data displayed, and play the beat annotation data in the activated state, so that the annotation personnel determines part of the annotation data from the beat annotation data based on the beat sound, and the electronic device determines part of the annotation data from the beat annotation data in response to the selection operation on the beat annotation data.

[0182] Among them, part of the annotation data can be the accent annotation data, accent, which refers to the strong beat of the music. Then the accent annotation data can also be displayed on the graphical user interface, and the graphical user interface also displays: the accent track.

[0183] In one embodiment of the present application, if the beat speed of the music audio is uniform, it can be determined that the measures after the measures determined as accent in the beat annotation data are also accent, that is, the annotation platform provides an automatic overlay operation for the accent track, which can batch implement accent annotation.

[0184] Figure 7 The interface diagram of the beat annotation provided by the embodiment of the present application is shown in Figure 7 As shown, the beat annotation data annotated and drawn by the AI algorithm model is annotated and drawn on the pre-annotation track and the beat track, the beat speed and the beat number are displayed at the speed track, the annotation personnel can calculate the beat speed through the BPM measurer, and adjust the beat speed corresponding to the beat number determined by the annotation platform according to the calculated beat speed, wherein *2, +2 indicates the adjustment control of adjusting the beat speed, which can be multiplied by 2 or added by 2 for adjustment.

[0185] For uniform speed segments, the beat annotation data (i.e. Figure 7The beat line displayed in the beat track) is displayed in the beat track, and the BPM and default tempo are calculated according to the beat annotation data. The non-uniform speed segment directly displays the beat annotation data in the beat track.

[0186] The beat annotation data of the beat track can be adjusted by the annotator to make the beat track play the beat sound to assist the annotator in beat annotation. A pop-up window can be opened by double-clicking to adjust the beat annotation in a measure unit. The subsequent part in the music audio paragraph can be checked to batch implement the beat annotation, and the beat annotation data is displayed at the beat track.

[0187] It should be noted that, Figure 7 The beat annotation interface shown can also display the lyrics of the music audio to be annotated and the comment track. If the preset lyrics are opened, the preset lyrics are displayed, and the comment information for the music audio to be annotated is displayed at the comment track. Figure 7 In addition, if the beat annotation task is completed, the save control and the submit control can be clicked in sequence to enter the next annotation task, which can be a composition structure annotation task.

[0188] In an optional implementation, the at least one preset music element includes lyrics. In response to an adjustment operation on the annotation data of the target music element, the annotation data of the target music element can be adjusted, which can include Figure 8 The steps shown.

[0189] Figure 8 The data annotation method provided by the embodiment of the application Figure Four As Figure 8 In response to an adjustment operation on the annotation data of the target music element, the annotation data of the target music element can be adjusted, which can include:

[0190] S401, in response to a first insertion operation input for a separator, a first separator is inserted in the lyrics data.

[0191] S402, in response to an adjustment operation on the first annotation data corresponding to the first separator in the lyrics annotation data, the first annotation data is adjusted.

[0192] The at least one preset music element includes lyrics, and the annotation data of the at least one preset music element includes lyrics annotation data. The annotator can input a first insertion operation for a separator, and the electronic device responds to the first insertion operation for the separator to insert a first separator in the lyrics annotation data.

[0193] Afterwards, the annotator can input an adjustment operation for the first annotation data corresponding to the first delimiter in the lyrics annotation data to adjust the first annotation data, for example, adding, deleting, or modifying part of the lyrics (i.e., the first annotation data) in the lyrics annotation data. The first annotation data corresponding to the first delimiter can be the annotation data before or after the first delimiter in the lyrics annotation data.

[0194] It should be noted that when the annotator adjusts the first annotation data in the lyrics annotation data, the annotation platform also provides real-time annotation error prompts, including but not limited to the time boundary between the lyrics and the song sentences not being aligned, the adjusted lyrics not being consistent with the actual ones, etc.

[0195] Figure 9 This is a schematic diagram of the interface for lyrics annotation provided in the embodiment of the present application, such as Figure 9 As shown, the lyrics pre-annotated by the AI ​​algorithm model are drawn on the lyrics track, and the sentences are composed of lyrics and drawn on the sentence track. The annotator can check and drag the boundary lines of the sentences and the boundary lines of the lyrics to align the boundary lines of the sentences composed of lyrics with the boundary lines of the last lyrics of the sentence. When dragging the boundary lines, the annotation platform can automatically absorb the boundary lines of the lyrics and the sentence, making it convenient for the annotator to drag the boundary lines to align.

[0196] It should be noted that Figure 9 The beat annotation interface shown can also display the lyrics of the music audio to be annotated and the annotation track. If the preset lyrics are turned on, the preset lyrics will be displayed, and the annotation track will display the annotation information for the music audio to be annotated. Figure 9 The annotation track shown is empty. In addition, if the beat marking task is completed, you can click the Save control and the Submit control in sequence to enter the next marking task, which can be a pitch marking task.

[0197] In an optional implementation, at least one preset music element includes: a musical score pitch, step S201, in response to an adjustment operation on the annotation data of a target music element, adjusting the annotation data of the target music element may include: Figure 10 Steps shown.

[0198] Figure 10 Schematic diagram of the data annotation method provided in this application embodiment Figure Five ,like Figure 10 As shown, in response to the adjustment operation on the annotation data of the target music element, adjusting the annotation data of the target music element may include:

[0199] S501: In response to an operation for inputting a separator, insert a second separator into the music score pitch marking data.

[0200] S502: In response to an adjustment operation on the second annotation data corresponding to the second separator in the score pitch annotation data, adjust the second annotation data.

[0201] Among them, at least one preset music element includes: music score pitch, and the annotation data of at least one preset music element includes: music score pitch annotation data. The annotator can input a second insertion operation for the separator, and the electronic device responds to the second insertion operation for the separator and inserts a second separator in the music score pitch annotation data.

[0202] Afterwards, the annotator can input an adjustment operation for the second annotation data corresponding to the second delimiter in the score pitch annotation data to adjust the second annotation data. For example, the annotator can add, delete, or modify part of the score pitch (second annotation data) in the score pitch annotation data, such as inserting a transposition into the second annotation data, to correct the score pitch annotation data pre-annotated by the AI ​​algorithm model. The second annotation data corresponding to the second delimiter can be the annotation data before or after the second delimiter in the score pitch annotation data.

[0203] In an optional embodiment, the method may further include:

[0204] Quantize the data accuracy of the music score pitch marking data to the data accuracy corresponding to the music audio.

[0205] Among them, the adjusted music score pitch annotation data, that is, the pitch free time value, can also perform structured processing on the music score pitch, which may specifically include: automatically quantizing the accuracy of the music score pitch data to the data accuracy corresponding to the music audio to be annotated.

[0206] In one embodiment of the present application, the annotation platform also provides an automatic quantization function. The annotator can customize the quantization accuracy. After turning on automatic quantization, automatic quantization can be performed. During the automatic quantization process, the dividing lines of all annotated cells can be aligned to the nearest note line from the front to the back in the pitch track, lyrics track (i.e., Chinese character track), and phrase track. Among them, the phrase track is used to display phrases. The phrases are sentences obtained by dividing the lyrics according to the musical sense. In other words, the phrases are the phrase annotation data obtained based on the musical sense on the lyrics annotation data.

[0207] Figure 11 A schematic diagram of the pitch marking interface provided in the embodiment of the present application is shown as follows: Figure 11 As shown in the figure, the pitch annotation data pre-annotated by the AI ​​algorithm model are plotted on the pitch track and phrase track. The annotators can check and correct the pitch annotation data. When the pitch segment is adjusted, dynamic feedback of the piano reference tone can be provided to assist the annotators in their work.

[0208] in, Figure 11The category of the phrase in the phrase category field is Other, indicating the "other" category, i.e., the category of the phrase that cannot be classified into a certain set type, Other1 indicates the first phrase called "other", Other2 indicates the second phrase called "other", Other3 indicates the third phrase called "other", and so on.

[0209] In addition, if the pitch labeling task is completed, the save control and the submit control can also be clicked in sequence to enter the pitch structured score labeling task. The pitch structured score labeling task will display the last stage pitch free time value (i.e., the score pitch labeling data), and then automatically quantize the pitch, lyrics and phrase to the accuracy requirement at the time of task publishing, such as 1 / 32 note, through calculation. The user checks and modifies the quantization result.

[0210] Among them, for lyrics, the modification of the quantization result can be to move the dividing line of the lyrics left and right, so as to adjust the end time point of the last lyric in the song to the end time point of the song; for pitch, the modification of the quantization result can be to adjust the pitch by selecting up and down with the mouse; for the phrase, the modification of the quantization result can be to move the dividing line of the phrase left and right to adjust the time range corresponding to the phrase, and the category of the phrase can be adjusted by selecting up and down with the mouse.

[0211] Figure 12 The interface diagram of the composition structure labeling provided by the embodiment of the present application is shown in Figure 12 As shown, the AI algorithm model pre-labeled composition structure labeling data is drawn on the composition structure track, the AI algorithm model pre-labeled paragraph sequence and the tonality of the song are drawn on the paragraph track, and the chord sequence is drawn on the chord track. The labeling personnel can check and modify whether the paragraph division is correct, whether the tonality to which each paragraph belongs is correct, and whether the chord pre-labeling is correct.

[0212] Among them, when each chord segment is selected, there is real-time dynamic feedback of piano chord sound, which assists the labeling personnel in the work.

[0213] In an optional implementation, the at least one preset music element includes a phrase. In response to an adjustment operation on the labeling data of the target music element, the labeling data of the target music element is adjusted, which can include Figure 13 as shown.

[0214] Figure 13 The flowchart of the data labeling method provided by the embodiment of the present application is shown in Figure Six As shown in Figure 10 In response to an adjustment operation on the labeling data of the target music element, the labeling data of the target music element is adjusted, which includes:

[0215] S601, in response to a third insertion operation input for the delimiter, inserting a third delimiter in the verse annotation data.

[0216] S602, inserting the selected preset lyrics into the verse annotation data.

[0217] The at least one preset music element includes a verse, and the annotation data of the at least one preset music element includes verse annotation data. The annotator can input a third insertion operation for the delimiter. The electronic device inserts a third delimiter in the verse annotation data in response to the third insertion operation for the delimiter.

[0218] Then, the selected preset lyrics can be inserted into the verse annotation data. The preset lyrics can be preset lyrics displayed on the verse annotation interface. The selected preset lyrics can be a verse pointed to by an arrow (for example, "in a place far from this" shown in the arrow in the middle). Figure 7

[0219] In this embodiment, the verse annotation data generated by calling the backend AI algorithm pre-annotation is provided with an automatic verse input function. After starting the automatic verse input, the lyrics of the verse pointed to by the arrow in the preset lyrics are automatically input into the corresponding interval (such as the interval after the delimiter) every time a delimiter (i.e., a carriage return interval) is added. After the lyrics are inserted, the arrow points to the next line of lyrics. If there is no next line, the automatic verse input is closed.

[0220] In an optional implementation, in step S102, the at least one audio element annotation model obtained by training is used to process the music audio to obtain at least one preset music element data, which can include Figure 14 as shown in the steps.

[0221] Figure 14 The data annotation method provided in the embodiment of the present application Figure Seven As shown in Figure 14 , the at least one audio element annotation model obtained by training is used to process the music audio to obtain at least one preset music element data, including:

[0222] S701, processing the music audio according to the first audio element annotation model to obtain first music element annotation data.

[0223] S702, processing the music audio and the first music element annotation data according to the second audio element annotation model to obtain second music element annotation data.

[0224] ​The music audio to be labeled is input into the first audio element labeling model to obtain first music element labeling data, and the first music element labeling data and the music audio are input into the second music element labeling model to obtain second music element labeling data, that is, the second music element labeling data depends on the first music element data, so that the accuracy of the second music element labeling data can be improved.

[0225] For example, the first music element labeling data is beat labeling data, and the second music element labeling data is composition structure labeling data, the first music element labeling data is lyrics labeling data, and the second music element labeling data is pitch labeling data.

[0226] Figure 15 The architecture diagram of data labeling provided by the embodiment of the application is shown in Figure 15 The music audio to be labeled is input into the beat labeling model to obtain beat labeling data, and the beat labeling data and the music audio to be labeled are input into the composition structure labeling model to obtain composition structure labeling data.

[0227] The music audio to be labeled is input into the lyrics labeling model to obtain lyrics labeling data, and the lyrics labeling data and the music audio to be labeled are input into the score pitch labeling model to obtain score pitch labeling data (i.e. pitch free time value). In addition, the score pitch labeling data can be quantified to obtain a pitch structured value.

[0228] In the implementation scheme, after each stage of labeling task is completed, it is checked and audited by a tester, and after the check passes, a manager uniformly stores the labeling data in a background system to provide unified data, and formats the music audio according to the labeling data to generate labeling data for AI algorithm model training.

[0229] Figure 16 The interface diagram of the labeling data export provided by the embodiment of the application is shown in Figure 16 After all labeling tasks are completed, the labeling content can be exported to a preset path, Figure 16 The format of the labeling content is.txt.

[0230] Figure 17 The structure diagram of the data labeling device provided by the embodiment of the application is shown in Figure 17 The device includes:

[0231] The acquisition module 801 is configured to acquire music audio to be labeled.

[0232] The processing module 802 is configured to process the music audio according to the at least one audio element labeling model to obtain at least one preset music element data, wherein the at least one audio element labeling model is trained according to a plurality of music audio samples, and the plurality of music audio samples are labeled with sample labeling data corresponding to the preset music element.

[0233] The display module 803 is configured to display the corresponding labeling data at the audio track of the at least one preset music element on the graphical user interface.

[0234] In an optional implementation, the processing module 802 is further configured to:

[0235] In response to an adjustment operation on the labeling data of the target music element, the labeling data of the target music element is adjusted.

[0236] In an optional implementation, the processing module 802 is further configured to:

[0237] In response to a playback operation on the labeling data of the target music element, the audio of the labeling data of the target music element is played.

[0238] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of beat, lyrics, arrangement structure, and score pitch.

[0239] In an optional implementation, the at least one preset music element includes beat.

[0240] The processing module 802 is further configured to determine a beat speed of the music audio according to the beat labeling data.

[0241] The display module 803 is further configured to display the beat speed at a speed track on the graphical user interface.

[0242] In an optional implementation, the display module 803 is further configured to:

[0243] In response to a selection operation on the beat speed measurer displayed on the graphical user interface, a beat touch window is displayed, and a beat sound of the beat labeling data in the music audio is played, so that a touch operation on the beat touch window is input based on the beat sound.

[0244] The processing module 802 is further configured to determine a touch speed corresponding to the touch operation on the beat touch window in response to the touch operation on the beat touch window.

[0245] The touch speed is determined as the adjusted beat speed.

[0246] In an optional implementation, the processing module 802 is further configured to:

[0247] playing the beat sound of the beat annotation data;

[0248] determining part of the annotation data from the beat annotation data in response to a selection operation on the beat annotation data.

[0249] In an optional implementation, the at least one preset music element includes lyrics, and the processing module 802 is specifically configured to:

[0250] inserting a first delimiter in the lyrics annotation data in response to a first insertion operation on the delimiter input;

[0251] adjusting first annotation data corresponding to the first delimiter in the lyrics annotation data in response to an adjustment operation on the first annotation data.

[0252] In an optional implementation, the at least one preset music element includes a musical score pitch, and the processing module 802 is specifically configured to:

[0253] inserting a second delimiter in the musical score pitch annotation data in response to a second insertion operation on the delimiter input;

[0254] adjusting second annotation data corresponding to the second delimiter in the musical score pitch annotation data in response to an adjustment operation on the second annotation data.

[0255] In an optional implementation, the processing module 802 is further configured to:

[0256] quantizing the precision of the musical score pitch annotation data to the data precision corresponding to the music audio.

[0257] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of a phrase, a syllable, a phoneme, a song pitch, and a sound length.

[0258] In an optional implementation, the at least one preset music element includes a phrase, and the processing module 802 is specifically configured to:

[0259] inserting a third delimiter in the phrase annotation data in response to a third insertion operation on the delimiter input;

[0260] inserting the selected preset lyrics into the phrase annotation data.

[0261] In an optional implementation, the processing module 802 is specifically configured to:

[0262] processing the music audio according to a first audio element annotation model to obtain first music element annotation data;

[0263] The music audio and the first music element annotation data are processed according to the second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model comprises the first audio element annotation model and the second audio element annotation model.

[0264] In the data annotation apparatus, the obtaining module is configured to obtain music audio to be annotated, the processing module is configured to process the music audio according to at least one trained audio element annotation model to obtain at least one preset music element data, wherein the at least one audio element annotation model is trained according to a plurality of music audio samples, the plurality of music audio samples are annotated with sample annotation data for corresponding preset music elements, and the display module is configured to display the corresponding annotation data at audio tracks of the at least one preset music element on a graphical user interface. The preset music element data for the music audio is automatically annotated by means of an algorithm, and the annotation efficiency is high and the accuracy is high.

[0265] Figure 18 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. Figure 18 As shown in FIG. 1, the device includes a processor 901, a memory 902, and a bus 903. The memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device is running, the processor 901 communicates with the memory 902 through the bus 903. The processor 901 executes the machine-readable instructions to perform the following steps:

[0266] Obtaining music audio to be annotated;

[0267] Processing the music audio according to at least one trained audio element annotation model to obtain annotation data of at least one preset music element, wherein the at least one audio element annotation model is trained according to a plurality of music audio samples, and the plurality of music audio samples are annotated with sample annotation data for corresponding preset music elements;

[0268] Displaying the corresponding annotation data at audio tracks of the at least one preset music element on a graphical user interface.

[0269] In an optional implementation, after the corresponding annotation data is displayed at the audio tracks of the at least one preset music element on the graphical user interface, the processor 901 is further configured to:

[0270] Adjusting the annotation data of the target music element in response to an adjustment operation on the annotation data of the target music element.

[0271] In an optional implementation, before the annotation data of the target music element is adjusted in response to the adjustment operation on the annotation data of the target music element, the processor 901 is further configured to:

[0272] in response to a play operation on the annotation data of the target music element, playing audio of the annotation data of the target music element.

[0273] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of beat, lyrics, arrangement structure, and score pitch.

[0274] In an optional implementation, the at least one preset music element includes beat.

[0275] After displaying the corresponding annotation data at the audio track of the at least one preset music element on the graphical user interface, the processor 901 is further configured to:

[0276] determining a beat speed of the music audio according to the beat annotation data;

[0277] displaying the beat speed at a speed track on the graphical user interface.

[0278] In an optional implementation, the processor 901 is further configured to:

[0279] in response to a selection operation on the beat speed measurer displayed on the graphical user interface, displaying a beat touch window and playing a beat sound of the beat annotation data of the music audio, so as to input a touch operation on the beat touch window based on the beat sound;

[0280] in response to the touch operation on the beat touch window, determining a touch speed corresponding to the touch operation;

[0281] determining the touch speed as the adjusted beat speed.

[0282] In an optional implementation, the processor 901 is further configured to:

[0283] playing the beat sound of the beat annotation data;

[0284] in response to a selection operation on the beat annotation data, determining part of the annotation data from the beat annotation data.

[0285] In an optional implementation, the at least one preset music element includes lyrics.

[0286] The processor 901 is specifically configured to:

[0287] in response to a first insertion operation on the input separator, inserting a first separator in the lyrics annotation data;

[0288] in response to an adjustment operation on first annotation data corresponding to the first separator in the lyrics annotation data, adjusting the first annotation data.

[0289] In an optional implementation, the at least one preset music element includes a musical pitch.

[0290] The processor 901 is specifically configured to:

[0291] insert a second delimiter in the musical pitch annotation data in response to a second insertion operation input for the delimiter;

[0292] adjust the second annotation data corresponding to the second delimiter in the musical pitch annotation data in response to an adjustment operation on the second annotation data.

[0293] In an optional implementation, the processor is further configured to:

[0294] quantize the precision of the musical pitch annotation data to the data precision corresponding to the music audio.

[0295] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of a verse, a syllable, a phoneme, a song pitch, and a sound length.

[0296] In an optional implementation, the at least one preset music element includes a verse.

[0297] The processor 901 is specifically configured to:

[0298] insert a third delimiter in the verse annotation data in response to a third insertion operation input for the delimiter;

[0299] insert the selected preset lyrics into the verse annotation data.

[0300] In an optional implementation, the processor 901 is specifically configured to:

[0301] process the music audio according to a first audio element annotation model to obtain first music element annotation data;

[0302] process the music audio and the first music element annotation data according to a second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes the first audio element annotation model and the second audio element annotation model.

[0303] In the electronic device of the embodiment, when the electronic device is running, the processor acquires music audio to be labeled, processes the music audio according to at least one audio element labeling model trained, to obtain labeling data of at least one preset music element, wherein the at least one audio element labeling model is trained according to a plurality of music audio samples, the plurality of music audio samples are labeled with sample labeling data corresponding to the preset music element, and the corresponding labeling data is displayed at the audio track of the at least one preset music element on a graphical user interface. The preset music element data corresponding to the music audio is automatically labeled by means of an algorithm, and the labeling efficiency is high and the accuracy is high.

[0304] The embodiment of the application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed when a processor is running.

[0305] Acquire music audio to be labeled.

[0306] Process the music audio according to at least one audio element labeling model trained, to obtain labeling data of at least one preset music element, wherein the at least one audio element labeling model is trained according to a plurality of music audio samples, the plurality of music audio samples are labeled with sample labeling data corresponding to the preset music element.

[0307] Display the corresponding labeling data at the audio track of the at least one preset music element on a graphical user interface.

[0308] In an optional implementation, after the corresponding labeling data is displayed at the audio track of the at least one preset music element on the graphical user interface, the method further includes:

[0309] Adjust the labeling data of the target music element in response to an adjustment operation on the labeling data of the target music element.

[0310] In an optional implementation, before the labeling data of the target music element is adjusted in response to the adjustment operation on the labeling data of the target music element, the method further includes:

[0311] Play the audio of the labeling data of the target music element in response to a play operation on the labeling data of the target music element.

[0312] In an optional implementation, the music audio includes score audio, and the at least one preset music element includes at least one of beat, lyrics, arrangement structure, and score pitch.

[0313] In an optional implementation, the at least one preset music element includes beat.

[0314] After displaying the corresponding annotation data at the track of the at least one preset music element respectively on the graphical user interface, the method further comprises:

[0315] According to the beat annotation data, determining a beat speed of the music audio;

[0316] Displaying the beat speed at a speed track on the graphical user interface.

[0317] In an optional implementation, the method further comprises:

[0318] In response to a selection operation on the beat speed measurer displayed on the graphical user interface, displaying a beat touch window and playing a beat sound of the beat annotation data in the music audio, so as to input a touch operation on the beat touch window based on the beat sound;

[0319] In response to the touch operation on the beat touch window, determining a touch speed corresponding to the touch operation;

[0320] Determining the touch speed as the adjusted beat speed.

[0321] In an optional implementation, the method further comprises:

[0322] Playing the beat sound of the beat annotation data;

[0323] In response to a selection operation on the beat annotation data, determining part of the annotation data from the beat annotation data.

[0324] In an optional implementation, the at least one preset music element comprises: lyrics;

[0325] In response to an adjustment operation on the annotation data of the target music element, adjusting the annotation data of the target music element, comprising:

[0326] In response to a first insertion operation on the input separator, inserting a first separator in the lyrics annotation data;

[0327] In response to an adjustment operation on the first annotation data corresponding to the first separator in the lyrics annotation data, adjusting the first annotation data.

[0328] In an optional implementation, the at least one preset music element comprises: musical score pitch;

[0329] In response to an adjustment operation on the annotation data of the target music element, adjusting the annotation data of the target music element, comprising:

[0330] In response to a second insertion operation on the input separator, inserting a second separator in the musical score pitch annotation data;

[0331] In response to the adjustment operation on the second annotation data corresponding to the second delimiter in the score pitch annotation data, the second annotation data is adjusted.

[0332] In an optional implementation, the method further includes:

[0333] The precision of the score pitch annotation data is quantified to the data precision corresponding to the music audio.

[0334] In an optional implementation, the music audio includes song audio, and the at least one preset music element includes at least one of a verse, a syllable, a phoneme, a song pitch, and a sound length.

[0335] In an optional implementation, the at least one preset music element includes a verse.

[0336] In response to the adjustment operation on the annotation data of the target music element, the annotation data of the target music element is adjusted, including:

[0337] In response to the third insertion operation input for the delimiter, a third delimiter is inserted in the verse annotation data.

[0338] The selected preset lyrics are inserted into the verse annotation data.

[0339] In an optional implementation, the music audio is processed according to the at least one trained audio element annotation model to obtain the annotation data of the at least one preset music element, including:

[0340] The music audio is processed according to the first audio element annotation model to obtain first music element annotation data.

[0341] The music audio and the first music element annotation data are processed according to the second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes the first audio element annotation model and the second audio element annotation model.

[0342] The computer readable storage medium is processed and run, and can execute the following steps: obtaining music audio to be annotated, processing the music audio according to at least one trained audio element annotation model to obtain annotation data of at least one preset music element, wherein the at least one audio element annotation model is trained according to a plurality of music audio samples, the plurality of music audio samples are annotated with sample annotation data corresponding to the preset music element, and the corresponding annotation data is displayed at the audio track of the at least one preset music element on a graphical user interface. The preset music element data of the music audio is automatically annotated by means of an algorithm, the annotation efficiency is high, and the accuracy is high.

[0343] In the embodiments of the present application, the computer program, when executed by the processor, can also execute other machine readable instructions to perform the methods as described in the embodiments, for the specific method steps and principles performed, refer to the description of the embodiments, which will not be described in detail here.

[0344] In the embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented by other means. The device embodiments described above are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, and for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, which can be electrical, mechanical or other forms.

[0345] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0346] In addition, each functional unit in the embodiments provided in the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0347] The functions can be realized in the form of software functional units and sold or used as independent products when the functions are realized in the form of software functional units and sold or used as independent products, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the technical solutions that essentially contribute to the prior art can be embodied in the form of software products, which are stored in a storage medium and include a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0348] It should be noted that like reference numerals and letters refer to like items throughout the accompanying drawings, and once an item is defined in one drawing, it should not be further defined and explained in subsequent drawings, and further, the terms "first", "second", "third" and the like are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0349] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present application, and are used to illustrate the technical solutions of the present application, but are not limiting, and the protection scope of the present application is not limited thereto, although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features, and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data annotation method, characterized in that: include: Get the music audio to be annotated; Processing the music audio according to at least one trained audio element annotation model to obtain annotation data of at least one preset music element, wherein the at least one audio element annotation model is trained based on a plurality of music audio samples, and the plurality of music audio samples are annotated with sample annotation data for the corresponding preset music element; Displaying the corresponding annotation data at the audio track of the at least one preset music element on the graphical user interface; The at least one audio element annotation model obtained through training processes the music audio respectively to obtain at least one preset music element data, including: Processing the music audio according to the first audio element annotation model to obtain first music element annotation data; The music audio and the first music element annotation data are processed according to a second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes: the first audio element annotation model and the second audio element annotation model.

2. The method according to claim 1, characterized in that After the corresponding annotation data is displayed at the audio track of the at least one preset music element on the graphical user interface, the method further includes: In response to an adjustment operation on the annotation data of the target music element, the annotation data of the target music element is adjusted.

3. The method according to claim 2, characterized in that The method further includes, before responding to the adjustment operation on the annotation data of the target music element and adjusting the annotation data of the target music element: In response to a play operation on the annotated data of the target music element, audio of the annotated data of the target music element is played.

4. The method according to claim 2, characterized in that The music audio includes: music score audio, and the at least one preset music element includes: at least one element of: beat, lyrics, arrangement structure, and music score pitch.

5. The method according to claim 1, wherein The at least one preset music element includes: beat; After displaying the corresponding annotation data at the audio track of the at least one preset music element on the graphical user interface, the method further includes: Determining the tempo of the music audio according to the tempo marking data; The tempo is displayed at a tempo track on the graphical user interface.

6. The method according to claim 5, characterized in that The method further comprises: In response to a selection operation on the beat speed meter displayed on the graphical user interface, a beat touch window is displayed, and a beat sound of the beat marking data in the music audio is played so that a touch operation on the beat touch window is input based on the beat sound; In response to a touch operation on the beat touch window, determining a touch speed corresponding to the touch operation; The touch speed is determined to be the adjusted beat speed.

7. The method according to claim 5, characterized in that The method further comprises: Playing the beat sound of the beat marking data; In response to a selection operation on the beat labeling data, partial labeling data is determined from the beat labeling data.

8. The method according to claim 2, characterized in that The at least one preset music element includes: lyrics; The step of adjusting the labeled data of the target music element in response to the adjustment operation on the labeled data of the target music element includes: In response to a first insert operation for a delimiter input, inserting a first delimiter into the lyrics annotation data; In response to an adjustment operation on the first annotation data corresponding to the first separator in the lyrics annotation data, the first annotation data is adjusted.

9. The method according to claim 2, characterized in that The at least one preset music element includes: music score pitch; The step of adjusting the labeled data of the target music element in response to the adjustment operation on the labeled data of the target music element includes: In response to a second insert operation for the separator input, inserting a second separator into the score pitch marking data; In response to an adjustment operation on the second marking data corresponding to the second separator in the score pitch marking data, the second marking data is adjusted.

10. The method according to claim 9, characterized in that The method further comprises: The accuracy of the music score pitch marking data is quantified to the data accuracy corresponding to the music audio.

11. The method according to claim 1, wherein The music audio includes singing audio, and the at least one preset music element includes at least one element of a phrase, a syllable, a phoneme, a singing pitch, and a note length.

12. The method according to claim 2, characterized in that The at least one preset music element includes: a verse; The step of adjusting the labeled data of the target music element in response to the adjustment operation on the labeled data of the target music element includes: In response to a third insert operation for a delimiter input, inserting a third delimiter into the phrase marking data; Insert the selected preset lyrics into the song phrase annotation data.

13. A data labeling device, characterized in that: include: An acquisition module, used to acquire the music audio to be annotated; a processing module, configured to process the music audio according to at least one trained audio element annotation model to obtain at least one preset music element data, wherein the at least one audio element annotation model is trained based on a plurality of music audio samples, and the plurality of music audio samples are annotated with sample annotation data for the corresponding preset music element; A display module, configured to display the corresponding annotation data at the audio track of the at least one preset music element on a graphical user interface; The processing module is specifically used to: Processing the music audio according to the first audio element annotation model to obtain first music element annotation data; The music audio and the first music element annotation data are processed according to a second audio element annotation model to obtain second music element annotation data, wherein the at least one audio element annotation model includes: the first audio element annotation model and the second audio element annotation model.

14. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the data labeling method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data labeling method according to any one of claims 1 to 12 is executed.

Citation Information

Patent Citations

  • Lyric timestamp generation method and storage medium

    CN112786020A