Information processing method, device, electronic device and storage medium
Through the pre-trained phrase division model and neural network, multi-level threshold processing is performed based on the singing breathing points, which solves the problem of automatic melody division and realizes the rapid multi-level phrase recognition and judgment of the melody.
Patent Information
- Application Number
- CN202110448290.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-04-25
AI Technical Summary
It is difficult to quickly realize the automatic multi-level phrase division of melody with the existing technology. Especially when the melody is automatically generated, it is difficult to make a reasonable division based on simple logic.
A pre-trained phrase segmentation model is used, with melody information and singing breathing points as the division moments. Melody segmentation is performed through multi-level thresholds, and a neural network is used for data-driven multi-level phrase segmentation.
It realizes the rapid and automated multi-level phrase division of melodies, can accurately identify primary and secondary phrases, adapt to various melodic situations, and can judge melodies that are difficult to divide.
Smart Images

Figure CN113158642B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital music, and in particular to an information processing method, device, electronic device and storage medium. Background Art
[0002] Since the information revolution, the way music and multimedia are disseminated has undergone tremendous changes in a short period of time. This qualitative change has led to an explosive growth in the market demand for all types of music: whether it is singles, albums, MVs, karaoke with pop music or artistic creation as the main elements, or short videos, advertisements, animations, promotional films and film and television works that use music as an auxiliary, or radio stations, anchors, and public space music that use music as background content, a large amount of original music is needed. In computer automatic composition technology, if the automatically created melody needs to be divided into melodic phrases (i.e., musical phrases) for subsequent applications such as lyrics, the ability to automatically divide musical phrases is very important in applications such as automatic composition and vocal synthesis. How to quickly realize the automatic division of musical phrases has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] The present application provides an information processing method, device, electronic device and storage medium.
[0004] According to one aspect of the present application, there is provided an information processing method, comprising:
[0005] According to melody information and a pre-trained phrase division model, the melody information is segmented based on a multi-level threshold to obtain multi-level phrase information constituting the melody information; wherein the annotation information used to train the phrase division model includes: phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment.
[0006] According to another aspect of the present application, there is provided an information processing device, comprising:
[0007] A phrase segmentation processing module is used to perform melody phrase segmentation processing on the melody information based on a multi-level threshold value according to the melody information and a pre-trained phrase segmentation model, so as to obtain multi-level phrase information constituting the melody information; wherein the annotation information used to train the phrase segmentation model includes: phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment.
[0008] According to another aspect of the present application, an electronic device is provided, including:
[0009] at least one processor; and
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by any embodiment of the present application.
[0012] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method provided by any embodiment of the present application.
[0013] By adopting the present application, the melody information can be segmented based on a multi-level threshold according to the melody information and a pre-trained phrase division model, thereby obtaining multi-level phrase information constituting the melody information; wherein, the annotation information used to train the phrase division model includes: the phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment, thereby quickly realizing the automatic division of phrases.
[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.
[0016] Figure 1 is a flowchart of an information processing method according to an embodiment of the present application;
[0017] Figure 2 is a schematic diagram of the structure of an information processing device according to an embodiment of the present application;
[0018] Figure 3 It is a block diagram of an electronic device used to implement the information processing method of an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.
[0021] In addition, numerous specific details are provided in the detailed description below to better illustrate the present application. Those skilled in the art will appreciate that the present application can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.
[0022] According to an embodiment of the present application, an information processing method is provided. Figure 1 This is a flow chart of an information processing method according to an embodiment of the present application. The method can be applied to an information processing device. For example, the device can be deployed in a terminal or server or other processing device to perform phrase division, etc. The terminal can be a user equipment (UE), a mobile device, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementations, the method can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1 As shown, including:
[0023] S101. Based on melody information and a pre-trained phrase division model, the melody information is segmented based on a multi-level threshold to obtain multi-level phrase information constituting the melody information; wherein the annotation information used to train the phrase division model includes: phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment.
[0024] In one example, if the automatically created melody information needs to be matched with lyrics to realize subsequent applications such as synthesized singing, the melody information needs to be segmented (or divided) into phrases. Phrase information is a basic structural unit with characteristics that constitutes a piece of music. For example, four lines of lyrics correspond to four phrases. Since the annotation information in the sample data set for training the phrase division model is the phrase annotation information obtained by annotating the phrases at the breathing point during the singing of the song, rather than annotating the phrases according to an obvious logical rule, in the obtained multi-level phrase information, if the two-level phrase information (i.e., the first-level phrase information and the second-level phrase information) is taken as an example, the breathing probability of the second-level phrase information will be lower than the breathing probability of the first-level phrase information. As a result, the pre-trained phrase division model obtained after training the phrase annotation information in this application has a significant difference in the probability distribution output by the model between the first-level phrase information and the second-level phrase information, making hierarchical phrase division possible. Finally, the above-mentioned melody phrase segmentation processing (such as multi-level phrase segmentation based on multi-level thresholds) can realize automatic multi-level phrase segmentation. Moreover, for melody information that is difficult to segment into phrases, a judgment can also be made.
[0025] By adopting the present application, the melody information can be segmented based on a multi-level threshold according to the melody information and a pre-trained phrase division model, thereby obtaining multi-level phrase information constituting the melody information; wherein, the annotation information used to train the phrase division model includes: the phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment, thereby quickly realizing the automatic division of phrases.
[0026] In one embodiment, the method further includes: obtaining music score information; extracting music segment construction information including phrase information from the music score information according to a preset beat; obtaining a music segment structure based on the music segment construction information, and collecting data based on the music segment structure to obtain a sample data set for training the phrase segmentation model; wherein the sample data set includes: the phrase annotation information.
[0027] In one embodiment, data collection based on the musical segment structure includes collecting melody representations transposed to predetermined positions within the musical segment structure, where the melody representations are used to describe the melody at different time intervals. For example, a musical segment structure may require collection of melody representations transposed to a predetermined position (e.g., after the key of C major or A minor).
[0028] In one embodiment, the music section structure includes a plurality of melody sequences obtained by dividing the melody information; wherein, when the measure is the first measure, the position of the first measure is determined according to the first chord at the beginning of the music section structure.
[0029] In one embodiment, the method further includes: in the process of training the phrase division model, obtaining multiple melody sequences obtained by dividing the melody information according to the sample data set, and multiple position sequences (Pos) corresponding to the multiple melody sequences (such as M); inputting the multiple melody sequences and the multiple position sequences into the phrase division model to obtain the probabilities of multiple vectors corresponding to the multiple melody sequences, wherein the probabilities are used to represent the probability that each melody sequence is the beginning of the multi-level phrase information; and performing back propagation of the loss function based on the probabilities until convergence to obtain the pre-trained phrase division model.
[0030] In one embodiment, the melody information is segmented based on a multi-level threshold value according to the melody information and a pre-trained phrase segmentation model to obtain the multi-level phrase information constituting the melody information, including: obtaining, based on the melody information and the pre-trained phrase segmentation model, probabilities of multiple vectors corresponding to multiple melody sequences; when the probabilities are greater than a first-level phrase threshold, extracting multiple first sub-melody sequences matching the current situation from the multiple melody sequences, and obtaining multiple first-level phrase information based on the multiple first sub-melody sequences; when the probabilities are greater than a second-level phrase threshold and less than the first-level phrase threshold, extracting multiple second sub-melody sequences matching the current situation from the multiple first sub-melody sequences, and obtaining multiple second-level phrase information based on the multiple second sub-melody sequences; and obtaining the multi-level phrase information based on the multiple first-level phrase information and the multiple second-level phrase information.
[0031] Application examples:
[0032] In computer-generated music composition technology, the ability to automatically segment melodies into phrases is crucial. Currently, no similar automatic melodic phrase segmentation technology exists; only automatic phrase segmentation technologies for speech and text exist. Hypothetically, rule-based automatic phrase segmentation technologies exist, for example, where conditional judgments on whether to segment a melody are made based on information such as the duration of the notes in the rhythm and the current phrase length. However, rule-based automatic phrase segmentation technologies struggle to cover a wide range of melodic scenarios. Especially for automatically generated melodies, phrases are often unclear, making them difficult to segment using simple logic. Automatic segmentation can sometimes fail, or the segmentation may be irrational. Furthermore, multi-level phrase segmentation is difficult.
[0033] In view of the above problems, the processing flow of the application example of the first embodiment of the present application can be used to automatically divide the melody into phrases through the pre-trained phrase division model, and a melody M=m0…m represented by N numerical values can be divided into two phrases. N-1 , to automatically divide the phrase, that is, we need to find a set of strictly increasing position subscripts P = p0…p K, indicating that the melody is divided into K first-level phrases, where p i is the position subscript of the beginning of the melody of the i∈{0,...,K-1}th sentence. In particular, for the convenience of representation, it is agreed that p K = N. The melody of the i∈{0,...,K-1}th sentence is And Q=q0…q K-1 ,p i i <p i+1 -1 means that the i-th phrase can be divided into two secondary phrases and For situations where a phrase cannot be classified as a primary or secondary phrase, an automatic judgment is made. Specifically, the following contents are included:
[0034] 1. Data Collection, Representation, Preprocessing and Model
[0035] Music scores were collected and recorded in segments to form a data set. A segment needs to record the melody after being transposed to C major or A minor. The melody needs to be divided into phrases based on the breathing points when singing the song. The first measure position of the melody and the total number of measures b are determined by the first chord at the beginning of the segment. The first beat of the first measure is defined as the zero moment of the segment. The difference between the position of the first note of the melody and the first measure is recorded as Δst. If the melody of the segment begins before the first measure, the melody of the advance part is called a weak melody, Δst < 0; if it starts on or after the first beat of the first measure, there is no weak melody, Δst ≥ 0. The longest allowed weak melody length in the dataset is ST max is 32, Δst ≥ -ST max .
[0036] Determine the beat quantization length SPQ to be 4. For example, for a 4 / 4 time song, each measure of the melody is represented by 16 values. The key point is the "equally spaced discretization representation" of the song melody, that is, no matter how complex the song melody is, each beat is evenly divided into 4 moments on the time scale. At the 4 moments, the following m can be used i The calculation formula records the melody at this moment. There are 4 moments per beat and 4 beats per measure, which is 16 moments, corresponding to the m of 16 melodies. i In addition, this application does not limit the time signature of the music section, whether to use transposition notation and the target key signature of the transposition, the SPQ value, and does not necessarily have to be expressed in equal parts, as shown below: i The calculation formula can also be changed, as long as a simple "quantization" method can be used to represent the melody, it is within the scope of protection of this application.
[0037] Using SPQ quantization and considering the weak start, the melody of a segment needs to be represented by N values, N = 16 × b + Δst, and only the segments that meet -32≤Δst<16, n≥4 are selected. The melody of a segment is represented by M = m0…m N-1 , where m i ,i∈{0,1,...,N-1} describes the moment If there is a note or rest whose starting time is not equal to a certain t i , then adjust its start time to be equal to a nearby t i The entire melody can be shifted octaves so that the melody satisfies the following formula:
[0038]
[0039] There are N types of melody values. m =62 categories.
[0040] A position sequence Pos=pos0...pos corresponding to the melody sequence can be obtained N-1 ,in:
[0041] POS i =i+Δst+ST max
[0042] The melody segmentation mark can be expressed as a vector s=s0…s corresponding to the melody sequence. N-1 .in,
[0043]
[0044] 2. Training Model
[0045] The model can be an attention network Transformer, which can calculate the probability sequence Y==y0…y of S from the sequence M and Pos N-1 ,y i For melody m i The probability of the beginning of a phrase
[0046] Y=Transformer(M,Pos)
[0047] Train the model, adjust the model parameters θ, and optimize the following loss function:
[0048]
[0049] in represents a data set, D represents the data of a music segment, and M represents the melody in D.
[0050] 3. Using the model for sentence segmentation
[0051] Calculate the Pos sequence for a given melody sequence M and Δst.
[0052] Using the trained model Transformer, the Y sequence is calculated by the formula Y = Transformer (M, Pos).
[0053] A first-level phrase threshold A=0.9 is set.
[0054] In the Y sequence, find all items whose probability values are greater than A. The total number of items is K, which means that K first-level phrases can be divided. The position subscripts of the items are obtained in increasing order p0…p K-1 , which satisfies:
[0055] y pi >A, i∈{0,1,...,K-1}
[0056] Combined with p K =N, and the required sequence P=p0…p k , where the first phrase is a first-level phrase
[0057] In the i-th sentence, find the position subscript with the highest probability in the Y sequence except the first one, that is, the beginning of the first-level phrase. The subscript is the required q i :
[0058]
[0059] A secondary phrase threshold B is set to 0.3.
[0060] like It can be considered that there is no secondary phrase in the i-th sentence, or the secondary phrase is not obvious.
[0061] like Then q i Subscript the position of the desired secondary phrase.
[0062] The first level phrase i∈{0,1,...,K-1} is divided into the second level phrases as described above, and Q=q0…q K-1 So far, the first-level phrase division P and the second-level phrase division Q of the melody sequence are obtained.
[0063] Set an acceptable maximum number of first-level phrases N based on the data set and application scenario. A_max , the minimum number of phrases N A_min .like It can be considered that the number of melody phrases does not meet the requirements.
[0064] Set a suspected phrase threshold C = 0.6 and a suspected ratio RC =0.5. If the proportion of Y sequence in the interval [A, C] is greater than R C , it can be considered that the melody phrase division is not obvious and it is difficult to divide the phrases.
[0065] This application uses a neural network to perform phrase segmentation, achieving data-driven hierarchical phrase segmentation. This is because the data is labeled based on breathing points during performance, rather than a clear logical rule. The probability of breathing in a secondary phrase is lower than that of a primary phrase. This allows the trained neural network to give significantly different probabilities for primary and secondary phrases, making hierarchical phrase segmentation possible. This enables rule-free automatic phrase segmentation, and allows for multi-level phrase segmentation based on set thresholds. Judgments can also be made for melodies that are difficult to segment into phrases.
[0066] According to an embodiment of the present application, an information processing device is provided. Figure 2 is a schematic diagram of the structure of an information processing device according to an embodiment of the present application. Figure 2 As shown, it includes: a sentence segmentation processing module 51, which is used to perform melody sentence segmentation processing on the melody information based on a multi-level threshold according to the melody information and a pre-trained phrase segmentation model, so as to obtain multi-level phrase information constituting the melody information; wherein, the annotation information used to train the phrase segmentation model includes: phrase annotation information obtained by taking the singing breathing point when singing a song based on the melody information as the division moment.
[0067] In one embodiment, the system further includes: a music score acquisition module for acquiring music score information; a music segment construction extraction module for extracting music segment construction information including phrase information from the music score information according to a preset beat; a sample set collection module for obtaining a music segment structure based on the music segment construction information, collecting data based on the music segment structure as a unit, and obtaining a sample data set for training the phrase segmentation model; wherein the sample data set includes the phrase annotation information.
[0068] In one embodiment, the sample set collection module is used to collect melody representations that are transposed to predetermined positions in the music section structure, where the melody representations are used to describe melody conditions at different divided moments.
[0069] In one embodiment, the music section structure includes multiple melody sequences obtained by dividing the melody information; and further includes: a judgment module for: when the measure is the first measure, the position of the first measure is judged according to the first chord at the beginning of the music section structure.
[0070] In one embodiment, it further includes: a first processing module, which is used to obtain, according to the sample data set, multiple melody sequences obtained by dividing the melody information and multiple position sequences corresponding to the multiple melody sequences in the process of training the phrase division model; a second processing module, which is used to input the multiple melody sequences and the multiple position sequences into the phrase division model to obtain the probabilities of multiple vectors corresponding to the multiple melody sequences, and the probabilities are used to represent the probability that each melody sequence is the beginning of the multi-level phrase information; a third processing module, which is used to perform backpropagation of the loss function based on the probabilities until convergence to obtain the pre-trained phrase division model.
[0071] In one embodiment, the sentence segmentation processing module is configured to: obtain, based on melody information and a pre-trained phrase segmentation model, the probabilities of multiple vectors corresponding to multiple melody sequences; extract, when the probabilities are greater than a first-level phrase threshold, multiple first sub-melody sequences matching the current situation from the multiple melody sequences, and obtain multiple first-level phrase information based on the multiple first sub-melody sequences; extract, when the probabilities are greater than a second-level phrase threshold and less than the first-level phrase threshold, multiple second sub-melody sequences matching the current situation from the multiple first sub-melody sequences, and obtain multiple second-level phrase information based on the multiple second sub-melody sequences; and obtain the multi-level phrase information based on the multiple first-level phrase information and the multiple second-level phrase information.
[0072] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method and will not be repeated here.
[0073] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0074] like Figure 3 , is a block diagram of an electronic device for implementing the information processing method of an embodiment of the present application. The electronic device may be the aforementioned deployment device or proxy device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0075] like Figure 3As shown, the electronic device includes: one or more processors 801, a memory 802, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 3 A processor 801 is taken as an example.
[0076] Memory 802 is the non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the information processing method provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions for causing a computer to perform the information processing method provided in this application.
[0077] Memory 802, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs, and modules, such as the program instructions / modules corresponding to the information processing methods in the embodiments of the present application. Processor 801 executes the non-transitory software programs, instructions, and modules stored in memory 802 to execute various functional applications and data processing of the server, thereby implementing the information processing methods in the above-mentioned method embodiments.
[0078] The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 802 may optionally include a memory remotely located relative to the processor 801, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0079] The electronic device of the information processing method may further include: an input device 803 and an output device 804. The processor 801, the memory 802, the input device 803 and the output device 804 may be connected via a bus or other means. Figure 3 The bus connection is taken as an example.
[0080] The input device 803 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device, such as input devices such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, and a joystick. The output device 804 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0081] Various implementations of the systems and techniques described herein can be realized in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0082] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0084] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0085] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0086] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.
[0087] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. An information processing method, characterized in that: The method comprises: Based on melody information and a pre-trained phrase segmentation model, the melody information is segmented based on a multi-level threshold to obtain multi-level phrase information constituting the melody information; wherein the annotation information used to train the phrase segmentation model includes: phrase annotation information obtained by using breathing points during singing of a song based on the melody information as segmentation moments; The method of performing melody phrase segmentation processing on the melody information based on a multi-level threshold value according to the melody information and a pre-trained phrase segmentation model to obtain multi-level phrase information constituting the melody information includes: According to the melody information and the pre-trained phrase segmentation model, the probabilities of multiple vectors corresponding to the multiple melody sequences are obtained; When the probability is greater than a first-level phrase threshold, extracting a plurality of first sub-melody sequences matching the current situation from the plurality of melody sequences, and obtaining a plurality of first-level phrase information based on the plurality of first sub-melody sequences; When the probability is greater than the secondary phrase threshold and less than the primary phrase threshold, extracting a plurality of second sub-melody sequences matching the current situation from the plurality of first sub-melody sequences, and obtaining a plurality of second-level phrase information based on the plurality of second sub-melody sequences; The multi-level phrase information is obtained according to the plurality of first-level phrase information and the plurality of second-level phrase information.
2. The method according to claim 1, characterized in that Also includes: Get music score information; extracting music segment construction information including phrase information from the music score information according to a preset beat; A music segment structure is obtained according to the music segment construction information, and data is collected based on the music segment structure to obtain a sample data set for training the music segmentation model; wherein the sample data set includes: the music segment annotation information.
3. The method according to claim 2, characterized in that The data collection based on the music segment structure includes: Melody representations that are transposed to predetermined positions in the music section structure are collected, and the melody representations are used to describe the melody conditions at different divided moments.
4. The method according to claim 2, characterized in that The music section structure includes a plurality of melody sequences obtained by dividing the melody information; wherein, In the case where the measure is the first measure, the position of the first measure is determined according to the first chord at the beginning of the music section structure.
5. The method according to claim 2, characterized in that Also includes: In the process of training the phrase division model, a plurality of melody sequences obtained by dividing the melody information and a plurality of position sequences respectively corresponding to the plurality of melody sequences are obtained according to the sample data set; Inputting the multiple melody sequences and the multiple position sequences into the phrase division model to obtain probabilities of multiple vectors corresponding to the multiple melody sequences, the probabilities being used to represent the probability that each melody sequence is the beginning of the multi-level phrase information; Back propagation of the loss function is performed based on the probability until convergence, thereby obtaining the pre-trained phrase segmentation model.
6. An information processing device, characterized in that The device comprises: a phrase segmentation processing module for performing melody phrase segmentation processing on the melody information based on a multi-level threshold value according to the melody information and a pre-trained phrase segmentation model, thereby obtaining multi-level phrase information constituting the melody information; wherein the annotation information used to train the phrase segmentation model includes phrase annotation information obtained by using breathing points during singing of a song based on the melody information as segmentation moments; Wherein, the sentence segmentation processing module is used to: According to the melody information and the pre-trained phrase segmentation model, the probabilities of multiple vectors corresponding to the multiple melody sequences are obtained; When the probability is greater than a first-level phrase threshold, extracting a plurality of first sub-melody sequences matching the current situation from the plurality of melody sequences, and obtaining a plurality of first-level phrase information based on the plurality of first sub-melody sequences; When the probability is greater than the secondary phrase threshold and less than the primary phrase threshold, extracting a plurality of second sub-melody sequences matching the current situation from the plurality of first sub-melody sequences, and obtaining a plurality of second-level phrase information based on the plurality of second sub-melody sequences; The multi-level phrase information is obtained according to the plurality of first-level phrase information and the plurality of second-level phrase information.
7. The device according to claim 6, characterized in that Also includes: Music score acquisition module, used to obtain music score information; a music segment construction extraction module, configured to extract music segment construction information including phrase information from the music score information according to a preset beat; The sample set collection module is used to obtain a music segment structure according to the music segment construction information, collect data based on the music segment structure, and obtain a sample data set for training the music segmentation model; wherein the sample data set includes: the music segment annotation information.
8. The device according to claim 7, characterized in that The sample set collection module is used to: Melody representations that are transposed to predetermined positions in the music section structure are collected, and the melody representations are used to describe the melody conditions at different divided moments.
9. The device according to claim 7, characterized in that The music section structure includes a plurality of melody sequences obtained by dividing the melody information; It also includes: a judgment module, which is used to: when the bar is the first bar, the position of the first bar is judged according to the first chord at the beginning of the music section structure.
10. The device according to claim 7, characterized in that Also includes: a first processing module configured to obtain, during training of the phrase division model, a plurality of melody sequences obtained by dividing the melody information according to the sample data set, and a plurality of position sequences corresponding to the plurality of melody sequences; a second processing module, configured to input the plurality of melody sequences and the plurality of position sequences into the phrase division model, and obtain probabilities of a plurality of vectors corresponding to the plurality of melody sequences, the probabilities being used to represent a probability that each melody sequence is the beginning of the multi-level phrase information; The third processing module is used to perform back propagation of the loss function based on the probability until convergence, so as to obtain the pre-trained phrase segmentation model.
11. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system device for retrieving songs based on voice modes
CN102053998A
Music rhythm sectionalized automatic marking method based on eigen-note
CN1737798A