Data processing method, computer device and computer readable storage medium

By analyzing the music score and combining the style sequence for fundamental frequency prediction processing, the style matching problem of the fundamental frequency sequence generation in singing synthesis is solved, and the style feature regulation and authenticity of the fundamental frequency sequence are improved.

CN116110368BActive Publication Date: 2025-05-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310164286.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-05-09
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

In singing vocal synthesis technology, how to effectively generate a basic frequency sequence to match the style characteristics of real singing vocals, especially the jitter frequency and jitter amplitude.

Method used

By obtaining the target music score and its corresponding style sequence, the music score is analyzed to obtain the pitch sequence and phoneme sequence, and the fundamental frequency prediction process is performed in combination with the style sequence to generate a fundamental frequency sequence that meets the real situation.

Benefits of technology

The style characteristics control of the basic frequency sequence is realized, making it more similar to the basic frequency sequence corresponding to the singing voices sung by a live person, including reasonable jitter, and improving the authenticity of singing synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110368B_ABST
    Figure CN116110368B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a data processing method, a computer device and a computer-readable storage medium, wherein the method includes: obtaining a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include jitter frequency and / or jitter amplitude; performing analytical processing on the target music score to obtain a pitch sequence and a phoneme sequence of the target music score; and performing fundamental frequency prediction processing according to the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singing melody of the singer. Through the embodiment of the present application, the fundamental frequency sequence of the music score can be directly generated according to the music score, and the addition of the style sequence makes the fundamental frequency sequence more in line with the actual situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, a computer device, and a computer-readable storage medium. Background Art

[0002] With the popularization of music applications, singing synthesis technology has also attracted more and more attention. Singing synthesis technology is a technology that synthesizes singing audio close to real people's singing based on music score information.

[0003] In singing synthesis, the fundamental frequency curve and energy curve usually determine the pitch and singing level of the synthesized singing voice, which is the key information of singing synthesis. Among them, the energy curve contains volume information, which determines the strength of the singing voice; the fundamental frequency curve contains pitch information, which is the key factor in determining whether the singing voice is out of tune. The fundamental frequency curve is usually generated based on the fundamental frequency sequence, but how to generate the fundamental frequency sequence is a problem to be solved. Summary of the invention

[0004] The embodiments of the present application provide a data processing method, a computer device, and a computer-readable storage medium, which can directly generate a fundamental frequency sequence of a music score based on the music score, and the addition of a style sequence makes the fundamental frequency sequence more consistent with the actual situation.

[0005] On the one hand, an embodiment of the present application provides a data processing method, the method comprising:

[0006] Acquire a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include jitter frequency and / or jitter amplitude;

[0007] Analyzing the target music score to obtain a pitch sequence and a phoneme sequence of the target music score;

[0008] A fundamental frequency prediction process is performed according to the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singing melody of the singer.

[0009] On the one hand, an embodiment of the present application provides a data processing device, the device comprising:

[0010] An acquisition unit, used for acquiring a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include a jitter frequency and / or a jitter amplitude;

[0011] A processing unit, used for analyzing the target music score to obtain a pitch sequence and a phoneme sequence of the target music score;

[0012] The processing unit is further used to perform fundamental frequency prediction processing based on the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singer's singing melody.

[0013] On the one hand, an embodiment of the present application provides a computer device, including: a processor and a memory, the memory storing executable program code, and the processor being used to call the executable program code to implement the data processing method provided in the embodiment of the present application.

[0014] Accordingly, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is executed on a computer, the computer implements the data processing method provided by the embodiment of the present application.

[0015] Accordingly, the embodiment of the present application further provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the data processing method provided in the embodiment of the present application.

[0016] In the present application, firstly, a target music score and a target style sequence corresponding to the target music score are obtained, the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include jitter frequency and / or jitter amplitude; the target music score is analyzed to obtain the pitch sequence and phoneme sequence of the target music score, and the fundamental frequency prediction process is performed according to the pitch sequence, phoneme sequence and target style sequence to obtain the target fundamental frequency sequence used to indicate the singer's singing melody. The data processing method provided in the present application can directly predict the corresponding fundamental frequency sequence according to the music score, and because the style sequence including the jitter frequency and / or jitter amplitude is added in the prediction process, the predicted fundamental frequency sequence contains reasonable jitter, which makes the fundamental frequency sequence more in line with the actual situation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a schematic diagram of a system architecture applicable to the data processing method provided in the embodiment of the present application;

[0019] Figure 2 It is a flowchart of a data processing method provided in an embodiment of the present application;

[0020] Figure 3 It is a schematic diagram of a music score provided in an embodiment of the present application;

[0021] Figure 4 It is a flowchart of a model training method provided in an embodiment of the present application;

[0022] Figure 5 It is a flow chart of a method for obtaining style parameters provided in an embodiment of the present application;

[0023] Figure 6 is a schematic diagram of a method for determining style parameters provided in an embodiment of the present application;

[0024] Figure 7 It is a structural schematic diagram of a fundamental frequency prediction model provided in an embodiment of the present application;

[0025] Figure 8 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;

[0026] Fig. 9 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0028] It should be noted that the descriptions of "first", "second", etc. involved in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined as "first" or "second" may explicitly or implicitly include at least one of the features.

[0029] Singing synthesis technology refers to the technology of synthesizing singing voice close to real people's singing by computer programs. In the process of synthesizing singing voice, it is usually necessary to input music score, fundamental frequency curve and other parameters. Music score is a regular combination of various written symbols that record information such as lyrics and pitch. The fundamental frequency curve is the sine wave with the lowest frequency among the multiple sine waves of different frequencies that make up the singing voice. The fundamental frequency curve determines the melody and is also the key factor in judging whether the singing is out of tune. The fundamental frequency curve is obtained based on the fundamental frequency sequence.

[0030] Based on this, the embodiment of the present application provides a data processing method, which can directly predict a detailed fundamental frequency sequence based on the score and style sequence, and the fundamental frequency sequence is more similar to the real situation. The data processing method provided in the embodiment of the present application can be applied to the field of artificial intelligence. Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0031] The data processing method provided in the embodiment of the present application can Figure 1 The data processing device 101 is implemented. In the process of the data processing device 101 performing fundamental frequency sequence prediction processing, the music score and the style sequence corresponding to the music score are first obtained, and this style sequence is used to indicate the style characteristics of the predicted fundamental frequency. The style sequence is determined according to the style parameters, and the style parameters include the jitter frequency and / or the jitter amplitude. Then, the music score is analyzed to obtain the pitch sequence and phoneme sequence of the music score. Finally, the fundamental frequency prediction processing is performed according to the pitch sequence, the phoneme sequence and the style sequence to obtain the fundamental frequency sequence corresponding to the music score. The fundamental frequency sequence is used to indicate the singing melody of the singer. The fundamental frequency sequence predicted according to the music score contains style characteristics, so the fundamental frequency sequence can reflect more singing details, which makes the fundamental frequency sequence more realistic and more similar to the fundamental frequency sequence of the singing voice of a real person.

[0032] In one embodiment, the data processing device 101 may be a terminal device. The terminal device may obtain the data to be processed input by the user through human-computer interaction with the user, and the data to be processed includes the music score and the style sequence corresponding to the music score, and then perform fundamental frequency prediction processing according to the data to be processed, so as to obtain the fundamental frequency sequence that meets the user's requirements. The terminal device may be a smart home appliance, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a vehicle terminal, an intelligent voice interaction device, a wearable device or other smart devices, etc., but is not limited thereto.

[0033] In one embodiment, the data processing device 101 may be a server. The terminal device performs human-computer interaction with the user, obtains the user's demand information, and processes the demand information to obtain the music score and the style sequence corresponding to the music score. The terminal device sends the music score and the style sequence to the server, and the server performs fundamental frequency prediction processing based on the received music score and style sequence to obtain the fundamental frequency sequence. The fundamental frequency sequence can be used to indicate the singer's singing melody. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0034] It is understandable that the system architecture diagram applicable to the data processing method described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Figure 1 The number of data processing devices 101 in the embodiment is only for illustration. Moreover, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0035] See also Figure 2 , which is a flow chart of a data processing method provided by an exemplary embodiment of the present application. The present application embodiment applies the method to Figure 1 The data processing device described above is used as an example to illustrate that the method can be applied to Figure 1 The data processing device shown in the figure, the data processing method may include but is not limited to the following steps:

[0036] S201. Obtain a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include jitter frequency and / or jitter amplitude.

[0037] In the embodiment of the present application, the target music score is the music score of the song to be synthesized. The target music score contains pitch information and phoneme information. The target style sequence is used to indicate the style characteristics of the predicted fundamental frequency. The target style sequence is determined according to the style parameters, and the style parameters include the jitter frequency and / or the jitter amplitude. The target style sequence can cause controllable jitter to appear in the fundamental frequency sequence, so that the fundamental frequency sequence is more in line with the actual situation.

[0038] In one embodiment, the style parameter may be a preset fixed value, and the style parameter may include a jitter frequency and / or a jitter amplitude. For example, the jitter frequency has a value range of [0.1, 0.2], and the jitter amplitude has a value range of [0.3, 0.5]. If the style parameter includes the jitter frequency and the jitter amplitude, the values ​​of the jitter frequency and the jitter amplitude are obtained, and the target style sequence is determined according to the style parameter including the jitter frequency and the jitter amplitude.

[0039] In one embodiment, the style parameters can be adjusted so that the predicted fundamental frequency sequence meets the user's expected requirements. Specifically, if the style features contained in the predicted target fundamental frequency sequence do not meet the user's expected requirements, and the style parameters include the jitter frequency and the jitter amplitude, the values ​​of the jitter frequency and the jitter amplitude are adjusted to obtain an adjusted target style sequence, and then the fundamental frequency prediction process is performed based on the adjusted style sequence and the music score to obtain a target fundamental frequency sequence that meets the expected requirements.

[0040] S202: Analyze the target music score to obtain a pitch sequence and a phoneme sequence of the target music score.

[0041] In the embodiment of the present application, the target music score includes the pitch information, phoneme information and song time information required for synthesizing the fundamental frequency. The pitch information includes the relevant information of the pitch change of the song, the phoneme information includes the relevant information of the song pronunciation, and the song time information includes the song set speed and the song time signature.

[0042] In one embodiment, the process of parsing the target music score to obtain the pitch sequence and the phoneme sequence can be: first, an extraction operation is performed on the target music score to obtain the pitch information, the phoneme information and the song time information, and then the pitch information and the phoneme information are framed according to the song time information to obtain the pitch sequence and the phoneme sequence of the target music score.

[0043] See also Figure 3 , Figure 3 A schematic diagram of a music score provided in an embodiment of the present application. The music score includes lyrics, song melody, and song time information, wherein the lyrics are "Beautiful Spring, Happy", and the song time information includes the song time signature, set speed, etc. First, the lyrics are extracted to obtain the phoneme information of the music score; then the song melody is extracted to obtain the pitch information of the music score; finally, the phoneme information and pitch information are framed according to the song time information to obtain the frame-level phoneme sequence and pitch sequence of the music score.

[0044] S203, performing fundamental frequency prediction processing according to the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singing melody of the singer.

[0045] In the embodiment of the present application, the target fundamental frequency prediction model can be called to perform fundamental frequency prediction processing on the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score, and the target fundamental frequency sequence is used to indicate the singing melody of the singer. The target fundamental frequency sequence can be used to indicate the singing melody, so that the singing melody can match the music score and achieve the effect of not being out of tune in the singing melody.

[0046] Based on the above embodiments, the beneficial effect of the present application is that the method provided in the embodiments of the present application can not only directly perform fundamental frequency prediction processing according to the target music score to obtain a target fundamental frequency sequence, but also adds a target style sequence determined according to style parameters in the process of fundamental frequency prediction processing. Because the style parameters include jitter frequency and / or jitter amplitude, the target fundamental frequency sequence determined according to the target style sequence has reasonable jitter, that is, the method provided in the present application realizes the regulation of the style characteristics of the target fundamental frequency sequence, so that the target fundamental frequency sequence is relatively similar to the fundamental frequency sequence corresponding to the singing of a real person.

[0047] See also Figure 4 , which is a flow chart of a model training method provided in an embodiment of the present application. The model training method can obtain a target fundamental frequency prediction model according to the training of the initial fundamental frequency prediction model. The model training method can be implemented by a model training device, which can be a data processing device in the above embodiment, or other computing devices different from the data processing device in the above embodiment. The model training method can include but is not limited to the following steps:

[0048] S401, acquiring the sample music score, and performing parsing processing on the sample music score to obtain a sample phoneme sequence and a reference pitch sequence of the sample music score.

[0049] In the embodiment of the present application, the sample music score includes sample phoneme information, sample pitch information and sample time information. The sample music score is extracted to obtain the sample phoneme information and sample pitch information. The sample phoneme information and sample pitch information are framed according to the sample time information to obtain a frame-level sample phoneme sequence and a reference pitch sequence.

[0050] S402: Determine the sample pitch sequence according to the reference pitch sequence.

[0051] In the embodiment of the present application, the reference pitch sequence can be offset to obtain a sample pitch sequence. By performing an offset operation on a reference pitch sequence, multiple sample pitch sequences can be obtained, thereby expanding the number of training samples of the fundamental frequency prediction model, expanding the sound range covered by the training samples of the fundamental frequency prediction model, and improving the generalization ability of the fundamental frequency prediction model.

[0052] In one embodiment, the reference pitch sequence may be directly determined as a sample pitch sequence without performing an offset operation on the reference pitch sequence. The sample pitch sequence matches the sample pitch information contained in the sample music score, so that the fundamental frequency sequence predicted by the model is more consistent with the music score, and the ability of the fundamental frequency prediction model to perform nonlinear modeling on the fundamental frequency sequence is enhanced.

[0053] S403: Acquire the sample singing voice data, perform fundamental frequency extraction processing on the sample singing voice data, and obtain a reference fundamental frequency sequence.

[0054] In the embodiment of the present application, the sample singing data is the singing audio matched with the sample music score. The process of obtaining the reference fundamental frequency sequence according to the sample singing data can be: first obtain the sample singing data, then perform frame processing on the sample singing data, obtain multiple frames of singing audio, then perform fundamental frequency extraction processing on each frame of singing audio in the multiple frames of singing audio, obtain the fundamental frequency value corresponding to each frame of singing audio, and finally arrange the fundamental frequency value corresponding to each frame of singing audio according to the sequence of the playback time corresponding to each frame of singing audio, and obtain the reference fundamental frequency sequence. In the model training process, the reference fundamental frequency sequence can be used to combine with the sample pitch sequence, and determine the style characteristics of the sample fundamental frequency (the style characteristics of the sample fundamental frequency include the jitter frequency and / or jitter amplitude of the sample fundamental frequency).

[0055] In one embodiment, the sample singing voice data may be a pure human voice singing voice audio. A plurality of human voice singing voice audios matching the sample music score are obtained, and the plurality of singing voice audios are combined to obtain the sample singing voice data. The sample singing voice data helps the fundamental frequency prediction model to model the style characteristics of the fundamental frequency sequence.

[0056] In one embodiment, the fundamental frequency extraction process is performed on each frame of singing audio, and the fundamental frequency extraction algorithm may be used for processing. The fundamental frequency extraction algorithm may be a probabilistic YIN algorithm. The probabilistic YIN (pYIN) algorithm is an improved algorithm of the YIN algorithm. The pYIN algorithm calculates each frame of singing audio using a preset function, obtains multiple target values ​​as candidates, and uses a preset model to construct the transfer law of the fundamental frequency, and finally obtains the fundamental frequency corresponding to the singing audio. Other fundamental frequency extraction algorithms, such as the DIO algorithm, the Harvest algorithm, and the YIN algorithm are also applicable to the data processing method proposed in this application.

[0057] S404: Determine the sample fundamental frequency sequence according to the reference fundamental frequency sequence.

[0058] In the embodiment of the present application, the fundamental frequency values ​​in the reference fundamental frequency sequence can be offset to obtain a sample fundamental frequency sequence. By offsetting a reference fundamental frequency sequence, multiple sample fundamental frequency sequences can be obtained, thereby expanding the training sample set and enhancing the generalization ability of the fundamental frequency prediction model.

[0059] In one embodiment, the reference fundamental frequency sequence may not be subjected to offset processing, and the reference fundamental frequency sequence may be directly determined as the sample fundamental frequency sequence. The fundamental frequency prediction model performs fundamental frequency prediction based on the sample style sequence, sample pitch sequence, and sample phoneme sequence contained in the sample fundamental frequency sequence, and the style characteristics of the predicted fundamental frequency sequence are more similar to the style characteristics of the sample singing data.

[0060] S405: Determine the sample style sequence according to the sample pitch sequence and the sample fundamental frequency sequence, wherein the sample style sequence is generated according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

[0061] In the embodiment of the present application, the musical note represents the relative duration of different lengths of sound. Figure 5 The implementation method of determining the sample style sequence according to the sample pitch sequence and the sample fundamental frequency sequence may include but is not limited to the following steps:

[0062] S501 , segmenting the sample pitch sequence to obtain a plurality of pitch subsequences, each pitch subsequence corresponding to a note.

[0063] In the embodiment of the present application, the sample pitch sequence is segmented according to the notes in the sample music score, and the sample pitch sequence corresponding to each note is divided into a pitch subsequence, thereby obtaining multiple pitch subsequences. The relative duration represented by each note is different, and the length of the corresponding pitch subsequence may also be different. The pitch subsequence is used to obtain the style parameters in each note.

[0064] S502: Segment the sample fundamental frequency sequence to obtain a plurality of fundamental frequency subsequences, each of which corresponds to a note.

[0065] In the embodiment of the present application, the sample fundamental frequency sequence is segmented according to the sample singing data, and the sample fundamental frequency sequence corresponding to each note is divided into a fundamental frequency subsequence, thereby obtaining multiple fundamental frequency subsequences. Because the sample fundamental frequency sequence is determined according to the sample singing data, the sample pitch sequence is determined according to the sample music score, and the sample song data matches the sample music score, the pitch subsequence and fundamental frequency subsequence corresponding to the note can be determined according to a note. For example: the sample pitch sequence is divided according to the notes in the sample music score, and the sample pitch sequence can be divided into 100 pitch subsequences; the sample fundamental frequency sequence is divided according to the notes in the sample singing data, and the sample fundamental frequency sequence can also be divided into 100 fundamental frequency subsequences; since the notes in the sample music score and the sample singing data are the same (the sample music score and the sample singing data correspond to the same song), therefore, the above 100 pitch subsequences and 100 fundamental frequency subsequences are one-to-one corresponding. For example, for the first note in the sample music score (or sample singing data), a pitch subsequence corresponding to the first note can be determined, and a fundamental frequency subsequence corresponding to the first note can also be determined.

[0066] S503: For any fundamental frequency subsequence among the multiple fundamental frequency subsequences, determine, from the multiple pitch subsequences, a matching pitch subsequence that matches the any fundamental frequency subsequence.

[0067] In the embodiment of the present application, because the sample fundamental frequency sequence is divided according to the sample singing data, and the sample pitch sequence is divided according to the sample music score, and the sample singing data and the sample music score correspond to the same song, the sample fundamental frequency sequence and the sample pitch sequence are divided according to the same method of notes, and the pitch subsequence and fundamental frequency subsequence corresponding to a note can be determined based on the note; when a fundamental frequency subsequence is determined, a matching pitch subsequence can be determined based on the note corresponding to the fundamental frequency subsequence.

[0068] S504: Determine a jitter frequency and / or a jitter amplitude between any fundamental frequency subsequence and the matching pitch subsequence.

[0069] In the embodiment of the present application, the jitter frequency and / or jitter amplitude can be determined according to any fundamental frequency subsequence and the matching pitch subsequence. The fundamental frequency subsequence can be subjected to a curve fitting operation to obtain a fundamental frequency subsequence curve. Curve fitting refers to fitting discrete data with an appropriate curve type to make it a continuous curve. The fundamental frequency subsequence curve can represent the pitch change process of the fundamental frequency subsequence in the note. Because the fundamental frequency subsequence is obtained according to the sample singing data, and the pitch in a note in the sample singing data is generally variable, the fundamental frequency subsequence curve is generally an undulating curve. And determine the line corresponding to the matching pitch subsequence, which represents the pitch change process of the matching pitch subsequence in the note. Because the matching pitch subsequence is obtained according to the sample music score, and the pitch in each note in the sample music score is generally constant, the line is generally a straight line, and then the jitter frequency and / or jitter amplitude are determined according to the fundamental frequency subsequence curve and the line corresponding to the matching pitch subsequence.

[0070] In one embodiment, the jitter frequency can be determined based on the fundamental frequency subsequence curve and the line corresponding to the matching pitch subsequence. First, the number of intersections between the fundamental frequency subsequence curve and the line corresponding to the matching pitch subsequence is determined, and then the jitter frequency is determined based on the number of intersections and the length of the note corresponding to the fundamental frequency subsequence.

[0071] In one embodiment, the jitter amplitude can be determined based on the fundamental frequency subsequence curve and the line corresponding to the matching pitch subsequence. Any point is selected from the fundamental frequency subsequence curve corresponding to any fundamental frequency subsequence as the first position point. A second position point matching the first position point is determined in the line corresponding to the matching pitch subsequence. The relative value of the first position point is determined based on the absolute value of the difference between the value of the first position point and the value of the second position point. M position points (M is a positive integer greater than 1) are selected in the fundamental frequency subsequence curve, and the relative values ​​of the M position points are calculated based on the fundamental frequency subsequence curve and the line corresponding to the matching pitch subsequence. The jitter amplitude of the fundamental frequency subsequence is then determined based on the M relative values.

[0072] In one embodiment, the jitter amplitude of the base frequency subsequence is determined according to M relative values ​​as follows: the M relative values ​​are sorted from large to small to obtain a relative value queue. The mean is calculated based on the relative values ​​in the last 90% of the relative values ​​in the relative value queue to obtain the jitter amplitude of the base frequency subsequence. Removing the largest 10% of the relative values ​​in the relative value queue can effectively reduce the interference caused by errors and improve the accuracy of the model in predicting the style features of the base frequency sequence. Of course, the percentages arranged in the last are only for illustrative purposes and can also be set to other values ​​according to actual needs.

[0073] See also Figure 6 , Figure 6A schematic diagram of a style parameter acquisition method provided by an embodiment of the present application. In the figure, the length of the note corresponding to the fundamental frequency subsequence is 100 frames, and the number of intersections between the curve corresponding to the fundamental frequency subsequence and the line corresponding to the matching pitch subsequence is 10. Then, according to the number of intersections and the length of the note, the jitter frequency is 0.1. When calculating the jitter amplitude, any point is selected from the fundamental frequency subsequence curve as the first position point. Determine the second position point that matches the first position point in the line corresponding to the matching pitch subsequence. According to the absolute value of the difference between the value of the first position point and the value of the second position point, determine the relative value of the first position point, and the relative value is 0.4. If the value of M is 100 here, 100 position points are selected from the fundamental frequency subsequence curve, and the relative values ​​of the 100 position points are calculated. And the 100 relative values ​​are sorted according to the rule from large to small to obtain a relative value queue. Finally, the average is calculated using the last 90 relative values ​​in the relative queue value, and the jitter amplitude of the fundamental frequency subsequence corresponding to the note is 0.3.

[0074] S505: Generate a jitter frequency subsequence according to the jitter frequency corresponding to each baseband subsequence in the multiple baseband subsequences, and / or generate a jitter amplitude subsequence according to the jitter amplitude corresponding to each baseband subsequence.

[0075] In an embodiment of the present application, the fundamental frequency subsequence is obtained by segmenting the sample fundamental frequency sequence according to the notes in the sample singing data, then the jitter frequency corresponding to each fundamental frequency subsequence in the multiple fundamental frequency subsequences is calculated, and then the jitter frequency is combined according to the order of the notes corresponding to the fundamental frequency subsequences to obtain the jitter frequency subsequence. The jitter amplitude corresponding to each fundamental frequency subsequence in the multiple fundamental frequency subsequences is calculated, and then the jitter frequency is combined according to the order of the notes corresponding to the fundamental frequency subsequences to obtain the jitter amplitude subsequence. When the jitter frequency subsequence and the jitter amplitude subsequence are determined according to the sample fundamental frequency sequence and the sample pitch sequence, the sequence length of the jitter frequency subsequence is the same as that of the jitter amplitude subsequence. At this time, the jitter frequency subsequence and the jitter amplitude subsequence both reflect the jitter situation in the fundamental frequency sequence. When the jitter frequency subsequence or the jitter amplitude subsequence is determined according to the sample fundamental frequency sequence and the sample pitch sequence, the present application does not limit the sequence length of the jitter frequency subsequence or the jitter amplitude subsequence; at this time, the jitter frequency subsequence or the jitter amplitude subsequence both reflect the jitter situation in the fundamental frequency sequence. Training the fundamental frequency prediction model using a style sequence containing a jitter frequency subsequence and / or a jitter amplitude subsequence helps the fundamental frequency prediction model to model the style features of the fundamental frequency sequence, so that the predicted fundamental frequency sequence contains features of the jitter frequency and jitter amplitude.

[0076] S506: Generate the sample style sequence according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

[0077] In an embodiment of the present application, a sample style sequence can be generated according to the jitter frequency subsequence and / or the jitter amplitude subsequence. The sequence length of the sample style sequence is the same as the sequence length of the sample pitch sequence. The sample style sequence contains the style features in the sample fundamental frequency sequence. During the model training process, the sample style sequence enables the model to focus on predicting the style feature details of the fundamental frequency curve in each note, thereby improving the model's detail modeling capability.

[0078] S406: Call the initial fundamental frequency prediction model to process the sample pitch sequence, the sample phoneme sequence and the sample style sequence to obtain the predicted fundamental frequency sequence.

[0079] In the embodiment of the present application, the sample pitch sequence, the sample phoneme sequence and the sample style sequence can be processed by calling the initial fundamental frequency prediction model to obtain a predicted fundamental frequency sequence.

[0080] See also Figure 7 , which is a schematic diagram of the structure of a fundamental frequency prediction model provided by an exemplary embodiment of the present application. In the figure, Conv represents a convolution layer, LeakyReLU represents an activation layer, T.Conv represents a transposed convolution layer, Dense represents a linear transformation layer, and Noise add represents a noise adding operation. Figure 7As shown, the fundamental frequency prediction model mainly includes three parts: a convolution module 71, an encoding module 72 and a decoding module 73; wherein the convolution module 71 includes three independent convolution layers, which are used to process the input pitch sequence, phoneme sequence and style sequence respectively; the encoding module 72 includes four convolution units (numbered 721-724 in the figure), each of which is composed of a convolution layer and an activation layer; the decoding module 73 includes two transposed convolution units (numbered 731 and 732 in the figure) and a linear transformation unit (numbered 733 in the figure), each of which is composed of a transposed convolution layer and an activation layer, and the linear transformation unit is composed of a transposed convolution layer and a linear transformation layer. When performing fundamental frequency prediction processing, the pitch sequence, phoneme sequence and style sequence are first input into the corresponding convolution layers in the convolution module 71 for processing, and the output results of the three convolution layers are fused to obtain the first feature sequence. The first feature sequence is input into the encoding module 72 for encoding processing to obtain a coding feature sequence; and the coding feature sequence is then input into the decoding module 73 for decoding processing. The intermediate features of the transposed convolution unit 731 are fused with the intermediate features of the convolution unit 723, and the fusion results of the two are fused with noise, and input into the transposed convolution unit 732 for processing. The intermediate features of the transposed convolution unit 732 are fused with the intermediate features of the convolution unit 722, and the fusion results of the two are fused with noise, and input into the linear transformation unit 733 for processing to obtain a decoding feature sequence. Finally, the decoding feature sequence is fused with the pitch sequence to obtain a predicted fundamental frequency sequence. The operation of adding noise in the decoding module 73 can enhance the model's ability to model the details of the fundamental frequency sequence, so that the predicted fundamental frequency sequence is more similar to the fundamental frequency sequence corresponding to the score, and the decoding feature sequence is fused with the pitch sequence to obtain a method for predicting the fundamental frequency sequence, which can further ensure the pitch accuracy of the predicted fundamental frequency sequence.

[0081] In one embodiment, the phoneme sequence can be represented by word embedding, so that the representation of the phoneme sequence matches the representation of the pitch sequence. Word embedding is an expression method that means each Chinese word is represented by a vector, and each dimension is an attribute. Figure 7The fundamental frequency prediction model shown in the figure is used for prediction processing, and the process of obtaining the target fundamental frequency sequence can be as follows: Assuming that the target fundamental frequency length is T, the pitch sequence, phoneme sequence and style sequence are obtained, wherein the dimension of the pitch sequence is [T, 1], the phoneme sequence can be represented by word embedding, the dimension corresponding to the phoneme sequence is [T, 128], and the dimension of the style sequence is [T, 2]. The convolution kernel size of the convolution layer in the convolution module 71 is 3, and the number of output channels is 128. First, the pitch sequence, phoneme sequence and style sequence are respectively input into the corresponding convolution layer in the convolution module 71 for processing, and then the outputs of the three convolution layers are fused to obtain a first feature sequence, and the dimension of the first feature sequence is [T, 128]. The first feature sequence is input into the encoding module 72 for encoding processing. The encoding module 72 includes 4 convolution units, each of which is composed of a one-dimensional convolution layer and an activation function layer. The convolution kernel size of the one-dimensional convolution layer is 5, the step size is 2, and the number of input channels and the number of output channels are both 128. After being processed by the encoding module 72, a coding feature sequence is obtained. The coding feature sequence is then input into the decoding module 73 for decoding. The decoding module 73 includes two transposed convolution modules 731 and 732, and a linear transformation module 733. Both the transposed convolution modules 731 and 732 include a transposed convolution layer and an activation function layer. The transposed convolution layer upsamples the input sequence so that the output sequence gradually restores the various features in the original sequence. The number of input channels and the number of output channels of the transposed convolution layer are both 128. The intermediate features of the transposed convolution unit 731 are fused with the intermediate features of the convolution unit 723, and the fusion results of the two are fused with Gaussian noise and input into the transposed convolution unit 732 for processing. The intermediate features of the transposed convolution unit 732 are fused with the intermediate features of the convolution unit 722, and the fusion results of the two are fused with Gaussian noise and input into the linear transformation unit 733 for processing. The linear transformation unit 733 includes a transposed convolution layer and a linear transformation layer. The number of input channels and the number of output channels of the transposed convolution layer are both 128, the number of input channels of the linear transformation layer is 128, and the number of output channels is 1. The linear transformation layer is used to perform a linear transformation on the sequence so that the sequence dimension is consistent with the dimension of the target baseband sequence. After being processed by the linear transformation unit 733, a decoding feature sequence is obtained, and the dimension of the decoding feature sequence is [T, 1], which is the same as the dimension of the target baseband sequence. Finally, the pitch sequence is fused with the decoding feature sequence to obtain a target baseband sequence, and the dimension of the target baseband sequence is [T, 1].

[0082] It should be noted that the structure of the fundamental frequency prediction model provided in the embodiment of the present application is only one possible structure. Other models that can realize fundamental frequency prediction are included in the coverage of the present application.

[0083] S407: Determine the first loss parameter according to the sample fundamental frequency sequence and the predicted fundamental frequency sequence.

[0084] In the embodiment of the present application, the first average error can be obtained according to the sample base frequency sequence and the predicted base frequency sequence. In a feasible implementation, the first average error can be directly used as the first loss parameter, or the first average error can be used as the first loss parameter after weighting.

[0085] In one embodiment, the first average error may be calculated by calculating the minimum mean square error between the sample fundamental frequency sequence and the predicted fundamental frequency sequence; or one or more calculation methods such as calculating absolute value loss, calculating logarithmic loss, calculating cross entropy loss, calculating exponential loss, etc. may be used to calculate the first average error.

[0086] S408: Determine the predicted style sequence according to the sample pitch sequence and the predicted fundamental frequency sequence, and determine the second loss parameter according to the sample style sequence and the predicted style sequence.

[0087] In an embodiment of the present application, the step of determining the predicted style sequence may be: segmenting the sample pitch sequence to obtain multiple pitch subsequences, each pitch subsequence corresponds to a note; and segmenting the predicted fundamental frequency sequence to obtain multiple fundamental frequency subsequences, each fundamental frequency subsequence corresponds to a note. Because the sample pitch sequence and the predicted fundamental frequency sequence are both obtained based on the sample music score, the multiple pitch subsequences obtained after the two are segmented according to the notes can correspond one to one with the multiple predicted fundamental frequency subsequences. For any predicted fundamental frequency subsequence in the multiple predicted fundamental frequency subsequences, a fundamental frequency matching pitch subsequence matching the predicted fundamental frequency subsequence is determined from the multiple pitch subsequences. The predicted fundamental frequency subsequence is curve fitted to obtain a predicted fundamental frequency curve. The jitter frequency and / or jitter amplitude are obtained according to the predicted fundamental frequency curve and the line corresponding to the fundamental frequency matching pitch subsequence. A jitter frequency subsequence is generated according to the jitter frequency corresponding to each predicted fundamental frequency subsequence in the multiple predicted fundamental frequency subsequences, and / or a jitter amplitude subsequence is generated according to the jitter amplitude corresponding to each predicted fundamental frequency subsequence. Finally, a prediction style sequence is generated according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

[0088] In the embodiment of the present application, the second average error can be obtained according to the sample style and the predicted style sequence. In a feasible implementation, the second average error can be directly used as the second loss parameter, or the second average error can be used as the second loss parameter after weighted processing.

[0089] In one embodiment, the second average error may be calculated by calculating the minimum mean square error of the sample style sequence and the predicted style sequence. The second average error may also be calculated by one or more of the following calculation methods: calculating absolute value loss, calculating logarithmic loss, calculating cross entropy loss, calculating exponential loss, etc.

[0090] In one embodiment, the second average error may be a minimum mean square error value between the sample style sequence and the predicted style sequence.

[0091] S409. Determine a target loss parameter according to the first loss parameter and the second loss parameter.

[0092] In an embodiment of the present application, a first weighted parameter corresponding to the first loss parameter can be obtained, and a second weighted parameter corresponding to the second loss parameter can be obtained. The first weighted parameter represents the weight value of the first loss parameter in the two loss parameters, and the second weighted parameter represents the weight value of the second loss parameter in the two loss parameters. Finally, the weighted first loss parameter and the weighted second loss parameter are summed to obtain a target loss parameter. For example: the first weighted parameter can be 3, and the second weighted parameter can be 2, then the first loss parameter can be weighted according to the first weighted parameter, the second loss parameter can be weighted according to the second weighted parameter, and the weighted first loss parameter and the weighted second loss parameter are summed to obtain the target loss parameter.

[0093] In the embodiment of the present application, the first loss parameter is determined based on the sample fundamental frequency sequence and the predicted fundamental frequency sequence. Therefore, the size of the first loss parameter can represent the difference between the predicted fundamental frequency sequence and the sample fundamental frequency sequence in terms of pitch features and phoneme features. Adjusting the model parameters of the fundamental frequency prediction model based on the first loss parameter can make the fundamental frequency prediction model pay more attention to the extraction of pitch features and phoneme features when predicting the fundamental frequency, thereby improving the accuracy of the predicted fundamental frequency sequence. The second loss parameter is determined based on the style sequence corresponding to the sample style sequence and the predicted fundamental frequency sequence. Therefore, the size of the second loss parameter can represent the difference between the predicted fundamental frequency sequence and the sample fundamental frequency sequence in terms of style features. Adjusting the parameters of the fundamental frequency prediction model based on the second loss parameter can make the fundamental frequency prediction model pay more attention to the extraction of style features of the fundamental frequency sequence in each note when predicting the fundamental frequency, thereby increasing the jitter details of the predicted fundamental frequency sequence. The integration of two loss parameters to adjust the parameters of the fundamental frequency prediction model can enhance the nonlinear modeling ability of the fundamental frequency prediction model for the fundamental frequency style characteristics, while improving the matching degree between the fundamental frequency prediction sequence and the music score, so that the singing audio synthesized according to the fundamental frequency sequence has a high similarity with the singing audio sung by real people.

[0094] In one embodiment, the second weighting parameter is greater than the first weighting parameter. For example, the first weighting parameter is 1 and the second weighting parameter is 5. The second weighting parameter is larger, so that the second loss parameter accounts for a larger proportion of the target loss parameter. The fundamental frequency prediction model is adjusted according to the target loss parameter, so that the fundamental frequency prediction model focuses on predicting the style characteristics of the fundamental frequency sequence, so that the predicted fundamental frequency sequence is more similar to the fundamental frequency sequence corresponding to the singing voice of a real person, and the jitter details are richer.

[0095] S410: Adjust model parameters of the initial fundamental frequency prediction model according to the target loss parameter to obtain the target fundamental frequency prediction model.

[0096] In the embodiment of the present application, adjusting the parameters of the initial fundamental frequency prediction model according to the target loss parameter can be to update the relevant parameters of the initial fundamental frequency prediction model according to the gradient back propagation algorithm. The gradient back propagation (BackwardPropagation) algorithm is an algorithm that uses the chain rule to recursively calculate the gradient of an expression. In the embodiment of the present application, the sample data includes a sample score and a sample singing data corresponding to the sample score. Then the sample data set containing multiple groups of sample data can be input into the initial fundamental frequency prediction model for fundamental frequency prediction, and multiple target loss parameters are obtained according to the result of the fundamental frequency prediction. The multiple target loss parameters are processed to obtain the final loss parameter. The parameters of the initial fundamental frequency prediction model are updated according to the final loss parameter using the gradient back propagation algorithm and the adaptive moment estimation (Adaptive Moment Estimation, Adam) optimizer. The Adam optimizer can realize the update of model parameters according to the historical gradient obtained by the gradient back propagation algorithm. Using the gradient back propagation algorithm and the Adam optimizer, the relevant parameters in the fundamental frequency prediction model can be accurately and quickly updated, thereby improving the prediction accuracy of the fundamental frequency prediction model.

[0097] It should be noted that the above-mentioned model training process can be repeated multiple times. The prediction accuracy of the fundamental frequency prediction model can be verified with a test data set. The test data set can be a part of the sample data set, or it can be another data set different from the sample data set. The test data set is input into the initial fundamental frequency prediction model after parameter adjustment to perform fundamental frequency prediction, and a test loss parameter is obtained according to the result of the fundamental frequency prediction. When the test loss parameter is less than a preset value, it can be considered that the model training of the fundamental frequency prediction model is completed, and the initial fundamental frequency prediction model after the parameter adjustment is used as the target fundamental frequency prediction model.

[0098] Based on the above embodiments, the beneficial effects of the present application are: in the process of model training, the fundamental frequency prediction model is trained according to the sample pitch sequence, the sample phoneme sequence and the sample style sequence, so that the model focuses on the detailed modeling of the fundamental frequency sequence in each note, thereby improving the modeling ability of the fundamental frequency prediction model; Gaussian noise is added to the model to make the model's prediction of details more realistic; the output feature sequence and the pitch sequence are fused to further improve the accuracy of the predicted fundamental frequency sequence and reduce the possibility of being out of tune. In the process of using the model, the fundamental frequency sequence corresponding to the score can be obtained directly according to the score, and because a style sequence including jitter frequency and / or jitter amplitude is added, the fundamental frequency sequence contains appropriate jitter, which is more similar to the fundamental frequency sequence corresponding to the singing voice of a real person.

[0099] See also Figure 8 , Figure 8 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application. Figure 8 The data processing device shown may specifically include:

[0100] The acquisition unit 801 is used to acquire a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include a jitter frequency and / or a jitter amplitude;

[0101] The processing unit 802 is used to parse the target music score to obtain a pitch sequence and a phoneme sequence of the target music score;

[0102] The processing unit 802 is further used to perform fundamental frequency prediction processing according to the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singer's singing melody.

[0103] In one embodiment, when the processing unit 802 is used to perform fundamental frequency prediction processing according to the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score, it is specifically used to: call the target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score; wherein the target fundamental frequency prediction model is obtained by training based on sample data, and the sample data includes a sample pitch sequence, a sample phoneme sequence, a sample style sequence and a sample fundamental frequency sequence; the sample pitch sequence and the sample phoneme sequence are determined based on the analysis result of the sample music score, and the sample style sequence is determined based on the sample music score. The target fundamental frequency prediction model is obtained by adjusting the model parameters of the initial fundamental frequency prediction model according to the first loss parameter and the second loss parameter, the first loss parameter is determined according to the sample fundamental frequency sequence and the predicted fundamental frequency sequence, the predicted fundamental frequency sequence is obtained by calling the initial fundamental frequency prediction model to process the sample pitch sequence, the sample phoneme sequence and the sample style sequence, the second loss parameter is determined according to the sample style sequence and the predicted style sequence, and the predicted style sequence is determined according to the sample pitch sequence and the predicted fundamental frequency sequence.

[0104] In one embodiment, the target fundamental frequency prediction model includes a convolution module, an encoding module and a decoding module. When the processing unit 802 calls the target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score, it is specifically used to: input the pitch sequence, the phoneme sequence and the target style sequence into the target fundamental frequency prediction model for processing respectively, and the convolution module performs convolution processing on the pitch sequence, the phoneme sequence and the target style sequence respectively, and fuses the convolution processing results corresponding to the pitch sequence, the phoneme sequence and the target style sequence respectively to obtain a first feature sequence; the encoding module encodes the first feature sequence to obtain a coded feature sequence, and the decoding module decodes the coded feature sequence to obtain a decoded feature sequence; the decoded feature sequence is fused with the pitch sequence to obtain a target fundamental frequency sequence corresponding to the target music score.

[0105] In one embodiment, the processing unit 802 is further used to: perform frame processing on the sample singing data corresponding to the sample music score to obtain multiple frames of singing audio; perform fundamental frequency extraction processing on each frame of singing audio in the multiple frames of singing audio to obtain the fundamental frequency value corresponding to each frame of singing audio; arrange the fundamental frequency values ​​corresponding to each frame of singing audio according to the sequence of playback time corresponding to each frame of singing audio to obtain a reference fundamental frequency sequence; determine the sample fundamental frequency sequence according to the reference fundamental frequency sequence.

[0106] In one embodiment, when the processing unit 802 is used to determine the sample fundamental frequency sequence according to the reference fundamental frequency sequence, it is specifically used to: determine the reference fundamental frequency sequence as the sample fundamental frequency sequence; or perform offset processing on the fundamental frequency value in the reference fundamental frequency sequence to obtain the sample fundamental frequency sequence.

[0107] In one embodiment, the processing unit 802 is further used to: parse the sample music score to obtain a reference pitch sequence of the sample music score, and determine the sample pitch sequence according to the reference pitch sequence; segment the sample pitch sequence to obtain multiple pitch subsequences, each pitch subsequence corresponding to a note; segment the sample fundamental frequency sequence to obtain multiple fundamental frequency subsequences, each fundamental frequency subsequence corresponding to a note; for any fundamental frequency subsequence in the multiple fundamental frequency subsequences, determine a matching pitch subsequence that matches the any fundamental frequency subsequence from the multiple pitch subsequences, and determine the jitter frequency and / or jitter amplitude between the any fundamental frequency subsequence and the matching pitch subsequence; generate a jitter frequency subsequence according to the jitter frequency corresponding to each fundamental frequency subsequence in the multiple fundamental frequency subsequences, and / or generate a jitter amplitude subsequence according to the jitter amplitude corresponding to each fundamental frequency subsequence; generate the sample style sequence according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

[0108] In one embodiment, when the processing unit 802 is used to determine the jitter frequency and jitter amplitude between any fundamental frequency subsequence and the matching pitch subsequence, it is specifically used to: determine the number of intersections between the curve corresponding to any fundamental frequency subsequence and the line corresponding to the matching pitch subsequence, and determine the jitter frequency according to the number of intersections and the length of the note corresponding to any fundamental frequency subsequence; determine the difference between the value of the first position point in the curve corresponding to any fundamental frequency subsequence and the value of the second position point in the line corresponding to the matching pitch subsequence, and determine the jitter amplitude according to the difference corresponding to each position point in the curve; the first position point is any position point on the curve, and the second position point is a position point in the line that matches the first position point.

[0109] In one embodiment, the above-mentioned processing unit 802 is also used to: calculate a first average error between the sample base frequency sequence and the predicted base frequency sequence, and determine the first loss parameter based on the first average error; when used to determine the second loss parameter based on the sample style sequence and the predicted style sequence, it is specifically used to: calculate a second average error between the sample style sequence and the predicted style sequence, and determine the second loss parameter based on the second average error.

[0110] In one embodiment, the processing unit 802 is further used to: obtain a first weighted parameter corresponding to the first loss parameter, and perform weighted processing on the first loss parameter using the first weighted parameter to obtain the weighted first loss parameter; obtain a second weighted parameter corresponding to the second loss parameter, and perform weighted processing on the second loss parameter using the second weighted parameter to obtain the weighted second loss parameter; the numerical value corresponding to the second loss parameter is greater than the numerical value corresponding to the first loss parameter; and sum the weighted first loss parameter and the weighted second loss parameter to obtain a target loss parameter.

[0111] It can be understood that the functions of each functional unit of the data processing device of the embodiment of the present application can be specifically implemented according to the data processing method in the above method embodiment, and its specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0112] Based on the above embodiments, the beneficial effects of the present application are: in the process of model training, the fundamental frequency prediction model is trained according to the sample pitch sequence, the sample phoneme sequence and the sample style sequence, so that the model focuses on the detailed modeling of the fundamental frequency sequence in each note, thereby improving the modeling ability of the fundamental frequency prediction model; Gaussian noise is added to the model to make the model's prediction of details more realistic; the output feature sequence and the pitch sequence are fused to further improve the accuracy of the predicted fundamental frequency sequence and reduce the possibility of being out of tune. In the process of using the model, the fundamental frequency sequence corresponding to the score can be obtained directly according to the score, and because a style sequence including jitter frequency and / or jitter amplitude is added, the fundamental frequency sequence contains appropriate jitter, which is more similar to the fundamental frequency sequence corresponding to the singing voice of a real person.

[0113] See also Fig. 9 , Fig. 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device described in the embodiment of the present application includes: a processor 901 and a memory 902. The processor 901 and the memory 902 may be connected via a bus or other means, and the embodiment of the present application takes the connection via a bus as an example.

[0114] Among them, the processor 901 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power on and off instructions sent by the user to the computer device, and control the computer device to perform power on and off operations; another example: the CPU can transmit various interactive data between the internal structures of the computer device, and so on. The memory 902 (Memory) is a memory device in the computer device, which is used to store programs and data. It can be understood that the memory 902 here can include the built-in memory of the computer device, and of course, it can also include the extended memory supported by the computer device. The memory 902 provides a storage space, which stores the operating system of the computer device, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc., and this application does not limit this.

[0115] In the embodiment of the present application, the processor 901 performs the following operations by running the executable program code in the memory 902:

[0116] Acquire a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is used to indicate the style characteristics of the predicted fundamental frequency, and the target style sequence is determined according to style parameters, and the style parameters include jitter frequency and / or jitter amplitude;

[0117] Analyzing the target music score to obtain a pitch sequence and a phoneme sequence of the target music score;

[0118] A fundamental frequency prediction process is performed according to the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singing melody of the singer.

[0119] In one embodiment, when the processor 901 is used to perform fundamental frequency prediction processing according to the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score, it is specifically used to: call the target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score; wherein the target fundamental frequency prediction model is obtained by training based on sample data, and the sample data includes a sample pitch sequence, a sample phoneme sequence, a sample style sequence and a sample fundamental frequency sequence; the sample pitch sequence and the sample phoneme sequence are determined based on the analysis result of the sample music score, and the sample style sequence is determined based on the sample The target fundamental frequency prediction model is obtained by adjusting the model parameters of the initial fundamental frequency prediction model according to the first loss parameter and the second loss parameter, the first loss parameter is determined according to the sample fundamental frequency sequence and the predicted fundamental frequency sequence, the predicted fundamental frequency sequence is obtained by calling the initial fundamental frequency prediction model to process the sample pitch sequence, the sample phoneme sequence and the sample style sequence, the second loss parameter is determined according to the sample style sequence and the predicted style sequence, and the predicted style sequence is determined according to the sample pitch sequence and the predicted fundamental frequency sequence.

[0120] In one embodiment, the target fundamental frequency prediction model includes a convolution module, an encoding module and a decoding module. When the processor 901 calls the target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain the target fundamental frequency sequence corresponding to the target music score, it is specifically used to: input the pitch sequence, the phoneme sequence and the target style sequence into the target fundamental frequency prediction model for processing respectively, and the convolution module performs convolution processing on the pitch sequence, the phoneme sequence and the target style sequence respectively, and fuses the convolution processing results corresponding to the pitch sequence, the phoneme sequence and the target style sequence respectively to obtain a first feature sequence; the encoding module encodes the first feature sequence to obtain a coded feature sequence, and the decoding module decodes the coded feature sequence to obtain a decoded feature sequence; the decoded feature sequence is fused with the pitch sequence to obtain a target fundamental frequency sequence corresponding to the target music score.

[0121] In one embodiment, the processor 901 is further used to: perform frame processing on the sample singing data corresponding to the sample music score to obtain multiple frames of singing audio; perform fundamental frequency extraction processing on each frame of singing audio in the multiple frames of singing audio to obtain the fundamental frequency value corresponding to each frame of singing audio; arrange the fundamental frequency values ​​corresponding to each frame of singing audio according to the sequence of playback time corresponding to each frame of singing audio to obtain a reference fundamental frequency sequence; determine the sample fundamental frequency sequence according to the reference fundamental frequency sequence.

[0122] In one embodiment, when the processor 901 is used to determine the sample fundamental frequency sequence according to the reference fundamental frequency sequence, it is specifically used to: determine the reference fundamental frequency sequence as the sample fundamental frequency sequence; or perform offset processing on the fundamental frequency value in the reference fundamental frequency sequence to obtain the sample fundamental frequency sequence.

[0123] In one embodiment, the processor 901 is further used to: parse the sample music score to obtain a reference pitch sequence of the sample music score, and determine the sample pitch sequence according to the reference pitch sequence; segment the sample pitch sequence to obtain multiple pitch subsequences, each pitch subsequence corresponds to a note; segment the sample fundamental frequency sequence to obtain multiple fundamental frequency subsequences, each fundamental frequency subsequence corresponds to a note; for any fundamental frequency subsequence in the multiple fundamental frequency subsequences, determine a matching pitch subsequence that matches the any fundamental frequency subsequence from the multiple pitch subsequences, and determine the jitter frequency and / or jitter amplitude between the any fundamental frequency subsequence and the matching pitch subsequence; generate a jitter frequency subsequence according to the jitter frequency corresponding to each fundamental frequency subsequence in the multiple fundamental frequency subsequences, and / or generate a jitter amplitude subsequence according to the jitter amplitude corresponding to each fundamental frequency subsequence; generate the sample style sequence according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

[0124] In one embodiment, when the processor 901 is used to determine the jitter frequency and jitter amplitude between any fundamental frequency subsequence and the matching pitch subsequence, it is specifically used to: determine the number of intersections between the curve corresponding to any fundamental frequency subsequence and the line corresponding to the matching pitch subsequence, and determine the jitter frequency according to the number of intersections and the length of the note corresponding to any fundamental frequency subsequence; determine the difference between the value of the first position point in the curve corresponding to any fundamental frequency subsequence and the value of the second position point in the line corresponding to the matching pitch subsequence, and determine the jitter amplitude according to the difference corresponding to each position point in the curve; the first position point is any position point on the curve, and the second position point is a position point in the line that matches the first position point.

[0125] In one embodiment, the processor 901 is further used to: calculate a first average error between the sample base frequency sequence and the predicted base frequency sequence, and determine the first loss parameter based on the first average error; when used to determine the second loss parameter based on the sample style sequence and the predicted style sequence, it is specifically used to: calculate a second average error between the sample style sequence and the predicted style sequence, and determine the second loss parameter based on the second average error.

[0126] In one embodiment, the processor 901 is further used to: obtain a first weighted parameter corresponding to the first loss parameter, and perform weighted processing on the first loss parameter using the first weighted parameter to obtain the weighted first loss parameter; obtain a second weighted parameter corresponding to the second loss parameter, and perform weighted processing on the second loss parameter using the second weighted parameter to obtain the weighted second loss parameter; the numerical value corresponding to the second loss parameter is greater than the numerical value corresponding to the first loss parameter; and sum the weighted first loss parameter and the weighted second loss parameter to obtain a target loss parameter.

[0127] In a specific implementation, the processor 901 and the memory 902 described in the embodiment of the present application can execute an implementation of a computer device described in a data processing method provided in an embodiment of the present application, and can also execute an implementation described in a data processing device provided in an embodiment of the present application, which will not be repeated here.

[0128] Based on the above embodiments, the beneficial effects of the present application are: in the process of model training, the fundamental frequency prediction model is trained according to the sample pitch sequence, the sample phoneme sequence and the sample style sequence, so that the model focuses on the detailed modeling of the fundamental frequency sequence in each note, thereby improving the modeling ability of the fundamental frequency prediction model; Gaussian noise is added to the model to make the model's prediction of details more realistic; the output feature sequence and the pitch sequence are fused to further improve the accuracy of the predicted fundamental frequency sequence and reduce the possibility of being out of tune. In the process of using the model, the fundamental frequency sequence corresponding to the score can be obtained directly according to the score, and because a style sequence including jitter frequency and / or jitter amplitude is added, the fundamental frequency sequence contains appropriate jitter, which is more similar to the fundamental frequency sequence corresponding to the singing voice of a real person.

[0129] In the several embodiments provided in the present application, it should be understood that the disclosed methods, devices and systems can be implemented in other ways. For example, the data processing device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0130] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is run on a computer, the computer executes the data processing method provided in the embodiment of the present application. The specific implementation method can be referred to the above description, and will not be repeated here.

[0131] The embodiment of the present application also provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method provided in the embodiment of the present application. The specific implementation method can be referred to the above description, which will not be repeated here.

[0132] It should be noted that, for the above-mentioned various method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0133] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0134] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire a target music score and a target style sequence corresponding to the target music score; wherein the target style sequence is determined according to style parameters, and the style parameters include a jitter frequency and / or a jitter amplitude; Analyzing the target music score to obtain a pitch sequence and a phoneme sequence of the target music score; Calling a target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score; the target fundamental frequency sequence is used to indicate the singer's singing melody; Among them, the target fundamental frequency prediction model is obtained by adjusting the model parameters of the initial fundamental frequency prediction model according to the first loss parameter and the second loss parameter, the first loss parameter is determined according to the sample fundamental frequency sequence and the predicted fundamental frequency sequence, the predicted fundamental frequency sequence is obtained by calling the initial fundamental frequency prediction model to process the sample pitch sequence, the sample phoneme sequence and the sample style sequence, the second loss parameter is determined according to the sample style sequence and the predicted style sequence, and the predicted style sequence is determined according to the sample pitch sequence and the predicted fundamental frequency sequence; the sample pitch sequence and the sample phoneme sequence are determined according to the analysis results of the sample music score, the sample style sequence is determined according to the sample pitch sequence and the sample fundamental frequency sequence, and the sample fundamental frequency sequence is determined according to the fundamental frequency of the sample singing data corresponding to the sample music score.

2. The method according to claim 1, characterized in that: The target fundamental frequency prediction model includes a convolution module, an encoding module and a decoding module; the calling of the target fundamental frequency prediction model to process the pitch sequence, the phoneme sequence and the target style sequence to obtain a target fundamental frequency sequence corresponding to the target music score includes: The pitch sequence, the phoneme sequence and the target style sequence are respectively input into the target fundamental frequency prediction model, the convolution module performs convolution processing on the pitch sequence, the phoneme sequence and the target style sequence respectively, and the convolution processing results corresponding to the pitch sequence, the phoneme sequence and the target style sequence are respectively fused to obtain a first feature sequence; The encoding module encodes the first feature sequence to obtain a coded feature sequence, and the decoding module decodes the coded feature sequence to obtain a decoded feature sequence; The decoded feature sequence is fused with the pitch sequence to obtain a target fundamental frequency sequence corresponding to the target music score.

3. The method according to claim 1, characterized in that The method further comprises: Performing frame processing on the sample singing data corresponding to the sample music score to obtain multiple frames of singing audio; Performing a fundamental frequency extraction process on each frame of singing audio in the multiple frames of singing audio to obtain a fundamental frequency value corresponding to each frame of singing audio; Arrange the fundamental frequency values ​​corresponding to each frame of the singing audio according to the sequence of the playback time corresponding to each frame of the singing audio to obtain a reference fundamental frequency sequence; The sample base frequency sequence is determined according to the reference base frequency sequence.

4. The method according to claim 3, characterized in that The determining the sample fundamental frequency sequence according to the reference fundamental frequency sequence comprises: Determine the reference base frequency sequence as the sample base frequency sequence; or, An offset process is performed on the fundamental frequency values ​​in the reference fundamental frequency sequence to obtain the sample fundamental frequency sequence.

5. The method according to claim 1, characterized in that The method further comprises: The sample music score is analyzed to obtain a reference pitch sequence of the sample music score, and the sample pitch sequence is determined based on the reference pitch sequence.

6. The method according to claim 1, characterized in that The method further comprises: Segmenting the sample pitch sequence to obtain a plurality of pitch subsequences, each pitch subsequence corresponding to a note; Segmenting the sample fundamental frequency sequence to obtain a plurality of fundamental frequency subsequences, each fundamental frequency subsequence corresponding to a note; For any fundamental frequency subsequence among the multiple fundamental frequency subsequences, determine a matching pitch subsequence that matches the any fundamental frequency subsequence from the multiple pitch subsequences, and determine a jitter frequency and / or jitter amplitude between the any fundamental frequency subsequence and the matching pitch subsequence; Generate a jitter frequency subsequence according to the jitter frequency corresponding to each base frequency subsequence in the multiple base frequency subsequences, and / or generate a jitter amplitude subsequence according to the jitter amplitude corresponding to each base frequency subsequence; The sample style sequence is generated according to the jitter frequency subsequence and / or the jitter amplitude subsequence.

7. The method according to claim 6, characterized in that The determining of the jitter frequency and jitter amplitude between any one of the base frequency subsequences and the matching pitch subsequence comprises: Determine the number of intersections between the curve corresponding to any fundamental frequency subsequence and the line corresponding to the matching pitch subsequence, and determine the jitter frequency according to the number of intersections and the length of the note corresponding to any fundamental frequency subsequence; Determine the difference between the value of the first position point in the curve corresponding to any fundamental frequency subsequence and the value of the second position point in the line corresponding to the matching pitch subsequence, and determine the jitter amplitude according to the difference corresponding to each position point in the curve; the first position point is any position point on the curve, and the second position point is a position point in the line that matches the first position point.

8. The method according to claim 1, characterized in that: The method further comprises: Calculating a first averaged error between the sample fundamental frequency sequence and the predicted fundamental frequency sequence, and determining the first loss parameter according to the first averaged error; A second average error between the sample style sequence and the predicted style sequence is calculated, and the second loss parameter is determined according to the second average error.

9. The method according to claim 1, characterized in that: The method further comprises: Obtaining a first weighting parameter corresponding to the first loss parameter, and performing weighted processing on the first loss parameter using the first weighting parameter to obtain a weighted first loss parameter; Obtaining a second weighting parameter corresponding to the second loss parameter, and performing weighted processing on the second loss parameter using the second weighting parameter to obtain a weighted second loss parameter; the value corresponding to the second loss parameter is greater than the value corresponding to the first loss parameter; The weighted first loss parameter and the weighted second loss parameter are summed to obtain a target loss parameter; The model parameters of the initial fundamental frequency prediction model are adjusted according to the target loss parameter to obtain the target fundamental frequency prediction model.

10. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores executable program code, and the processor is used to call the executable program code to implement the data processing method according to any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which, when executed on a computer, enable the computer to implement the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice style migration method and device, electronic equipment and storage medium

    CN113963679A

  • Voice processing method and device, equipment and storage medium

    CN114495956A