An audio processing method, apparatus, device and storage medium

By acquiring user singing ability representation data and using reinforcement learning models to adjust audio parameters, accompaniment audio suitable for the user's level is generated, solving the problem that existing singing software cannot dynamically adjust the difficulty of songs, and improving the user's singing experience and ability.

CN122116858APending Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-11-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing singing software cannot dynamically adjust the difficulty of songs based on the user's singing ability, resulting in a poor user singing experience.

Method used

By acquiring user singing ability representation data, a reinforcement learning model is used to adjust the audio parameters of standard accompaniment audio to generate target accompaniment audio suitable for the user's level.

Benefits of technology

It improves the user's singing experience and ability, making the difficulty of songs closer to the user's current level, and enhancing performance in long-term singing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116858A_ABST
    Figure CN122116858A_ABST
Patent Text Reader

Abstract

The application discloses an audio processing method, device, equipment and storage medium. Specifically, the method comprises: obtaining singing ability representation data of a target object and standard accompaniment audio of a song to be sung; determining parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data; inputting the parameter adaptation data into a preset audio parameter adjustment model to obtain audio parameter adjustment data; the preset audio parameter adjustment model is obtained by performing real-time audio parameter adjustment training on a preset reinforcement learning model based on historical parameter adaptation data of historical songs, historical audio parameter adjustment data of the historical songs and historical singing performance data of the target object singing the historical songs; and based on the audio parameter adjustment data, performing parameter configuration on the standard accompaniment audio to obtain target accompaniment audio of the song to be sung. The technical solution provided in the application can adaptively adjust the audio parameters of the song to be sung, so that the singing difficulty of the song to be sung can be closer to the current singing level of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, device and storage medium. Background Technology

[0002] Existing singing software typically provides all users with the same fixed services and functions. When a user selects a song, they can only sing according to the standard original audio parameters of that song. If the user's singing ability does not meet the singing skills requirements of the selected song, the corresponding singing performance will be unsatisfactory, affecting the user's singing experience. Summary of the Invention

[0003] This application provides an audio processing method, apparatus, device, and storage medium that can adaptively adjust audio parameters to make the difficulty of the song to be sung more closely match the current singing level of the target, thereby improving the target's singing ability and enhancing the target's singing experience over a long period of time. The technical solution of this application is as follows:

[0004] On the one hand, an audio processing method is provided, the method comprising:

[0005] Obtain the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object;

[0006] Determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data;

[0007] The parameter adaptation data is input into the preset audio parameter tuning model, and the standard accompaniment audio is subjected to audio parameter tuning analysis to obtain audio parameter tuning data. The preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of the historical songs, and the historical performance data of the target object singing the historical songs.

[0008] Based on the audio parameter tuning data, the standard accompaniment audio is configured with parameters to obtain the target accompaniment audio corresponding to the song to be sung.

[0009] On the other hand, an audio processing apparatus is provided, the apparatus comprising:

[0010] The standard accompaniment audio acquisition module is used to acquire the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object.

[0011] The parameter adaptation data determination module is used to determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability characterization data.

[0012] The audio parameter tuning and analysis module is used to input the parameter adaptation data into the preset audio parameter tuning model, perform audio parameter tuning analysis on the standard accompaniment audio, and obtain audio parameter tuning data. The preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of the historical songs, and the historical performance data of the target object singing the historical songs.

[0013] The first parameter configuration module is used to configure the parameters of the standard accompaniment audio based on the audio parameter tuning data to obtain the target accompaniment audio corresponding to the song to be sung.

[0014] On the other hand, an audio processing device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the audio processing method as described in the first aspect.

[0015] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the audio processing method as described in the first aspect.

[0016] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio processing method as described in the first aspect.

[0017] The audio processing method, apparatus, device, and storage medium provided in this application have the following technical effects:

[0018] In the application scenario of song accompaniment audio processing based on reinforcement learning, this application uses historical parameter adaptation data of historical songs relative to the historical singing ability of the target object, historical audio parameter tuning data of historical songs, and historical singing performance data of the target object singing historical songs to train a preset reinforcement learning model for audio parameter tuning, so that the model can decide on the accompaniment audio adjustment parameters that are adapted to the singing level of the target object. Then, it determines the parameter adaptation data of the standard accompaniment audio of the song to be sung by the target object relative to the singing ability representation data of the target object, and inputs the parameter adaptation data into the preset audio parameter tuning model to obtain audio parameter tuning data. Based on the audio parameter tuning data, the parameters of the standard accompaniment audio are configured to obtain the target accompaniment audio of the song to be sung. Through adaptive audio parameter adjustment, the singing difficulty of the song to be sung can be made closer to the current singing level of the target object, so as to improve the singing ability of the target object in the long-term singing process, thereby improving the singing experience of the target object. Attached Figure Description

[0019] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;

[0021] Figure 2 This is a schematic flowchart of an audio processing method provided in an embodiment of this application;

[0022] Figures 3a-3c This is a schematic diagram of the pitch error distribution histogram provided in the embodiments of this application;

[0023] Figure 4 This is a flowchart illustrating the training process of a preset audio parameter tuning model provided in an embodiment of this application;

[0024] Figure 5 This is a flowchart illustrating a historical audio parameter tuning data decision-making process provided in an embodiment of this application;

[0025] Figure 6 This is a flowchart illustrating a historical performance data calibration process provided in an embodiment of this application;

[0026] Figure 7 This is a schematic diagram of a framework for an audio parameter tuning and training process based on reinforcement learning, provided in an embodiment of this application.

[0027] Figure 8 This is a flowchart illustrating how parameter adaptation data is input into a preset audio parameter tuning model to perform audio parameter tuning analysis on standard accompaniment audio and obtain audio parameter tuning data, as provided in this application embodiment.

[0028] Figures 9a-9b This is a flowchart illustrating another audio processing method provided in an embodiment of this application;

[0029] Figure 10a This is a schematic diagram of an audio parameter adjustment scheme for a karaoke scenario provided in an embodiment of this application;

[0030] Figure 10b This is a schematic diagram of a karaoke interface provided in an embodiment of this application;

[0031] Figure 10c This is a schematic diagram of a karaoke scoring interface provided in an embodiment of this application;

[0032] Figure 11 This is a block diagram of an audio processing device provided in an embodiment of this application;

[0033] Figure 12 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0035] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products or devices.

[0036] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0037] To facilitate understanding of the embodiments of this application, several concepts will be briefly introduced below:

[0038] Reinforcement learning, a crucial branch of machine learning, is a machine learning method that uses environmental feedback as input and learns autonomously through continuous exploration and trial to map environmental states to actions. Reinforcement learning awards a reward for each action attempt, optimizing the action by accumulating the maximum reward. Unlike supervised learning, reinforcement learning does not require pre-provided training samples; it is an online learning technique. The reinforcement learning agent only needs to memorize its environmental state and current policy knowledge, acquiring the optimal decision for the current environment through accumulated exploration experience.

[0039] Online karaoke software is music entertainment software installed on personal devices. Leveraging the device's recording and playback capabilities, users select songs from a rich song library, record themselves singing along to the background music, and can transmit audio data in real-time or offline via the internet for multi-person duets and song sharing. Online karaoke software also offers features such as singing scoring, sound effects processing, sound quality enhancement, and community interaction, allowing users to enjoy the singing experience while improving their singing skills.

[0040] Karaoke scoring systems are intelligent systems that analyze a user's singing performance to provide scores and feedback. These systems are typically integrated into karaoke software, KTV song selection systems, and some music games. The main purpose of karaoke scoring systems is to help users better understand their singing skills and improve them while having fun.

[0041] Pitch-based scoring algorithms: These algorithms primarily focus on the pitch accuracy of a user's singing. They analyze the user's audio signal, extract pitch information, and compare it to the pitch of the original song. A score is calculated based on the degree of match between the user's pitch and the original song's pitch. The advantage of these algorithms is their simplicity and ease of implementation, but they may not be able to comprehensively evaluate a user's singing performance.

[0042] Rhythm-based scoring algorithms: These algorithms focus on the rhythmic accuracy of a user's singing. They analyze the rhythmic information in the user's audio signal and compare it to the rhythm of the original song. A score is calculated based on the degree of match between the user's rhythm and the original song's rhythm. These algorithms can assist pitch scoring algorithms, providing a more comprehensive evaluation.

[0043] Acoustic model-based scoring algorithms: These algorithms use acoustic models (such as MFCC, PLP, etc.) to extract features from the user's audio signal and compare them with the features of the original song. A score is calculated based on feature similarity. These algorithms can evaluate a user's timbre, pronunciation, and other performance aspects, but they are computationally complex.

[0044] Deep learning-based scoring algorithms: These algorithms use deep learning techniques (such as convolutional neural networks and recurrent neural networks) to analyze user audio signals, automatically extract useful features, and compare them with the original song. These algorithms can evaluate a user's singing across multiple dimensions (such as pitch, rhythm, and timbre), but require a large amount of training data and computational resources.

[0045] Comprehensive scoring algorithm: This type of algorithm takes into account multiple dimensions (such as pitch, rhythm, timbre, etc.) and calculates the user's total score through weighted summation and other methods.

[0046] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application. This application environment may include a client 10 and a server 20, which can be indirectly connected via wireless communication. The client 10 sends an audio processing command to the server 20. In response to the audio processing command, the server 20 obtains the standard accompaniment audio corresponding to the song to be sung by the target object and the target object's singing ability representation data. It determines the parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data, and inputs the parameter adaptation data into a preset audio parameter tuning model. The server performs audio parameter tuning analysis on the standard accompaniment audio to obtain audio parameter tuning data, which is then fed back to the client 10. Based on the audio parameter tuning data, the client 10 configures the parameters of the standard accompaniment audio to obtain the target accompaniment audio corresponding to the song to be sung. The preset audio parameter tuning model is obtained by the server 20 through real-time audio parameter tuning training of a preset reinforcement learning model based on historical parameter adaptation data of historical songs, historical audio parameter tuning data of historical songs, and historical singing performance data of the target object singing historical songs. It should be noted that... Figure 1 This is just one example.

[0047] The client can be a physical device such as a smartphone, computer (e.g., desktop computer, tablet computer, laptop computer), digital assistant, smart voice interaction device (e.g., smart speaker), smart wearable device, in-vehicle terminal, etc., or it can be software running on the physical device, such as a computer program. The operating system corresponding to the first client can be Android, iOS (a mobile operating system developed by Apple), Linux, Microsoft Windows, etc.

[0048] The server side can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may include network communication units, processors, and memory, etc. The server side can provide backend services to the corresponding clients.

[0049] The aforementioned client 10 and server 20 can be used to build an audio processing system, which can be a distributed system.

[0050] It should be noted that the audio processing method provided in this application can be applied to both the client and the server, and is not limited to the embodiments described above.

[0051] The following describes a specific embodiment of an audio processing method provided in this application. Figure 2 This is a flowchart illustrating an audio processing method provided in an embodiment of this application. This application provides the operational steps of the method described in the embodiment or flowchart, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only execution order. In actual systems or products, the method can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiment or drawings. Specifically, as... Figure 2 As shown, the method may include:

[0052] S201, Obtain the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object.

[0053] In the embodiments of this specification, the target object can be a user of a preset application or a user account. Specifically, the preset application can be a backing track application that provides singing performance analysis functions. Optionally, the preset application can be a client application or a web application. In a specific embodiment, the song to be sung can be a song selected by the target object from multiple candidate songs provided by the preset application, and the standard backing track audio can be the original backing track audio provided by the preset application without parameter tuning. Optionally, the standard backing track audio can be obtained from the local music library or a remote music library (backend server) of the preset application.

[0054] In a specific embodiment, the singing ability representation data of the target object can be used to represent the singing ability of the target object in multiple dimensions. Specifically, the multiple dimensions may include, but are not limited to, pitch expression ability, rhythm expression ability and breath expression ability. Correspondingly, the singing ability representation data may include pitch expression ability representation data, rhythm expression ability representation data and breath expression ability representation data.

[0055] In one specific embodiment, pitch accuracy representation data can be used to characterize the pitch expression ability of a target object. Specifically, pitch accuracy representation data may include: upper treble limit data and / or lower treble limit data, where upper treble limit data can be used to characterize the target object's upper range ability, and lower treble limit data can be used to characterize the target object's lower range ability.

[0056] In a specific embodiment, the upper treble limit data and the lower treble limit data can be determined as follows:

[0057] S301: Collect the vocal signal from the target object's performance audio of a historical song, and convert the vocal signal into pitch to obtain the actual pitch data.

[0058] In one specific embodiment, the historical songs can be songs that the target object has sung before the song to be sung. Optionally, the historical songs can be determined from the target object's historical playlist.

[0059] Specifically, the fundamental frequency can be analyzed in the singing audio using a fundamental frequency detection algorithm to obtain the fundamental frequency at each singing moment, and the actual pitch value at each singing moment can be calculated based on the fundamental frequency to pitch conversion formula.

[0060] In a specific embodiment, the formula for converting fundamental frequency to pitch can be expressed as follows:

[0061] T = 12 × log2(F / 440) + 69, where F represents the fundamental frequency and T represents the actual pitch.

[0062] As an illustration, the fundamental frequency detection algorithm can employ the autocorrelation method and the Fourier transform method, and this application does not impose any particular limitations on this.

[0063] S302, compare the actual pitch data with the standard pitch data corresponding to historical songs to obtain pitch error distribution data.

[0064] In one specific embodiment, pitch error distribution data can characterize the distribution of the number of pitch errors of a target object within the pitch statistical range. Optionally, the pitch error distribution data can be represented in forms including, but not limited to, lists, histograms, etc.

[0065] In one specific embodiment, the standard pitch data corresponding to the historical songs can be extracted from the standard pitch file of the historical songs provided by the preset application. The standard pitch file can represent the standard pitch value corresponding to each singing moment of the original singer in the whole song.

[0066] In one specific embodiment, the actual pitch data may include: actual pitch values ​​corresponding to multiple singing moments, and the standard pitch file may include: standard pitch values ​​corresponding to multiple singing moments. The above comparison of the actual pitch data with the standard pitch data corresponding to historical songs to obtain pitch error distribution data may include:

[0067] S3021, traversing multiple performance moments;

[0068] S3022, If the difference between the actual pitch value and the standard pitch value at the current traversal singing moment meets the preset inaccuracy condition, increment the number of pitch errors corresponding to the standard pitch value at the current traversal singing moment by 1;

[0069] S3023 After the traversal is completed, the number of pitch errors within the pitch statistics range is statistically analyzed to obtain pitch error distribution data. The pitch statistics range is determined based on the standard pitch values ​​corresponding to multiple singing moments.

[0070] S303, based on pitch error distribution data, determines the high-pitched and low-pitched pitch error boundary data of the target object.

[0071] Specifically, by statistically analyzing the target's past pitch errors in singing, the boundaries of the target's low and high pitch inaccuracies are determined, thereby defining the target's range of high and low pitch capabilities.

[0072] To illustrate, let's take a histogram representing the distribution of pitch errors as an example. Figure 3a This application provides a histogram of pitch error distribution for high-pitched range misalignment of object A (i.e., user A) in an embodiment of the present application. Pitch intervals with pitch values ​​greater than the median of the pitch statistical range can be defined as high-pitched intervals. Figure 3aAs shown, the median of the pitch statistical range is 64, meaning the pitch range with a pitch value greater than 64 is the high-pitched range. The number of pitch errors in the high-pitched range [70, 73) is significantly higher than in other high-pitched ranges. Therefore, the left value of this range, 70, can be used as the high-pitched inaccuracy boundary data Tlimit_Up for object A. When the standard pitch value is higher than Tlimit_Up, the number of pitch errors increases significantly. In practical applications, to quantify the above description, the high-pitched statistical range within the pitch statistical range can be determined based on the median of the pitch range. The pitch value with the most pitch errors in the high-pitched statistical range can be used as the high-pitched pitch inaccuracy boundary Tlimit_Up for that object. Optionally, the difference between the number of errors at the high-pitched pitch inaccuracy boundary and the average number of errors in other high-pitched ranges can be limited to meet a preset difference condition to ensure that the number of errors at the high-pitched pitch inaccuracy boundary is significantly higher than the number of pitch errors in other high-pitched ranges. For example, the preset difference condition could be that the difference needs to be greater than 50% of the average number of errors in other high-pitched ranges. Figure 3a As shown, the number of pitch errors in each high-pitched interval ([64,67), [67,70), [70,73), [73,76), ≥76) within the high-pitched statistical range are 17, 13, 55, 24, and 1, respectively. The number of pitch errors (55) at [70,73) is the highest, and the average number of errors in the other high-pitched intervals is 14. The difference between 55 and 14 is greater than 50% of 14. Therefore, the left value 70 of the high-pitched interval [70,73) is taken as the high-pitched pitch error boundary Tlimit_Up.

[0073] Figure 3b This application provides a histogram of pitch error distribution for low-frequency range misalignment of object B (i.e., user B) in an embodiment of the application. Pitch ranges with pitch values ​​less than the median of the pitch statistical range can be defined as low-frequency ranges. Figure 3bAs shown, the median of the pitch statistical range is 64, meaning that pitch intervals with pitch values ​​less than 64 are considered low-pitched intervals. The low-pitched interval [55, 58) has a significantly higher number of pitch errors than other low-pitched intervals. Therefore, the rightmost value of this interval, 58, can be used as the low-pitched pitch inaccuracy boundary data Tlimit_Dw for object B. When the standard pitch value is lower than Tlimit_Up, the number of pitch errors increases significantly. In practical applications, to quantify the above description, the low-pitched statistical range within the pitch statistical range can be determined based on the median of the pitch intervals. The pitch value with the most pitch errors in the low-pitched statistical range can be used as the low-pitched pitch inaccuracy boundary Tlimit_Dw for that object. Optionally, the difference between the number of errors at the low-pitched pitch inaccuracy boundary and the average number of errors in other low-pitched intervals can be limited to meet a preset difference condition to ensure that the number of errors at the low-pitched pitch inaccuracy boundary is significantly higher than the number of pitch errors in other low-pitched intervals. For example, the preset difference condition could be that the difference needs to be greater than 50% of the average number of errors in other low-pitched intervals. Figure 3a As shown, the number of pitch errors in each bass interval (≤52, [52,55), [55,58), [58,61), [61,64)) within the bass statistical range are 0, 32, 65, 22, and 10, respectively. The number of pitch errors (65) at [55,58) is the highest, while the average number of errors in the other bass intervals is 16. The difference between 65 and 16 is greater than 50% of 16. Therefore, the right value 58 of the bass interval [55,58) is taken as the bass pitch inaccuracy boundary Tlimit_Dw.

[0074] Figure 3c This is a pitch error distribution histogram of object C (i.e. user C) with high and low pitch inaccuracies provided in an embodiment of this application. Based on the same method of analysis, the high pitch inaccuracy boundary Tlimit_Up of object C within the pitch statistical range is the left value 70 of the pitch interval [70, 73), while the low pitch inaccuracy boundary Tlimit_Dw is the right value 58 of the pitch interval [55, 58).

[0075] S304 uses the high-frequency misalignment boundary data as the upper limit of high frequencies and the low-frequency misalignment boundary data as the lower limit of low frequencies.

[0076] In one specific embodiment, rhythmic ability representation data can be used to characterize the rhythmic expression ability of a target object. Specifically, rhythmic ability representation data can be measured based on the maximum lyric density data that the target object can handle. Lyric density data can refer to the ratio of the number of lyrics to the duration of a song segment of a certain length. Optionally, the number of lyrics can refer to the number of syllables or words in the lyrics. For example, if a 10-second song segment has 20 syllables, then the lyric density is 20 / 10 = 2 syllables / second.

[0077] In one specific embodiment, the target object's mastery of the lyric density corresponding to the song segment can be measured based on the target object's performance data for that segment. Specifically, the performance data may include: a score given by a preset application to the target object's audio recording. Illustratively, the scoring algorithm used by the preset application may include, but is not limited to: pitch-based scoring algorithms, rhythm-based scoring algorithms, acoustic model-based scoring algorithms, deep learning-based scoring algorithms, comprehensive scoring algorithms, etc., and this application does not impose any particular limitations on this. In one specific embodiment, the maximum lyric density among multiple lyric densities where the target object's performance score is greater than a preset score threshold can be used as the target object's rhythmic ability representation data. For example, taking a preset score threshold of 85 points as an example, in the target object's historical song playlists, the maximum lyric density with a performance score higher than 85 points is 2.5 syllables / second, and the performance scores of other lyric densities higher than this rhythm value are all lower than 85 points. Therefore, the target object's rhythmic ability representation data Slimit is 2.5 syllables / second.

[0078] In one specific embodiment, breath control performance data can be used to characterize a target's ability to sing sustained long note segments. Specifically, breath control performance data can be measured based on the duration of stable note segments that the target can control, where a stable note segment can refer to a note segment with a continuously stable fundamental frequency, minimal inter-frame fundamental frequency differences, and a basically unchanged pitch value.

[0079] In one specific embodiment, the target's mastery of the duration of notes corresponding to stable note segments can be measured based on the target's performance data for stable note segments. Specifically, the performance data may include a pre-defined score given to the target's audio recordings by a preset application. In one specific embodiment, the longest duration of a note among multiple durations of notes with a performance score greater than a preset threshold can be used as the target's breath control ability representation data. For example, taking a preset threshold of 85 points as an example, if the longest duration of a stable note segment with a score higher than 85 points in the target's historical song list is 3 seconds, and the performance scores of other stable note segments with durations greater than 3 seconds are all lower than 85 points, then the target's breath control ability representation data Klimit is 3 seconds.

[0080] S202, determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data.

[0081] In one specific embodiment, parameter adaptation data can characterize the degree of adaptation between the audio parameters of the standard accompaniment audio and the singing ability characterization data of the target object.

[0082] In a specific embodiment, the aforementioned singing ability representation data may include: pitch accuracy representation data, rhythmic ability representation data, and breath control representation data; the aforementioned parameter adaptation data may include: pitch adaptation data, rhythmic adaptation data, and note duration adaptation data; and the aforementioned parameter adaptation data for determining the standard accompaniment audio relative to the singing ability representation data may include:

[0083] S2021, audio features are extracted from the standard accompaniment audio to obtain standard pitch data, lyric density data and note duration data of the standard accompaniment audio.

[0084] S2022, determine the pitch adaptation data between pitch accuracy representation data and standard pitch data.

[0085] In one specific embodiment, pitch adaptation data can characterize the degree of compatibility between the standard pitch data of the standard accompaniment audio and the pitch accuracy representation data of the target object. Specifically, the pitch accuracy representation data of the target object may include: upper treble limit data and / or lower treble limit data.

[0086] In an optional embodiment, when the target object has a high-pitched upper limit data (i.e., the target object's high-pitched expressive ability is limited), the highest value of the standard pitch data of the high-pitched segment in the standard accompaniment audio can be compared with the target object's high-pitched upper limit data. Specifically, the pitch matching data can be determined based on the pitch difference between the highest value of the standard pitch data and the high-pitched upper limit data. For example, the pitch difference Tdif = the highest value of the standard pitch data Tmax – the high-pitched upper limit data Tlimit_Up.

[0087] In an optional embodiment, if the target object has a lower bass limit (i.e., the target object's bass expression capability is limited), the lowest value of the standard pitch data of the bass segment in the standard accompaniment audio can be compared with the lower bass limit data of the target object. Specifically, the pitch matching data can be determined based on the pitch difference between the lower bass limit data and the lowest value of the standard pitch data. Schematic, the pitch difference Tdif = lower bass limit data Tlimit_Dw – lowest value of the standard pitch data Tmin.

[0088] In an optional embodiment, the pitch difference can be scaled down to obtain pitch-adapted data, thereby improving the learning efficiency and stability of the subsequent reinforcement learning model. Specifically, the pitch-adapted data can range from 0 to M1, mapping the pitch difference to the interval 0 to M1 using a linear transformation. M1 can be preset based on the learning efficiency and stability requirements of the reinforcement learning model in practical applications. For example, assuming M1 = 5, if the pitch difference ≤ 0, then the pitch-adapted data = 0; if 0 < pitch difference < 15, then the pitch-adapted data = [pitch difference / 3], where [] represents the rounding operation; if the pitch difference ≥ 15, then the pitch-adapted data = 5.

[0089] S2023, determine the rhythmic fit data between rhythmic ability representation data and lyric density data.

[0090] In one specific embodiment, rhythm adaptation data can characterize the degree of adaptation between the lyrics density data of the standard accompaniment audio and the rhythmic capability characterization data of the target object.

[0091] In one specific embodiment, the rhythmic capability representation data of the target object may include: the maximum lyric density data that the target object can handle; correspondingly, the rhythmic adaptation data may be determined based on the lyric density difference between the lyric density data of the standard accompaniment audio and the maximum lyric density data.

[0092] In an optional embodiment, the lyrics density difference can be scaled down to obtain rhythm-adapted data, thereby improving the learning efficiency and stability of the subsequent reinforcement learning model. Specifically, the value range of the rhythm-adapted data can be 0 to M2. The lyrics density difference is mapped to the interval 0 to M2 through a linear transformation. M2 can be preset based on the learning efficiency and stability requirements of the reinforcement learning model in practical applications. For example, assuming M2 = 5, if the lyrics density difference ≤ 0, then the rhythm-adapted data = 0; if 0 < 8 × lyrics density difference < 5, then the rhythm-adapted data = 8 × lyrics density difference; if 8 × lyrics density difference ≥ 5, then the rhythm-adapted data = 5.

[0093] S2024, determine the note duration adaptation data between breath ability representation data and note duration data.

[0094] In one specific embodiment, note duration adaptation data can characterize the degree of adaptation between the note duration data of the standard accompaniment audio and the breath control data of the target object.

[0095] In one specific embodiment, the breath control performance data of the target object may include: the longest note duration that the target object can handle. Accordingly, note duration adaptation data can be determined based on the note duration difference between the note duration data of the standard accompaniment audio and the longest note duration.

[0096] In an optional embodiment, the note duration difference can be scaled down to obtain note duration adaptation data, thereby improving the learning efficiency and stability of the subsequent reinforcement learning model. Specifically, the note duration adaptation data can range from 0 to M3, mapping the note duration difference to the interval 0 to M3 using a linear transformation. M3 can be preset based on the learning efficiency and stability requirements of the reinforcement learning model in practical applications. For example, assuming M3 = 5, if the note duration difference is ≤ 0, then the note duration adaptation data = 0; if 0 < 8 × note duration difference < 5, then the note duration adaptation data = 8 × note duration difference; if 8 × note duration difference ≥ 5, then the note duration adaptation data = 5.

[0097] As can be seen from the above embodiments, by analyzing the target subject's pitch expression ability, rhythm expression ability, and breath expression ability based on the target subject's historical singing performance, pitch ability representation data, rhythm ability representation data, and breath ability representation data are obtained. At the same time, audio features are extracted from the standard accompaniment audio of the song to be sung, obtaining standard pitch data, lyric density data, and note duration data of the standard accompaniment audio. Then, the pitch adaptation data between the pitch ability representation data and the standard pitch data, the rhythm adaptation data between the rhythm ability representation data and the lyric density data, and the note duration adaptation data between the breath ability representation data and the note duration data can be determined. This can improve the accuracy of parameter adaptation data representation, thereby improving the accuracy of subsequent accompaniment audio parameter adjustment.

[0098] S203, input the parameter adaptation data into the preset audio parameter tuning model, perform audio parameter tuning analysis on the standard accompaniment audio, and obtain audio parameter tuning data; the preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of historical songs, and the historical performance data of the target object singing historical songs.

[0099] In one specific embodiment, audio parameter adjustment data can be used to adjust the configuration parameters of the standard accompaniment audio. Specifically, the configuration parameters of the standard accompaniment audio may include, but are not limited to: pitch parameters, accompaniment playback speed parameters, lyrics display speed parameters, and note duration parameters.

[0100] In one specific embodiment, the audio parameter tuning data can be one of a plurality of preset audio parameter tuning data, and each preset audio parameter tuning data can be used to adjust at least one of the following parameters: pitch parameter, accompaniment playback speed parameter, lyrics display speed parameter, and note duration parameter.

[0101] In a specific embodiment, each preset audio parameter adjustment data can be a set of preset parameter adjustment operations. Specifically, a set of preset parameter adjustment operations can be a combination of multiple types of parameter adjustment operations. These multiple types of parameter adjustment operations can be divided into pitch adjustment operations, rhythm adjustment operations, and note duration adjustment operations. Specifically, multi-level pitch adjustment operations, multi-level rhythm adjustment operations, and multi-level note duration adjustment operations can be preset. The multi-level pitch adjustment operations, multi-level rhythm adjustment operations, and multi-level note duration adjustment operations are enumerated and combined to obtain multiple sets of preset parameter adjustment operations.

[0102] In a specific embodiment, the multi-level pitch adjustment operation may include at least one pitch-lowering operation or at least one pitch-uppering operation. Specifically, when the high-pitched segments of the song to be sung exceed the high-pitched expressive ability of the target (i.e., the highest value of the standard pitch data of the song to be sung is greater than the upper limit of the high pitch data of the target), the at least one pitch-lowering operation can be used to lower the pitch of notes in the standard accompaniment audio whose corresponding pitch values ​​exceed the upper limit of the high pitch data to different degrees. When the low-pitched segments of the song to be sung exceed the low-pitched expressive ability of the target (i.e., the lowest value of the standard pitch data of the song to be sung is less than the lower limit of the low pitch data of the target), the at least one pitch-uppering operation can be used to raise the pitch of notes in the standard accompaniment audio whose corresponding pitch values ​​are lower than the lower limit of the low pitch data to different degrees. Pitch adjustment can reduce the probability of the target singing out of tune, making singing easier. For illustration, the adjustment range of pitch lowering operation is 0 to (int)(1.2×(Tmax–Tlimit_Up)), where Tmax represents the highest value of the standard pitch data of the song to be sung, and Tlimit_Up represents the upper limit of the high pitch data of the target object. The adjustment range of pitch raising operation is -(int)(0.8×(Tlimit_Dw–Tmin)) to 0, where Tmin represents the lowest value of the standard pitch data of the song to be sung, and Tlimit_Dw represents the lower limit of the low pitch data of the target object. For example, if Tmax = 73 and Tlimit_Up = 70, then the adjustment range of pitch lowering operation is 0 to (int)(1.2×(73-70)) = 4, and the unit is semitone. If there are 4 levels of pitch lowering operation, it can be represented as pitch lowering operation 0 (keeping the original pitch, no change in pitch), pitch lowering operation 1 (lowering the pitch by 1 semitone), pitch lowering operation 2 (lowering the pitch by 2 semitones), and pitch lowering operation 3 (lowering the pitch by 3 semitones).

[0103] In one specific embodiment, when there is a segment in the standard accompaniment audio where the corresponding lyric density exceeds the target audience's rhythmic expression ability (i.e., the lyric density of a certain segment exceeds the maximum lyric density the target audience can handle), a multi-level rhythm adjustment operation can be used to adjust the singing rhythm of the standard accompaniment audio to different degrees. In one specific embodiment, the multi-level rhythm adjustment operation may include: a multi-level accompaniment slowdown operation and a multi-level lyric display line number adjustment operation. The multi-level accompaniment slowdown operation can be used to adjust the standard accompaniment audio to different playback speeds, and the multi-level lyric display line number adjustment operation can be used to adjust the lyrics to different display lines. It is understood that if the target audience has a weak rhythmic expression ability for a certain song, in addition to the accompaniment speed being too fast, it may also be because the target audience is not familiar enough with the lyrics, making it impossible to keep up with the singing speed. Therefore, in addition to the accompaniment slowdown operation, a lyric display line number adjustment operation can also be provided. By displaying more lyric information in advance, the target audience who is not familiar with the lyrics can prepare in advance, which helps the target audience keep up with the singing rhythm. For illustration purposes, if there are 5 levels of accompaniment speed reduction operation, they can be represented as accompaniment speed reduction operation 0 (no playback speed adjustment), accompaniment speed reduction operation 1 (adjust playback speed to 0.95x), accompaniment speed reduction operation 2 (adjust playback speed to 0.9x), accompaniment speed reduction operation 3 (adjust playback speed to 0.8x), and accompaniment speed reduction operation 4 (adjust playback speed to 0.7x). If there are 3 levels of lyrics display line number adjustment operation, they can be represented as lyrics line number operation 0 (display 1 line of lyrics, i.e., display the current line), lyrics line number operation 1 (display 2 lines of lyrics, i.e., display the current line and the next line of lyrics), and lyrics line number operation 2 (display 3 lines of lyrics, i.e., display the current line and the next two lines of lyrics).

[0104] In a specific embodiment, when there is a stable note segment in the standard accompaniment audio whose duration exceeds the target subject's breath control ability (i.e., the duration of the stable note segment is longer than the longest note duration the target subject can handle), multi-level note duration adjustment operations can be used to shorten the duration of the stable note segment to different degrees. Specifically, shortening the note duration can reduce the singing difficulty of the stable note segment and help improve the target subject's breath control. Illustratively, if there are 5 levels of note duration adjustment operations, they can be represented as duration shortening operation 0 (no note duration adjustment), duration shortening operation 1 (shortening the note duration to 0.95 times), duration shortening operation 2 (shortening the note duration to 0.9 times), duration shortening operation 3 (shortening the note duration to 0.8 times), and duration shortening operation 4 (shortening the note duration to 0.7 times).

[0105] In a specific embodiment, such as Figure 4 As shown, the above-mentioned preset audio parameter tuning model is trained in the following way:

[0106] S401, obtain historical parameter adaptation data, historical audio parameter tuning data, historical performance data, and decision reward representation information corresponding to the preset reinforcement learning model for the historical songs. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data.

[0107] In one specific embodiment, the historical song can be a song that the target object has sung before the song to be sung. The historical parameter adaptation data can be the degree of adaptation between the audio parameters of the standard accompaniment audio corresponding to the historical song and the historical singing ability representation data of the target object. The historical singing ability representation data of the target object can be obtained by analyzing the singing ability based on the historical singing data of the target object before singing the historical song. The singing ability analysis method and the method for determining the historical parameter adaptation data can be found in the detailed content of step S201, and will not be repeated here.

[0108] In one specific embodiment, historical audio parameter tuning data can be parameter tuning data obtained by performing audio parameter tuning analysis on the standard accompaniment audio of a historical song based on historical parameter adaptation data. In another specific embodiment, historical parameter adaptation data corresponding to a historical song can be input into a preset reinforcement learning model. The preset reinforcement learning model then uses the historical parameter adaptation data to determine a preset audio parameter tuning data from multiple preset audio parameter tuning data sets, which is then used as the historical audio parameter tuning data.

[0109] In a specific embodiment, the decision reward representation information corresponding to the preset reinforcement learning model is used to represent the long-term reward data obtained by the preset reinforcement learning model from different preset audio parameter tuning data under different parameter adaptation data. The long-term reward data is used to represent the feedback effect of the parameter tuning decisions accumulated since the audio parameter tuning training of the preset reinforcement learning model began. Specifically, the decision reward representation information can be represented by a decision reward table, the size of which can be E×F, where E represents the number of possible values ​​within the range of values ​​corresponding to the parameter adaptation data, and F represents the number of multiple preset audio parameter tuning data (i.e., the number of multiple sets of preset parameter tuning operations). Each element Q(Si, Aj) in the decision reward table is used to represent the long-term feedback effect obtained by the preset reinforcement learning model from deciding on preset audio parameter tuning data Ai (performing a corresponding set of preset parameter tuning operations) under parameter adaptation data Si, i = 1, ..., E, j = 1, ..., F.

[0110] In a specific embodiment, such as Figure 5 As shown, the above historical audio parameter tuning data was obtained through the following decision-making process:

[0111] S501, Obtain parameter tuning method indication information. The parameter tuning method indication information is used to indicate that the probability determined by the historical audio parameter tuning data through reinforcement learning is the first probability, and the probability determined by random selection is the second probability.

[0112] Specifically, the parameter tuning decisions for the pre-defined reinforcement learning model are primarily based on the magnitude of the long-term reward data corresponding to each pre-defined audio parameter tuning data under the historical parameter adaptation data in the decision reward representation information. However, in the early stages of training the pre-defined reinforcement learning model, the training samples are few, and the obtained long-term reward data may be inaccurate. If the decision reward representation information is used as the sole basis for decision-making, it is easy to cause decision errors. In addition, after the pre-defined reinforcement learning model has been trained for a period of time, if the decision reward representation information is used as the sole basis for decision-making, the model may have relatively fixed parameter tuning decisions and be unable to effectively explore the target object (the user performing the song). Therefore, the epsilon-greedy algorithm can be considered. That is, during the training process of the pre-defined reinforcement learning model, in each round of parameter tuning decision-making, a first probability is used to randomly select from multiple pre-defined audio parameter tuning data, and a second probability is used to select the pre-defined audio parameter tuning data with the best long-term reward data under the parameter adaptation data corresponding to the current round according to the reinforcement learning method.

[0113] Optionally, the first probability should gradually decrease as the number of training rounds increases, while the second probability should gradually increase as the number of training rounds increases.

[0114] S502, based on the parameter tuning method instruction information and the decision reward representation information corresponding to the preset reinforcement learning model, determine the historical audio parameter tuning data from a variety of preset audio parameter tuning data under the historical parameter adaptation data.

[0115] Specifically, based on the parameter tuning method instructions, when it is determined that the historical audio parameter tuning data was determined through reinforcement learning, the preset audio parameter tuning data with the best long-term reward data under the historical parameter fitting data in the current decision reward representation information is used as the historical audio parameter tuning data corresponding to the historical parameter fitting data; when it is determined that the historical audio parameter tuning data was determined through random selection, a preset audio parameter tuning data is randomly selected from multiple preset audio parameter tuning data under the historical parameter fitting data as the historical audio parameter tuning data corresponding to the historical parameter fitting data.

[0116] As can be seen from the above embodiments, in the training process of the preset reinforcement learning model, in each round of parameter tuning decision, the preset audio parameter tuning data with the best long-term reward data under the parameter adaptation data corresponding to the current round is randomly selected with a first probability and with a second probability according to the reinforcement learning method. This can avoid the model parameter tuning decision being relatively fixed and improve the training effect of the reinforcement learning model.

[0117] In one specific embodiment, historical performance data can be used to evaluate the target object's actual performance on a historical song. Specifically, historical performance data can be represented as historical performance scores.

[0118] In a specific embodiment, such as Figure 6 As shown, historical performance data can be determined in the following ways:

[0119] S601, Obtain the historical singing audio of the target object based on historical audio parameter adjustment data.

[0120] Specifically, parameters can be configured for the standard accompaniment audio corresponding to a historical song based on historical audio parameter tuning data to obtain the target accompaniment audio corresponding to the historical song. During the process of the target object singing the song based on the target accompaniment audio corresponding to the historical song, the historical singing audio can be collected.

[0121] S602 analyzes singing performance based on historical singing audio to obtain historical parameter-adjusted singing performance data.

[0122] In one specific embodiment, historical performance data can be used to evaluate the performance of a target subject singing a song based on the target accompaniment audio (the accompaniment audio after parameter tuning) corresponding to a historical song. Specifically, historical performance data can include: a score given by a preset application for the historical performance audio. Illustratively, the scoring algorithm used by the preset application may include, but is not limited to: pitch-based scoring algorithms, rhythm-based scoring algorithms, acoustic model-based scoring algorithms, deep learning-based scoring algorithms, comprehensive scoring algorithms, etc., and this application does not impose any particular limitations on these.

[0123] S603, based on historical audio parameter tuning data, determines the historical parameter deviation data between the target accompaniment audio corresponding to a historical song and the standard accompaniment audio corresponding to a historical song.

[0124] In one specific embodiment, historical parameter deviation data can be used to characterize the difference in audio configuration parameters between the target accompaniment audio corresponding to the historical song and the standard accompaniment audio corresponding to the historical song, that is, the difference in configuration parameters before and after audio adjustment of the historical song.

[0125] In a specific embodiment, historical audio parameter adjustment data may include: historical pitch adjustment data, historical rhythm adjustment data, and historical note duration adjustment data. Specifically, historical pitch adjustment data may represent pitch adjustment operations for a historical song, historical rhythm adjustment data may represent rhythm adjustment operations for a historical song, and historical note duration adjustment data may represent note duration adjustment operations for a historical song. Correspondingly, based on the historical audio parameter adjustment data, the historical parameter deviation data between the adjusted accompaniment audio corresponding to the historical song and the standard accompaniment audio corresponding to the historical song is determined.

[0126] S6031 integrates historical pitch adjustment data, historical rhythm adjustment data, and historical note duration adjustment data to obtain historical tuning parameter scores.

[0127] S6032 uses the ratio between the historical parameter score and the total number of notes in the historical song as the historical parameter deviation data.

[0128] To illustrate, suppose a historical song has 200 notes. Notes 3-10 undergo rhythm adjustment (level 4); notes 80-90 undergo duration adjustment (level 3); and notes 90-100 undergo pitch adjustment (level 2).

[0129] Historical parameter adjustment score = (10-3+1)×4+(90-80+1)×3+(100-90+1)×2=87;

[0130] Historical parameter deviation data = 100 × 87 / 200 = 0.485.

[0131] In an optional embodiment, the historical parameter deviation data can also be scaled up using a preset magnification factor. The preset magnification factor can be set according to the parameter deviation magnification accuracy in actual applications. For example, the preset magnification factor can be 100, which can correspondingly magnify the historical parameter deviation data of 0.485 to 48.5.

[0132] S604 calibrates historical parameter deviation data on historical performance data to obtain historical performance data.

[0133] In a specific embodiment, the calibration algorithm corresponding to historical performance data can be expressed as the following formula:

[0134] Score2 = a1 × Score 1 - a2 × SoundDeviation, where Score2 represents historical performance data, Score1 represents historical performance data with adjusted parameters, SoundDeviation represents historical parameter deviation data, and a1 and a2 are two weighting coefficients, with a1 being greater than a2.

[0135] Specifically, the weighting coefficients a1 and a2 can be preset according to the accuracy requirements of the singing performance calibration in actual applications. For illustration, a1 can be 1.0 and a2 can be 0.2.

[0136] Specifically, the difference between the weighted result of historical parameter-adjusted singing performance data and the weighted result of historical parameter deviation data is used as historical singing performance data. Based on the singing evaluation obtained by the target subject using parameter-adjusted accompaniment audio to sing a song, the performance bonus brought about by the reduced singing difficulty after parameter adjustment of the accompaniment audio is removed, and the actual singing performance evaluation corresponding to the actual singing performance ability of the target subject is obtained.

[0137] As can be seen from the above embodiments, the parameters of the standard accompaniment audio corresponding to the historical songs are configured based on the historical audio parameter tuning data of the historical songs. After the target object sings the song based on the configured target accompaniment audio, the singing performance is analyzed to obtain historical parameter tuning singing performance data. Based on the historical audio parameter tuning data, the audio configuration parameter differences between the target accompaniment audio and the standard accompaniment audio corresponding to the historical songs are determined. Then, based on the differences in audio configuration parameters, the historical parameter tuning singing performance data is calibrated. This can remove the performance bonus brought about by the reduction in singing difficulty after accompaniment audio parameter tuning, and obtain the actual singing performance evaluation corresponding to the actual singing performance ability of the target object. Subsequently, the decision reward representation information of the model is updated based on the actual singing performance evaluation. This can help the target object gradually reduce the intensity of accompaniment audio parameter tuning and improve the actual singing performance of the target object.

[0138] S402, based on historical performance data, determine the instant reward data obtained from historical audio parameter tuning data under historical parameter adaptation data, and the magnitude of the instant reward data is negatively correlated with the magnitude of the historical performance data.

[0139] In one specific embodiment, immediate reward data can be used to evaluate the immediate feedback effect of historical audio hyperparameter tuning data (i.e., performing historical audio hyperparameter tuning operations). Specifically, immediate reward is an important feedback mechanism in the reinforcement learning process, helping the reinforcement learning model evaluate the immediate effect of its behavior and thus adjust its strategy to obtain more rewards.

[0140] Specifically, in the reinforcement learning-based audio processing scenario involved in this application embodiment, the purpose of the reward is to enable the target object to gradually match the standard singing method or standard singing difficulty of the song after executing the optimal parameter tuning operation decided by the preset reinforcement learning model, so that the target object can sing the song smoothly. Therefore, after the preset reinforcement learning model decides the corresponding audio parameter tuning data based on the input parameter adaptation data, if the value of the target object's calibrated singing performance data is higher (indicating that the target object's actual singing performance is better), then the corresponding value of the immediate reward data should be smaller, thereby reducing the adjustment intensity of the accompaniment audio under the parameter adaptation data; if the value of the target object's calibrated singing performance data is lower (indicating that the target object's actual singing performance is worse), then the corresponding value of the immediate reward data should be larger, thereby increasing the adjustment intensity of the accompaniment audio under the parameter adaptation data.

[0141] Specifically, the exact value of the instant reward data is determined by a preset reward function. The preset reward function represents the instant reward obtained by making a decision on a certain preset audio parameter tuning data (i.e., performing a certain set of preset parameter tuning operations) under a certain parameter adaptation data. Schematic, the preset reward function can be expressed as the following formula:

[0142] R=func(Score2)=func(a1×Score 1-a2×SoundDeviation)

[0143] Here, Score2 represents the historical performance data (i.e., the performance data after refining the historical performance data Score1 based on the historical parameter deviation data SoundDeviation), and func() represents a monotonically decreasing function. When Score2 is less than the preset minimum score, func outputs the maximum reward value, and when Score2 is greater than the preset maximum score, func outputs the minimum reward value. The preset minimum score is less than the preset maximum score, and the maximum reward value is greater than the minimum reward value.

[0144] S403, based on real-time reward data, updates the decision reward representation information corresponding to the preset reinforcement learning model to obtain the preset audio parameter tuning model.

[0145] In a specific embodiment, the above-mentioned updating of the decision reward representation information corresponding to the preset reinforcement learning model based on real-time reward data, resulting in a preset audio parameter tuning model, can be expressed as the following formula:

[0146]

[0147] Where Q(S,A) represents the decision reward representation information corresponding to the preset audio parameter tuning model, Si represents the i-th parameter adaptation data in E parameter adaptation data, Aj represents the j-th preset audio parameter tuning data in F preset audio parameter tuning data, and S 输入 A represents the historical parameter adaptation data of the input preset reinforcement learning model. 输出 This indicates that the pre-defined reinforcement learning model is based on S. 输入 The historical audio parameter tuning data obtained from the decision, R represents the S 输入 Make decision A 输出 The obtained instant reward data, where α is the learning rate.

[0148] Specifically, the learning rate can be used to adjust the sensitivity of a preset reinforcement learning model to the learning effect of the target object's feedback when updating the decision reward representation information. In an optional embodiment, the learning rate ranges from 0 to 1; illustratively, the learning rate can be 0.1.

[0149] It can be understood that the parameter tuning decision of the preset reinforcement learning model is mainly based on the magnitude of the long-term reward data corresponding to each preset audio parameter tuning data under the historical parameter adaptation data in the decision reward representation information. Therefore, the decision reward representation information can represent the audio parameter tuning strategy of the preset reinforcement learning model. By updating the decision reward representation information, the preset reinforcement learning model can be updated to obtain the preset audio parameter tuning model.

[0150] In a specific embodiment, the aforementioned preset reinforcement learning model can be obtained by performing t rounds of audio parameter tuning training based on the sample data collected before the current target object sings the historical song. The historical parameter adaptation data, historical audio parameter tuning data, and historical singing performance data corresponding to the historical song are used as the sample data for the (t+1)th round of training. The preset reinforcement learning model is then trained for the (t+1)th round of audio parameter tuning to obtain the preset audio parameter tuning model.

[0151] Indicative, Figure 7 This is a schematic diagram of a framework for an audio parameter tuning and training process based on reinforcement learning, as provided in an embodiment of this application. Specifically, as shown... Figure 7As shown, reinforcement learning consists of two main parts: a pre-set reinforcement learning model and a target object (i.e., the learning environment). The model pre-sets the range of parameter adaptation data, multiple pre-set audio parameter tuning data sets, and a decision reward table. The size of the decision reward table can be E×F (where E represents the number of possible values ​​within the range of parameter adaptation data, and F represents the number of pre-set audio parameter tuning data sets). The decision reward table is initialized with a default initial value of 0. During multiple rounds of audio parameter tuning training, the pre-set reinforcement learning model interacts with the learning environment and continuously maintains and updates the decision reward table. Taking the t-th round of audio parameter tuning training as an example, the training process of the pre-set reinforcement learning model can include:

[0152] S701, the preset reinforcement learning model obtains parameter adaptation data (S) of the target object for the standard accompaniment audio of the current song. t ) and the decision reward table (Q) for the t-th training round t ).

[0153] S702, the preset reinforcement learning model is based on the adaptation data (S t ) and decision-making reward table (Q t Make audio parameter tuning decisions for the current song and output audio parameter tuning data (A). t ).

[0154] S703, based on audio parameter tuning data (A t The parameters of the standard accompaniment audio of the current song are adjusted to obtain the adjusted accompaniment audio. After the target object sings the current song based on the adjusted accompaniment audio, the adjusted singing performance data and singing ability characterization data of the target object are obtained. Based on the audio parameter deviation data between the adjusted accompaniment audio and the standard accompaniment audio, the adjusted singing performance data is calibrated to obtain the singing performance data of the target object.

[0155] S704, Based on the target object's singing performance data, determine the instant reward data (R) obtained from the decision audio parameter tuning data. t+1 ), and based on instant reward data (R t+1 ) Decision reward table (Q) for the t-th training round t Update the data to obtain the decision reward table (Q) for the (t+1)th training round. t+1 ).

[0156] S705, based on the target object's singing ability representation data, determine the parameter adaptation data (S) for the target object for the standard accompaniment audio of the next song. t+1 ).

[0157] As can be seen from the above embodiments, based on historical performance data, the instant reward data obtained by making decisions on historical audio parameter tuning data under historical parameter adaptation data is determined. The magnitude of the instant reward data is negatively correlated with the magnitude of the historical performance data. Based on the instant reward data, the decision reward representation information corresponding to the preset reinforcement learning model is updated to obtain the preset audio parameter tuning model. By establishing a scientific reinforcement learning evaluation mechanism, the parameter tuning strategy of the model is used to improve the actual performance of the target object while helping the target object gradually reduce the intensity of accompaniment audio parameter tuning.

[0158] In a specific embodiment, the above-mentioned input of parameter adaptation data into a preset audio parameter tuning model to perform audio parameter tuning analysis on standard accompaniment audio, and the resulting audio parameter tuning data may include:

[0159] S2031, Input the parameter adaptation data into the preset audio parameter tuning model. Based on the decision reward representation information corresponding to the preset audio parameter tuning model, determine the target audio parameter tuning data with the largest long-term reward data from a variety of preset audio parameter tuning data under the parameter adaptation data. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data. The long-term reward data is used to represent the feedback effect of the parameter tuning decisions accumulated since the audio parameter tuning training of the preset reinforcement learning model.

[0160] Specifically, the parameter tuning decision for the preset audio parameter tuning model is based on the magnitude of the long-term reward data corresponding to each preset audio parameter tuning data under the corresponding parameter adaptation data in the decision reward representation information. Each time, the preset audio parameter tuning data with the largest long-term reward data under the current input parameter adaptation data is selected as the target audio parameter tuning data. It can be understood that since the preset audio parameter tuning model is trained from a preset reinforcement learning model, the decision reward representation information used by the preset audio parameter tuning model is the decision reward representation information updated by the preset reinforcement learning model (from the initial audio parameter tuning training to the point where the preset audio parameter tuning model is obtained). Correspondingly, the long-term reward data corresponding to the preset audio parameter tuning model is used to represent the feedback effect of the cumulative parameter tuning decisions of the preset reinforcement learning model (from the initial audio parameter tuning training to the point where the preset audio parameter tuning model is obtained).

[0161] S2032, Determine audio parameter tuning data based on target audio parameter tuning data.

[0162] Specifically, the target audio parameter tuning data is used as the audio parameter tuning data.

[0163] As can be seen from the above embodiments, the parameter adaptation data is input into the preset audio parameter tuning model, and the target audio parameter tuning data with the largest corresponding long-term reward data is determined from the various preset audio parameter tuning data under the parameter adaptation data. This target audio parameter tuning data is used as the audio parameter tuning data of the song to be sung. By adaptively adjusting the audio parameters, the singing difficulty of the song to be sung can be closer to the current singing level of the target.

[0164] In a specific embodiment, the parameter adaptation data may include: pitch adaptation data, rhythm adaptation data, and note duration adaptation data. Correspondingly, the preset audio parameter tuning model may include: a pitch adjustment model, a rhythm adjustment model, and a note duration adjustment model. The pitch adjustment model can be used to determine the pitch adjustment data corresponding to the pitch adaptation data, the rhythm adjustment model can be used to determine the rhythm adjustment data corresponding to the rhythm adaptation data, and the note duration adjustment model can be used to determine the note duration adjustment data corresponding to the note duration adaptation data. The decision results of the above three models are combined to obtain the audio parameter tuning data. Specifically, as shown... Figure 8 As shown, the parameter adaptation data is input into the preset audio tuning model, and audio tuning analysis is performed on the standard accompaniment audio to obtain the audio tuning data, including:

[0165] S801 inputs pitch adaptation data into the pitch adjustment model, performs pitch adjustment analysis on the standard accompaniment audio, and obtains the pitch adjustment data corresponding to the pitch adaptation data.

[0166] In a specific embodiment, the above-mentioned inputting pitch adaptation data into the pitch adjustment model and performing pitch adjustment analysis on the standard accompaniment audio to obtain pitch adjustment data corresponding to the pitch adaptation data may include:

[0167] S8011, input the pitch adaptation data into the pitch adjustment model, and based on the first decision reward representation information corresponding to the pitch adjustment model, determine the target pitch adjustment data with the largest long-term reward data from a variety of preset pitch adjustment data under the pitch adaptation data. The first decision reward representation information is used to represent the long-term reward data obtained by deciding different preset audio tuning parameter data under different parameter adaptation data.

[0168] S8012 uses the target pitch adjustment data as the pitch adjustment data.

[0169] In a specific embodiment, the pitch adjustment model is obtained by real-time pitch adjustment training of the first reinforcement learning model based on the historical pitch adaptation data corresponding to the historical songs of the target object, the historical pitch adjustment data corresponding to the historical pitch adaptation data, and the historical singing performance data of the target object for the historical songs. Specifically, the training process of the pitch adjustment model is similar to the training process of the above-mentioned preset audio parameter adjustment model. For details, please refer to the detailed content of steps S401 to S403, which will not be repeated here.

[0170] S802 inputs the rhythm adaptation data into the rhythm adjustment model, performs rhythm adjustment analysis on the standard accompaniment audio, and obtains the rhythm adjustment data corresponding to the rhythm adaptation data.

[0171] S8021, input the rhythm adaptation data into the rhythm adjustment model, and based on the second decision reward representation information corresponding to the rhythm adjustment model, determine the target rhythm adjustment data with the largest long-term reward data from a variety of preset rhythm adjustment data under the rhythm adaptation data. The second decision reward representation information is used to represent the long-term reward data obtained by deciding different preset rhythm adjustment data under different rhythm adaptation data.

[0172] S8022, use the target rhythm adjustment data as the rhythm adjustment data.

[0173] In a specific embodiment, the rhythm adjustment model is obtained by real-time rhythm adjustment training of the second reinforcement learning model based on the historical rhythm adaptation data corresponding to the historical songs of the target object, the historical rhythm adjustment data corresponding to the historical rhythm adaptation data, and the historical performance data of the target object for the historical songs. Specifically, the training process of the rhythm adjustment model is similar to the training process of the above-mentioned preset audio parameter adjustment model. For details, please refer to the detailed content of steps S401 to S403, which will not be repeated here.

[0174] S803 inputs the note duration adaptation data into the note duration adjustment model, performs note duration adjustment analysis on the standard accompaniment audio, and obtains the note duration adjustment data corresponding to the note duration adaptation data.

[0175] S8031, input the note duration adaptation data into the note duration adjustment model, and based on the third decision reward representation information corresponding to the note duration adjustment model, determine the target duration adjustment data with the largest long-term reward data from a variety of preset duration adjustment data under the note duration adaptation data. The third decision reward representation information is used to represent the long-term reward data obtained by deciding different preset duration adjustment parameter data under different note duration adaptation data.

[0176] S8032 uses the target duration adjustment data as the note duration adjustment data.

[0177] In a specific embodiment, the note duration adjustment model is obtained by training the third reinforcement learning model in real time with note duration adjustment based on the historical note duration adaptation data corresponding to the historical songs of the target object, the historical note duration adjustment data corresponding to the historical note duration adaptation data, and the historical performance data of the target object for the historical songs. Specifically, the training process of the note duration adjustment model is similar to the training process of the above-mentioned preset audio parameter adjustment model. For details, please refer to the detailed content of steps S401 to S403, which will not be repeated here.

[0178] S804 obtains audio parameter adjustment data based on pitch adjustment data, rhythm adjustment data, and note duration adjustment data.

[0179] Specifically, the pitch adjustment data, rhythm adjustment data, and note duration adjustment data are combined to obtain audio parameter adjustment data.

[0180] As can be seen from the above embodiments, by using the pitch adjustment model, rhythm adjustment model, and note duration adjustment model to determine the pitch adjustment data corresponding to the pitch adaptation data, the rhythm adjustment data corresponding to the rhythm adaptation data, and the note duration adjustment data corresponding to the note duration adaptation data, and combining the decision results of the three models to obtain audio parameter adjustment data, the accuracy of accompaniment audio parameter adjustment can be improved on the basis of improving the accuracy of pitch adjustment, rhythm adjustment, and note duration adjustment.

[0181] S204, based on audio parameter tuning data, configures the parameters of the standard accompaniment audio to obtain the target accompaniment audio corresponding to the song to be sung.

[0182] Specifically, the target accompaniment audio is a parameter-tuned accompaniment audio obtained by configuring the parameters of the standard accompaniment audio based on audio parameter tuning data. In a specific embodiment, the audio parameter tuning data can be a set of preset parameter tuning operations. Performing this set of preset parameter tuning operations on the standard accompaniment audio yields the target accompaniment audio.

[0183] As can be seen from the above embodiments, based on historical parameter adaptation data of historical songs relative to the historical singing ability of the target object, historical audio parameter tuning data of historical songs, and historical singing performance data of the target object singing historical songs, a preset audio parameter tuning model is trained on a preset reinforcement learning model to obtain a preset audio parameter tuning model. This allows the model to determine the accompaniment audio adjustment parameters that are adapted to the singing level of the target object. Then, the parameter adaptation data of the standard accompaniment audio of the song to be sung by the target object relative to the singing ability representation data of the target object is determined, and the parameter adaptation data is input into the preset audio parameter tuning model to obtain audio parameter tuning data. Based on the audio parameter tuning data, the parameters of the standard accompaniment audio are configured to obtain the target accompaniment audio of the song to be sung. Through adaptive audio parameter adjustment, the singing difficulty of the song to be sung can be made closer to the current singing level of the target object, so as to improve the singing ability of the target object in the long-term singing process, thereby improving the singing experience of the target object.

[0184] In a specific embodiment, such as Figure 9a As shown, after configuring the parameters of the standard accompaniment audio based on the audio parameter tuning data to obtain the target accompaniment audio corresponding to the song to be sung, the above method may further include:

[0185] S205: During the process of the target object singing the song to be sung based on the target accompaniment audio, the current singing audio is collected.

[0186] S206: Analyze the singing performance based on the current singing audio to obtain parameter-adjusted singing performance data.

[0187] In one specific embodiment, the parametric-tuned performance data can be used to evaluate the performance of a target subject singing a song based on the parametrically tuned target accompaniment audio. Specifically, the parametric-tuned performance data may include: a performance score from a preset application for the current performance audio. Illustratively, the performance scoring algorithm used by the preset application may include, but is not limited to: pitch-based scoring algorithms, rhythm-based scoring algorithms, acoustic model-based scoring algorithms, deep learning-based scoring algorithms, comprehensive scoring algorithms, etc., and this application does not impose any particular limitations on these.

[0188] S207, based on audio parameter tuning data, determines the audio parameter deviation data between the target accompaniment audio and the standard accompaniment audio.

[0189] In a specific embodiment, audio parameter deviation data can be used to characterize the difference in audio configuration parameters between the target accompaniment audio and the standard accompaniment audio corresponding to the song to be sung, that is, the difference in configuration parameters before and after audio adjustment of the song to be sung. Specifically, the detailed content of "determining the audio parameter deviation data between the target accompaniment audio and the standard accompaniment audio based on audio parameter adjustment data" in step S207 is similar to the detailed content of "determining the historical parameter deviation data between the target accompaniment audio and the standard accompaniment audio corresponding to a historical song based on historical audio parameter adjustment data" in step S4013, and will not be repeated here.

[0190] S208 calibrates the performance data based on audio parameter deviation data to obtain the current performance data.

[0191] In one specific embodiment, the current performance data can be used to evaluate the target object's actual performance on the song to be sung. Specifically, the current performance data can be represented as a current performance score.

[0192] Specifically, the detailed content of "calibrating the performance data based on the audio parameter deviation data to obtain the current performance data" in step S208 is similar to the detailed content of "calibrating the historical performance data based on the historical parameter deviation data to obtain the historical performance data" in step S4014, and will not be repeated here.

[0193] S209, based on parameter adaptation data, audio parameter tuning data and current singing performance data, performs real-time audio parameter tuning training on the preset reinforcement learning model.

[0194] Specifically, the detailed content of "conducting real-time audio parameter tuning training on the preset reinforcement learning model based on parameter adaptation data, audio parameter tuning data and current singing performance data" in step S209 is similar to the detailed content of steps S402 to S403, and will not be repeated here.

[0195] As can be seen from the above embodiments, the accompaniment audio parameters of the song to be sung by the target object are adjusted based on the reinforcement learning method, and a scientific reinforcement learning evaluation mechanism is established. This makes the difficulty of the adjusted song more closely match the target object's own singing ability, and gradually improves the singing ability under the guidance of reinforcement learning.

[0196] In a specific embodiment, such as Figure 9b As shown, the above method may further include:

[0197] S901, obtain the song style data of the song to be sung.

[0198] Specifically, song style data can be used to characterize the musical style of the song to be sung. In a specific embodiment, song style data can be represented as a style tag. For example, style tags can include, but are not limited to, pop, rock, folk, etc.

[0199] S902 inputs the song style data into the preset sound effect parameter tuning model, performs sound effect parameter tuning analysis on the standard accompaniment audio, and obtains sound effect parameter tuning data.

[0200] In one specific embodiment, the audio effect tuning data can be used to adjust the audio effect parameters of the standard accompaniment audio. Specifically, the audio effect parameters of the standard accompaniment audio may include, but are not limited to: reverb parameters, compression parameters, and equalization parameters. Reverb simulates the sound reflection effects of different rooms and environments to create a realistic ambient sound effect; compression adjusts the dynamic range of the audio signal to make the intensity of the audio signal more balanced, thereby enhancing the clarity and audibility of the song; equalization adjusts the frequency response of the audio signal to make the timbre of the audio signal more balanced and natural, thereby enhancing the musicality and artistry of the song.

[0201] In one specific embodiment, the audio effect tuning data can be one of a plurality of preset audio effect tuning data, and each preset audio effect tuning data can be used to adjust at least one of the reverberation parameter, compression parameter and equalization parameter.

[0202] In a specific embodiment, the above-mentioned input of song style data into a preset sound effect parameter tuning model, and the sound effect parameter tuning analysis of standard accompaniment audio, to obtain sound effect parameter tuning data may include:

[0203] S9021, input the song style data into the preset sound effect parameter tuning model, and based on the decision reward representation information corresponding to the preset sound effect parameter tuning model, determine the target sound effect parameter tuning data with the largest long-term reward data from multiple preset sound effect parameter tuning data under the song style data. The decision reward representation information corresponding to the preset sound effect parameter tuning model is used to represent the long-term reward data obtained by making decisions on different preset sound effect parameter tuning data under different song style data.

[0204] Specifically, the parameter tuning decision corresponding to the preset sound effect tuning model is based on the magnitude of the long-term reward data corresponding to each preset sound effect tuning data under the corresponding song style data in the decision reward representation information. Each time, the preset sound effect tuning data with the largest long-term reward data under the currently input song style data is selected as the target sound effect tuning data.

[0205] S9022 uses the target audio parameter tuning data as the audio parameter tuning data.

[0206] In a specific embodiment, the preset sound effect parameter tuning model can be trained in the following way:

[0207] S1, obtain historical song style data, historical sound effect parameter tuning data, and historical performance data corresponding to historical sound effect parameter tuning data; the historical sound effect parameter tuning data is obtained by making sound effect parameter tuning decisions on historical song style data based on the fourth decision reward representation information corresponding to the fourth reinforcement learning model; the fourth decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset sound effect parameter tuning data under different song style data.

[0208] Specifically, sound effects can be configured on the standard accompaniment audio corresponding to a historical song based on historical sound effect parameter tuning data to obtain the target sound effect audio for the historical song. During the performance of the song by the target subject based on the target sound effect audio, the song's performance audio is collected, and performance performance is analyzed based on the performance audio to obtain historical performance scoring data. For illustrative purposes, the performance performance analysis here can use a performance scoring algorithm; this application does not impose any specific limitations on this.

[0209] S2, based on historical singing score data, determines the real-time sound effect reward data obtained from historical sound effect parameter adjustment data under historical song style data. The value of the real-time sound effect reward data is positively correlated with the value of the historical singing score data.

[0210] Specifically, real-time sound effect reward data can be used to evaluate the real-time performance feedback effect brought about by decision-making historical sound effect parameter tuning data (i.e., performing historical sound effect parameter tuning operations).

[0211] S3, based on real-time sound effect reward data, updates the fourth decision reward representation information corresponding to the fourth reinforcement learning model to obtain the preset sound effect parameter tuning model.

[0212] Specifically, the detailed content of "updating the fourth decision reward representation information corresponding to the fourth reinforcement learning model based on the real-time sound effect reward data to obtain the preset sound effect parameter tuning model" is similar to the detailed content of "updating the decision reward representation information corresponding to the preset reinforcement learning model based on the real-time reward data to obtain the preset audio parameter tuning model" in step S403, and will not be repeated here.

[0213] S903, based on audio parameter tuning data and sound effect parameter tuning data, configures the parameters of the standard accompaniment audio to obtain the target accompaniment audio corresponding to the song to be sung.

[0214] As can be seen from the above embodiments, different sound effect parameters can be adaptively configured through reinforcement learning and the current song style tag, making the target's singing effect more immersive, the sound clearer and more infectious, thereby further enhancing the target's singing experience.

[0215] See Figure 10a , Figure 10a This is a schematic diagram of an audio parameter adjustment scheme for a karaoke scenario provided in an embodiment of this application. Specifically, firstly, the standard accompaniment audio of the current song of the karaoke user is configured based on initial configuration parameters to obtain the initial accompaniment audio. These initial configuration parameters can be the default audio parameters of the online karaoke software when the karaoke user sings for the first time, or they can be the configuration parameters obtained from the audio parameter tuning model trained based on reinforcement learning methods for the karaoke user's most recent audio parameter tuning operation. The karaoke user can then adjust the parameters as follows: Figure 10b The karaoke singing interface shown demonstrates singing the current song based on the initial accompaniment audio. After the user finishes singing, the karaoke scoring system analyzes the user's actual performance, obtaining a karaoke score and multi-dimensional quantitative results of singing ability (including: quantitative results of pitch expression ability, quantitative results of rhythm expression ability, and quantitative results of breath expression ability). The karaoke score can be displayed as follows: Figure 10cThe karaoke scoring interface is shown. Using prior knowledge from the policy library, initial parameter tuning strategies for the next round of reinforcement learning are selected, and a set of audio parameter tuning operations corresponding to these strategies is established. The standard accompaniment audio for the user's next song is obtained, and the matching degree between the audio parameter features of the next song and the parameters of the three expressive ability quantification results is analyzed. Based on the parameter matching degree, the audio parameter tuning model uses reinforcement learning to determine the target audio parameter tuning operation suitable for the next song from the set of audio parameter tuning operations. This target audio parameter tuning operation is then performed on the standard accompaniment audio of the next song to obtain the tuned accompaniment audio. The karaoke scoring system then comprehensively analyzes the user's performance of the next song based on the tuned accompaniment audio, obtaining a karaoke score and multi-dimensional singing ability quantification results. Finally, using the karaoke score as the evaluation criterion, the audio parameter tuning model is trained and converged. Specifically, based on the audio parameter deviation value (audio... The parameter differences before and after adjustment are used to calibrate the karaoke score, resulting in a calibration score. The calibration score is negatively correlated with the evaluation (instant reward) of the reinforcement learning decision mentioned above; the higher the calibration score, the lower the instant reward. This makes the difficulty of singing the song after parameter adjustment closer to the karaoke user's own singing ability, and the initial parameter adjustment strategy is updated based on the instant reward. Similarly, during the karaoke user's subsequent song performances, based on a multi-dimensional analysis of the user's singing ability, and combined with prior knowledge from the policy library and reinforcement learning methods, more parameters built into the song are adaptively adjusted. Through comprehensive scoring and evaluation feedback, the current adjustment strategy is tentatively updated and modified, thus continuously converging to obtain a karaoke audio parameter adjustment strategy that matches the karaoke user. Through automatic parameter configuration based on reinforcement learning, karaoke users can more easily immerse themselves in singing, improving both the user's singing experience and their singing level.

[0216] As can be seen from the technical solutions provided in the embodiments of this application above, in the application scenario of song accompaniment audio processing based on reinforcement learning, the historical audio parameter tuning data of the target object is used to calibrate the historical parameter tuning performance data of the target object for historical songs, thereby obtaining the historical performance data of the target object. Based on the historical parameter adaptation data of the target object and historical songs, the historical audio parameter tuning data, and the historical performance data, the preset reinforcement learning model is trained to obtain the preset audio parameter tuning model. This allows the model's parameter tuning strategy to improve the target object's singing performance while helping the target object gradually reduce the intensity of the accompaniment audio parameter tuning. Then, the parameter adaptation data of the standard accompaniment audio of the song to be sung by the target object relative to the target object's singing ability representation data is determined, and the parameter adaptation data is input into the preset audio parameter tuning model to obtain the audio parameter tuning data. Based on the audio parameter tuning data, the parameters of the standard accompaniment audio are configured to obtain the target accompaniment audio of the song to be sung. Through adaptive audio parameter adjustment, the singing difficulty of the song to be sung can be made closer to the target object's current singing level. At the same time, during the long-term singing process of the target object, it helps the target object improve its singing ability while gradually reducing the intensity of the accompaniment audio parameter tuning, thereby improving the target object's singing experience.

[0217] This application also provides an audio processing device, such as... Figure 11 As shown, the audio processing device may include:

[0218] The standard accompaniment audio acquisition module 1110 is used to acquire the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object.

[0219] The parameter adaptation data determination module 1120 is used to determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data.

[0220] The audio parameter tuning and analysis module 1130 is used to input parameter adaptation data into the preset audio parameter tuning model, perform audio parameter tuning analysis on the standard accompaniment audio, and obtain audio parameter tuning data. The preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of historical songs, and the historical performance data of the target object singing historical songs.

[0221] The first parameter configuration module 1140 is used to configure the parameters of the standard accompaniment audio based on the audio parameter adjustment data to obtain the target accompaniment audio corresponding to the song to be sung.

[0222] In a specific embodiment, the aforementioned singing ability representation data may include: pitch accuracy representation data, rhythm ability representation data, and breath control ability representation data; the aforementioned parameter adaptation data may include: pitch adaptation data, rhythm adaptation data, and note duration adaptation data; and the aforementioned parameter adaptation data determination module 1120 may include:

[0223] The audio feature extraction unit is used to extract audio features from the standard accompaniment audio to obtain standard pitch data, lyric density data, and note duration data of the standard accompaniment audio.

[0224] The pitch adaptation data determination unit is used to determine the pitch adaptation data between the pitch accuracy representation data and the standard pitch data;

[0225] The rhythm adaptation data determination unit is used to determine the rhythm adaptation data between the rhythm capability representation data and the lyrics density data.

[0226] The note duration adaptation data determination unit is used to determine the note duration adaptation data between the breath ability representation data and the note duration data.

[0227] In one specific embodiment, the audio parameter tuning analysis module 1130 described above may include:

[0228] The target audio parameter tuning data determination unit is used to input parameter adaptation data into the preset audio parameter tuning model. Based on the decision reward representation information corresponding to the preset audio parameter tuning model, it determines the target audio parameter tuning data with the largest long-term reward data from a variety of preset audio parameter tuning data under the parameter adaptation data. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data. The long-term reward data is used to represent the feedback effect of parameter tuning decisions accumulated since the audio parameter tuning training of the preset reinforcement learning model.

[0229] The first parameter tuning data determination unit is used to determine audio parameter tuning data based on the target audio parameter tuning data.

[0230] In a specific embodiment, the aforementioned parameter adaptation data may include: pitch adaptation data, rhythm adaptation data, and note duration adaptation data; the aforementioned preset audio parameter adjustment model may include: pitch adjustment model, rhythm adjustment model, and note duration adjustment model; and the aforementioned audio parameter adjustment analysis module 1130 includes:

[0231] The pitch adjustment unit is used to input pitch adaptation data into the pitch adjustment model, perform pitch adjustment analysis on the standard accompaniment audio, and obtain pitch adjustment data corresponding to the pitch adaptation data.

[0232] The rhythm adjustment unit is used to input rhythm adaptation data into the rhythm adjustment model, perform rhythm adjustment analysis on the standard accompaniment audio, and obtain the rhythm adjustment data corresponding to the rhythm adaptation data.

[0233] The note duration adjustment unit is used to input note duration adaptation data into the note duration adjustment model, perform note duration adjustment analysis on the standard accompaniment audio, and obtain the note duration adjustment data corresponding to the note duration adaptation data.

[0234] The second parameter adjustment data determination unit is used to obtain audio parameter adjustment data based on pitch adjustment data, rhythm adjustment data, and note duration adjustment data.

[0235] In one specific embodiment, the above-mentioned preset audio parameter tuning model is trained using the following device:

[0236] The historical data acquisition module is used to acquire historical parameter adaptation data, historical audio parameter tuning data, historical performance data, and decision reward representation information corresponding to the preset reinforcement learning model for historical songs. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data.

[0237] The instant reward data determination module is used to determine the instant reward data obtained from historical audio parameter tuning data under historical parameter adaptation data based on historical performance data. The magnitude of the instant reward data is negatively correlated with the magnitude of the historical performance data.

[0238] The reward update module is used to update the decision reward representation information corresponding to the preset reinforcement learning model based on real-time reward data, so as to obtain the preset audio parameter tuning model.

[0239] In one specific embodiment, the aforementioned historical audio parameter tuning data is obtained through the following decision-making process:

[0240] The parameter tuning method indication information acquisition module is used to acquire parameter tuning method indication information. The parameter tuning method indication information is used to indicate that the probability determined by the historical audio parameter tuning data through reinforcement learning is the first probability, and the probability determined by random selection is the second probability.

[0241] The historical audio parameter tuning data determination module is used to determine historical audio parameter tuning data from a variety of preset audio parameter tuning data under historical parameter adaptation data, based on the parameter tuning method indication information and the decision reward representation information corresponding to the preset reinforcement learning model.

[0242] In one specific embodiment, the aforementioned historical performance data is obtained through a decision made by the following device:

[0243] The historical performance audio acquisition module is used to acquire historical performance audio of a target object singing a historical song based on historical audio parameter adjustment data;

[0244] The historical performance analysis module is used to analyze performance based on historical performance audio and obtain historical parameter-adjusted performance data.

[0245] The historical parameter deviation data determination module is used to determine the historical parameter deviation data between the target accompaniment audio corresponding to the historical song and the standard accompaniment audio corresponding to the historical song based on historical audio parameter tuning data.

[0246] The historical data calibration module is used to calibrate historical parameter adjustment performance data based on historical parameter deviation data, thereby obtaining historical performance data.

[0247] In one specific embodiment, the above-described apparatus may further include:

[0248] The current singing audio acquisition module is used to acquire the current singing audio while the target object is singing the song based on the target accompaniment audio;

[0249] The performance analysis module is used to analyze the performance based on the current singing audio and obtain performance data for parameter adjustment.

[0250] The audio parameter deviation data determination module is used to determine the audio parameter deviation data between the target accompaniment audio and the standard accompaniment audio based on the audio parameter tuning data.

[0251] The data calibration module is used to calibrate the performance data based on the audio parameter deviation data to obtain the current performance data.

[0252] The real-time model training module is used to perform real-time audio parameter tuning training on a preset reinforcement learning model based on parameter adaptation data, audio parameter tuning data, and current performance data.

[0253] In one specific embodiment, the above-described apparatus may further include:

[0254] The song style data acquisition module is used to acquire the song style data of the song to be sung.

[0255] The audio effect tuning and analysis module is used to input song style data into a preset audio effect tuning model, perform audio effect tuning and analysis on standard accompaniment audio, and obtain audio effect tuning data.

[0256] The second parameter configuration module is used to configure the parameters of the standard accompaniment audio based on the audio parameter tuning data and sound effect parameter tuning data, so as to obtain the target accompaniment audio corresponding to the song to be sung.

[0257] It should be noted that the apparatus in the device embodiment and the method embodiment are based on the same inventive concept.

[0258] This application provides an audio processing device, which includes a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the audio processing method provided in the above method embodiments.

[0259] Furthermore, Figure 12 A schematic diagram of the hardware structure of an audio processing device for implementing the audio processing method provided in the embodiments of this application is shown. The audio processing device may participate in or include the audio processing apparatus provided in the embodiments of this application. Figure 12 As shown, the audio processing device 120 may include one or more processors 1202 (shown as 1202a, 1202b, ..., 1202n in the figure) 1202 (processor 1202 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1204 for storing data, and a transmission device 1206 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 12 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, the audio processing device 120 may also include... Figure 12 The more or fewer components shown, or having the same Figure 12 The different configurations shown.

[0260] It should be noted that the aforementioned one or more processors 1202 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the audio processing device 120 (or mobile device). As involved in the embodiments of this application, the data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0261] The memory 1204 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the audio processing method described in the embodiments of this application. The processor 1202 executes various functional applications and data processing by running the software programs and modules stored in the memory 1204, thereby implementing the aforementioned audio processing method. The memory 1204 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1204 may further include memory remotely located relative to the processor 1202, and these remote memories can be connected to the audio processing device 120 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0262] The transmission device 1206 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the audio processing device 120. In one example, the transmission device 1206 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In one embodiment, the transmission device 1206 may be a radio frequency (RF) module for wireless communication with the Internet.

[0263] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the audio processing device 120 (or mobile device).

[0264] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an audio processing device to store at least one instruction or at least one program related to implementing the audio processing method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the audio processing method provided in the above method embodiment.

[0265] Optionally, in this embodiment, the storage medium may be located in at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0266] Embodiments of this application also provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio processing method provided in the method embodiments.

[0267] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0268] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0269] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and apparatus embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0270] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0271] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Obtain the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object; Determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability representation data; The parameter adaptation data is input into the preset audio parameter tuning model, and the standard accompaniment audio is subjected to audio parameter tuning analysis to obtain audio parameter tuning data. The preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of the historical songs, and the historical performance data of the target object singing the historical songs. Based on the audio parameter tuning data, the standard accompaniment audio is configured with parameters to obtain the target accompaniment audio corresponding to the song to be sung.

2. The method according to claim 1, characterized in that, The step of inputting the parameter adaptation data into a preset audio parameter tuning model and performing audio parameter tuning analysis on the standard accompaniment audio to obtain audio parameter tuning data includes: The parameter adaptation data is input into the preset audio parameter tuning model. Based on the decision reward representation information corresponding to the preset audio parameter tuning model, the target audio parameter tuning data with the largest long-term reward data is determined from a variety of preset audio parameter tuning data under the parameter adaptation data. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data. The long-term reward data is used to represent the feedback effect of parameter tuning decisions accumulated from the start of audio parameter tuning training of the preset reinforcement learning model. Based on the target audio parameter tuning data, the audio parameter tuning data is determined.

3. The method according to claim 1, characterized in that, The method further includes: The historical parameter adaptation data, historical audio parameter tuning data, historical performance data, and decision reward representation information corresponding to the preset reinforcement learning model are obtained. The decision reward representation information is used to represent the long-term reward data obtained by making decisions on different preset audio parameter tuning data under different parameter adaptation data. Based on the historical performance data, the instant reward data obtained by the decision-making of the historical audio parameter tuning data under the historical parameter adaptation data is determined. The value of the instant reward data is negatively correlated with the value of the historical performance data. Based on the real-time reward data, the decision reward representation information corresponding to the preset reinforcement learning model is updated to obtain the preset audio parameter tuning model.

4. The method according to claim 3, characterized in that, The historical performance data was obtained through the following methods: Obtain the historical audio recording of the target object singing the historical song based on the historical audio parameter adjustment data; Based on the historical singing audio, singing performance analysis was performed to obtain historical parameter-adjusted singing performance data; Based on the historical audio parameter tuning data, determine the historical parameter deviation data between the target accompaniment audio corresponding to the historical song and the standard accompaniment audio corresponding to the historical song; Based on the historical parameter deviation data, the historical parameter-adjusted singing performance data is calibrated to obtain the historical singing performance data.

5. The method according to claim 3, characterized in that, The historical audio parameter tuning data was obtained in the following way: Obtain parameter tuning method indication information, wherein the parameter tuning method indication information is used to indicate that the probability determined by the historical audio parameter tuning data through reinforcement learning is the first probability, and the probability determined by random selection is the second probability; Based on the parameter tuning method indication information and the decision reward representation information corresponding to the preset reinforcement learning model, the historical audio parameter tuning data is determined from a variety of preset audio parameter tuning data under the historical parameter adaptation data.

6. The method according to claim 1, characterized in that, After configuring the parameters of the standard accompaniment audio based on the audio parameter tuning data to obtain the target accompaniment audio corresponding to the song to be sung, the method further includes: During the process of the target object singing the song based on the target accompaniment audio, the current singing audio is captured; Based on the current singing audio, analyze the singing performance to obtain parameter-adjusted singing performance data; Based on the audio parameter tuning data, determine the audio parameter deviation data between the target accompaniment audio and the standard accompaniment audio; Based on the audio parameter deviation data, the adjusted singing performance data is calibrated to obtain the current singing performance data; Based on the parameter adaptation data, the audio parameter tuning data, and the current singing performance data, the preset reinforcement learning model is trained in real-time audio parameter tuning.

7. The method according to any one of claims 1 to 6, characterized in that, The singing ability representation data includes: pitch accuracy representation data, rhythmic ability representation data, and breath control representation data. The parameter adaptation data includes: pitch adaptation data, rhythmic adaptation data, and note duration adaptation data. The parameter adaptation data for determining the standard accompaniment audio relative to the singing ability representation data includes: Audio features are extracted from the standard accompaniment audio to obtain standard pitch data, lyric density data, and note duration data of the standard accompaniment audio; Determine the pitch adaptation data between the pitch accuracy representation data and the standard pitch data; Determine the rhythm adaptation data between the rhythm capability representation data and the lyrics density data; Determine the note duration adaptation data between the breath control performance data and the note duration data.

8. The method according to any one of claims 1 to 6, characterized in that, The parameter adaptation data includes: pitch adaptation data, rhythm adaptation data, and note duration adaptation data. The preset audio parameter tuning model includes: pitch adjustment model, rhythm adjustment model, and note duration adjustment model. The step of inputting the parameter adaptation data into the preset audio parameter tuning model and performing audio parameter tuning analysis on the standard accompaniment audio to obtain audio parameter tuning data includes: The pitch adaptation data is input into the pitch adjustment model, and the standard accompaniment audio is subjected to pitch adjustment analysis to obtain the pitch adjustment data corresponding to the pitch adaptation data. The rhythm adaptation data is input into the rhythm adjustment model, and the standard accompaniment audio is subjected to rhythm adjustment analysis to obtain the rhythm adjustment data corresponding to the rhythm adaptation data. The note duration adaptation data is input into the note duration adjustment model, and the note duration adjustment analysis is performed on the standard accompaniment audio to obtain the note duration adjustment data corresponding to the note duration adaptation data. The audio parameter data is obtained based on the pitch adjustment data, the rhythm adjustment data, and the note duration adjustment data.

9. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the song style data of the song to be sung; The song style data is input into a preset sound effect parameter tuning model, and the standard accompaniment audio is analyzed for sound effect parameter tuning to obtain sound effect parameter tuning data. Based on the audio parameter tuning data and the sound effect parameter tuning data, the standard accompaniment audio is configured with parameters to obtain the target accompaniment audio corresponding to the song to be sung.

10. An audio processing apparatus, characterized in that, The device includes: The standard accompaniment audio acquisition module is used to acquire the standard accompaniment audio corresponding to the song to be sung by the target object and the singing ability representation data of the target object. The parameter adaptation data determination module is used to determine the parameter adaptation data of the standard accompaniment audio relative to the singing ability characterization data. The audio parameter tuning and analysis module is used to input the parameter adaptation data into the preset audio parameter tuning model, perform audio parameter tuning analysis on the standard accompaniment audio, and obtain audio parameter tuning data. The preset audio parameter tuning model is obtained by real-time audio parameter tuning training of the preset reinforcement learning model based on the historical parameter adaptation data of historical songs, the historical audio parameter tuning data of the historical songs, and the historical performance data of the target object singing the historical songs. The first parameter configuration module is used to configure the parameters of the standard accompaniment audio based on the audio parameter tuning data to obtain the target accompaniment audio corresponding to the song to be sung.

11. An audio processing device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the audio processing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the audio processing method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the audio processing method as described in any one of claims 1 to 9.