Karaoke method, computer device, readable storage medium and computer program product
Patent Information
- Application Number
- CN202411342195.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-09-25
AI Technical Summary
[0003]传统技术以手机作为麦克风在电视上实现K歌的方案受制于手机的音频采集效果,若手机的音频采集效果不如专业麦克风,手机采集的音频在电视端的演唱效果会出现明显下降,同时,在密闭环境下若电视音量大,实际K歌过程中的演唱音频容易产生啸叫,影响K歌效果
[0039] The aforementioned karaoke method, device, computer equipment, computer-readable storage medium, and computer program product, in original sound mode, collect the user's current singing voice in real time, and, in conjunction with the current singing voice and user operation, detect whether the original sound mode meets preset switching conditions. When the original sound mode meets the preset switching conditions, it switches to artificial intelligence mode. In artificial intelligence mode, an artificial intelligence-generated voice simulating the user singing the target song replaces the user's singing voice for the subsequent stages, achieving intelligent karaoke. Based on the user's singing voice and user operation in original sound mode, it can flexibly select either the user's original singing voice or the artificial intelligence voice as the output voice for karaoke, avoiding the influence of factors such as song singing quality on the output voice and karaoke effect, thereby improving the karaoke singing experience.
Smart Images

Figure CN119207346B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a karaoke method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] Karaoke, a common form of entertainment, allows participants to sing songs with pre-recorded music accompaniment using a microphone. Participants can use audio capture devices such as mobile phones as microphones and audio playback devices such as televisions to perform karaoke.
[0003] Traditional karaoke solutions that use mobile phones as microphones on TVs are limited by the audio capture quality of mobile phones. If the audio capture quality of a mobile phone is not as good as that of a professional microphone, the singing effect of the audio captured by the mobile phone on the TV will be significantly reduced. At the same time, in a closed environment, if the TV volume is high, the singing audio during the actual karaoke process is prone to feedback, which will affect the karaoke effect.
[0004] Therefore, traditional methods suffer from poor singing quality in vocal audio. Summary of the Invention
[0005] Therefore, it is necessary to provide a karaoke method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the singing effect of karaoke, in response to the above-mentioned technical problems.
[0006] Firstly, this application provides a karaoke mode switching method, applied to an audio processing device, comprising:
[0007] The current singing voice of a user singing a target song, captured by an audio acquisition device in original mode, is used as the current singing voice of the target song.
[0008] When the original sound mode meets the first preset switching condition, the audio acquisition device is switched from the original sound mode to the artificial intelligence mode; wherein the original sound mode meets the first preset switching condition, including: the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0009] The audio acquisition device acquires an AI-generated singing voice obtained by simulating the user singing the target song at a subsequent stage using AI technology in the AI mode, and plays the AI-generated singing voice; wherein the subsequent stage is the stage after the current stage.
[0010] In one embodiment, acquiring the AI singing voice obtained by the audio acquisition device using AI technology to simulate the user singing the target song in the AI mode includes:
[0011] In response to switching the original sound mode to the artificial intelligence mode, a mode switching command is sent to the audio acquisition device; the mode switching command is used to control the audio acquisition device to send the artificial intelligence singing voice to the audio processing device.
[0012] The system receives the AI-generated singing voice sent by the audio acquisition device and plays the AI-generated singing voice according to the singing progress of the target song.
[0013] In one embodiment, the method further includes:
[0014] Obtain the current singing voice of the user singing the target song;
[0015] Obtain the feedback detection result corresponding to the current singing voice, and determine the singing quality of the current singing voice based on the feedback detection result; the feedback detection result is used to characterize whether there is feedback in the current singing voice.
[0016] In one embodiment, the method further includes:
[0017] Obtain the current singing voice of the user singing the target song;
[0018] Obtain the singing score of the currently sung song, and determine the song singing quality of the currently sung song based on the singing score.
[0019] In one embodiment, the method further includes:
[0020] In the artificial intelligence mode, the audio acquisition device acquires the current singing progress of the user singing the target song as the current singing voice of the target song.
[0021] When the artificial intelligence mode meets the second preset switching conditions, the artificial intelligence mode is switched to the original sound mode; wherein the artificial intelligence mode meets the second preset switching conditions, including: the singing quality of the current singing voice is higher than or equal to the preset standard quality, or the song segment corresponding to the current singing voice is not a specified song segment in the target song, or the audio processing device responds to the second mode switching operation triggered by the user.
[0022] Play the current singing voice of the user singing the target song, and acquire and play the subsequent singing voice of the user singing the target song in the original sound mode, which was acquired by the audio acquisition device.
[0023] In one embodiment, the audio processing device and the audio acquisition device are integrated into the same audio device.
[0024] Secondly, this application also provides a karaoke method applied to an audio acquisition device, the method comprising:
[0025] In the original audio mode, the current singing progress of the user singing the target song is captured as the current singing voice of the target song;
[0026] The current singing voice is sent to the audio processing device; the audio processing device is used to switch the audio acquisition device from the original sound mode to the artificial intelligence mode when the original sound mode meets the first preset switching condition; wherein the original sound mode meets the first preset switching condition includes: the singing quality of the current singing voice is lower than the preset standard quality, or the song segment corresponding to the current singing voice is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user;
[0027] In the AI mode, AI technology is used to simulate the user's singing of the target song at a later stage to obtain an AI singing voice, which is then sent to an audio processing device for playback; the later stage refers to the stage after the current stage.
[0028] In one embodiment, the step of using artificial intelligence technology to simulate the user's singing of the target song at subsequent stages to obtain an AI-generated singing voice, and then sending the AI-generated singing voice to an audio processing device for playback, includes:
[0029] Based on the timestamp corresponding to the current singing of the target song, determine the timestamp corresponding to the subsequent progress of the target song;
[0030] The AI singing voice is obtained by simulating the user's singing of the target song after the timestamp corresponding to the subsequent progress of the song using artificial intelligence technology.
[0031] The AI-generated singing voice is sent to the audio processing device for playback; the audio processing device is used to play the AI-generated singing voice starting from the timestamp corresponding to the subsequent progress of the target song.
[0032] Thirdly, this application also provides a karaoke device, applied to an audio processing equipment, comprising:
[0033] The acquisition module is used to acquire the current progress of the singing of the target song by the user, which is acquired by the audio acquisition device in the original mode, and use it as the current singing voice of the target song.
[0034] A switching module is used to switch the audio acquisition device from the original sound mode to the artificial intelligence mode when the original sound mode meets the first preset switching conditions; wherein the original sound mode meets the first preset switching conditions, including: the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0035] The playback module is used to acquire the AI singing voice obtained by the audio acquisition device in the AI mode using AI technology to simulate the user singing the target song at a subsequent stage, and to play the AI singing voice; wherein the subsequent stage is the stage after the current stage.
[0036] Fourthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the steps of the method described above.
[0037] Fifthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0038] Sixthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.
[0039] The aforementioned karaoke method, device, computer equipment, computer-readable storage medium, and computer program product, in original sound mode, collect the user's current singing voice in real time, and, in conjunction with the current singing voice and user operation, detect whether the original sound mode meets preset switching conditions. When the original sound mode meets the preset switching conditions, it switches to artificial intelligence mode. In artificial intelligence mode, an artificial intelligence-generated voice simulating the user singing the target song replaces the user's singing voice for the subsequent stages, achieving intelligent karaoke. Based on the user's singing voice and user operation in original sound mode, it can flexibly select either the user's original singing voice or the artificial intelligence voice as the output voice for karaoke, avoiding the influence of factors such as song singing quality on the output voice and karaoke effect, thereby improving the karaoke singing experience. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is an application environment diagram of a karaoke method in one embodiment;
[0042] Figure 2 This is a flowchart illustrating a karaoke method in one embodiment;
[0043] Figure 3 This is a schematic diagram of a process for optimizing the karaoke experience using a mobile phone microphone using AI technology in one embodiment;
[0044] Figure 4 This is a structural block diagram of a karaoke device in one embodiment;
[0045] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0047] The karaoke method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, the audio processing device 102 communicates with the audio acquisition device 104 via a network. A data storage system can store the data that the audio acquisition device 104 needs to process. The data storage system can be integrated into the audio acquisition device 104 or placed in the cloud or on another network server. The audio processing device 102 acquires the current progress of the user singing the target song, acquired by the audio acquisition device in original mode, as the current singing voice of the target song. When the original mode meets a first preset switching condition, the audio processing device 102 switches the audio acquisition device from original mode to artificial intelligence mode. The first preset switching condition includes: the singing quality of the current singing voice is lower than a preset standard quality, or the song segment corresponding to the current singing voice is a specified song segment in the target song, or the audio processing device responds to a first mode switching operation triggered by the user. The audio processing device 102 acquires the artificial intelligence voice obtained by simulating the user singing the target song at a subsequent progress point using artificial intelligence technology in artificial intelligence mode, and plays the artificial intelligence voice. The subsequent progress point refers to the progress point after the current progress point. The audio processing device 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The audio acquisition device 104 can include, but is not limited to, microphones and smartphones.
[0048] In one exemplary embodiment, such as Figure 2 As shown, a karaoke method is provided. Taking the application of this method to an audio processing device as an example, it includes the following steps S202 to S206. Wherein:
[0049] Step S202: Obtain the current progress of the user singing the target song, which is captured by the audio acquisition device in the original sound mode, as the current singing voice of the target song.
[0050] Audio acquisition devices can refer to hardware devices used to capture sound and convert it into electronic signals. In practical applications, audio acquisition devices can include, but are not limited to, microphones and mobile terminals (such as mobile phones).
[0051] Among them, the original sound mode can refer to the mode used to play the user's real voice. In practical applications, the singing voice played by the audio processing device in the original sound mode is the user's real voice or the real voice after voice processing.
[0052] In this context, "user" refers to the singer. In practical applications, "user" can include any person singing a song through an audio capture device.
[0053] The target song can refer to the song sung by the user during karaoke. In practical applications, the target song can include any song in a preset song database.
[0054] The current progress can refer to the current singing progress of the target song at the current moment. In practical applications, the current progress can be determined by the timestamp of the song being sung by the user at the current moment.
[0055] As an example, in the original sound mode, the singing voice played by the audio processing device is the real human voice captured by the audio acquisition device. In order to ensure a good karaoke effect, the audio processing device can obtain the current progress of the user singing the target song captured by the audio acquisition device in the original sound mode. As the current singing voice of the target song in the original sound mode, the audio processing device can combine the singing quality corresponding to the current singing voice, the song segment, and the user operation to determine whether to play the current singing voice.
[0056] Step S204: If the original sound mode meets the first preset switching condition, the audio acquisition device is switched from the original sound mode to the artificial intelligence mode; wherein the original sound mode meets the first preset switching condition, including: the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0057] The first preset switching condition can refer to information used to determine whether to switch from the original sound mode to the artificial intelligence mode. In practical applications, the first preset switching condition may include, but is not limited to, the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0058] Among them, the artificial intelligence mode can refer to a mode that uses artificial intelligence technology to simulate the user's singing voice and replace the user's real voice with it. In practical applications, the artificial intelligence mode can be regarded as an AI mode.
[0059] As an example, the audio processing device, in its original sound mode, combines the singing quality of the user's current vocal performance, the song segment, and user actions to determine in real time whether the original sound mode meets the first preset switching condition. If the original sound mode meets the first preset switching condition, the audio acquisition device switches from original sound mode to artificial intelligence mode. In artificial intelligence mode, the audio processing device can use artificial intelligence technology to simulate the user's singing voice, replacing the user's real voice, and then play it. In practical applications, if the singing quality of the current vocal performance is lower than the preset standard quality, it indicates that playing the current vocal performance will affect the karaoke experience, and the audio processing device can determine that the original sound mode meets the first preset switching condition. Users can also specify a song segment within the target song. If the song segment corresponding to the current vocal performance is the specified song segment within the target song, the audio processing device can determine that the original sound mode meets the first preset switching condition. Users can also control the audio processing device to perform a first mode switching operation. When the audio processing device responds to the user-triggered first mode switching operation, the audio processing device can determine that the original sound mode meets the first preset switching condition. In practical applications, when the original sound mode does not meet the first preset switching condition, the audio processing device does not need to switch the audio acquisition device from the original sound mode to the artificial intelligence mode. The audio acquisition device remains in the original sound mode, and the audio processing device can play the currently sung voice acquired by the audio acquisition device.
[0060] Step S206: Obtain the AI singing voice obtained by the audio acquisition device in AI mode using AI technology to simulate the subsequent progress of the user singing the target song, and play the AI singing voice.
[0061] The subsequent progress can refer to the progress after the current progress. In practical applications, the current progress can be determined by the timestamp of the song sung by the user at the current moment, and the subsequent progress can be determined by the timestamp after the timestamp corresponding to the current progress.
[0062] Among them, AI singing voice can refer to the singing voice generated by using artificial intelligence technology to simulate a user singing a target song. In practical applications, the timbre of AI singing voice can be the same as that of the user's timbre.
[0063] As an example, in AI mode, the audio acquisition device can use AI technology to simulate the user's singing of the target song at a later stage to obtain an AI singing voice. The AI mode uses AI technology to simulate the user's singing voice to replace the user's real voice and play it. Therefore, the audio processing device can play the AI singing voice. At this time, the content played by the audio processing device can include the AI singing voice and the corresponding accompaniment audio in the target song.
[0064] In the aforementioned karaoke method, in the original vocal mode, the user's current singing voice is collected in real time. The system then combines the current singing voice with the user's actions to detect whether the original vocal mode meets the preset switching conditions. If the original vocal mode meets the preset switching conditions, it is switched to the artificial intelligence mode. In the artificial intelligence mode, an AI-generated voice, simulated by the user singing the target song, replaces the user's singing voice for the subsequent parts of the song, achieving intelligent karaoke. Based on the user's singing voice and user actions in the original vocal mode, the system can flexibly select either the user's original vocal voice or the AI-generated voice as the output vocal voice for karaoke, avoiding the influence of factors such as song performance quality on the output vocal voice and karaoke effect, thereby improving the karaoke performance.
[0065] In an exemplary embodiment, acquiring an AI-generated vocal track obtained by simulating the subsequent progress of a user singing a target song using AI technology in an AI mode via an audio acquisition device includes: in response to switching from the original audio mode to the AI mode, sending a mode switching instruction to the audio acquisition device; the mode switching instruction being used to control the audio acquisition device to send the AI-generated vocal track to the audio processing device; receiving the AI-generated vocal track sent by the audio acquisition device, and playing the AI-generated vocal track according to the singing progress of the target song.
[0066] Among them, the mode switching command can refer to the information used to control the audio acquisition device to send the AI singing voice to the audio processing device.
[0067] As an example, in response to switching from original vocal mode to AI mode, the audio processing device can send a mode switching command to the audio acquisition device. Upon receiving the command, the audio acquisition device can send the AI-generated vocal track, derived in real-time using AI technology, to the audio processing device. The audio processing device receives the AI-generated vocal track and plays it according to the singing progress. Understandably, the audio acquisition device can generate an AI-generated vocal track for the entire target song and / or a specific segment of the target song in real-time, based on real-time captured user vocals and / or pre-captured user vocals. In practical applications, if the current karaoke mode is original vocal mode, the user can click the "AI mode button" / "mode switching button" to trigger a mode switching operation, and the audio processing device can switch the current mode (original vocal mode) to AI mode; conversely, if the current karaoke mode is AI mode, the user can click the "original vocal mode button" / "mode switching button" to trigger a mode switching operation, and the audio processing device can switch the current mode (AI mode) back to original vocal mode.
[0068] In an exemplary embodiment, the method further includes: obtaining the current singing voice of the user singing the target song; obtaining the feedback detection result corresponding to the current singing voice, and determining the singing quality of the current singing voice based on the feedback detection result; the feedback detection result is used to characterize whether there is feedback in the current singing voice.
[0069] Among them, the feedback detection result can refer to the information used to characterize whether there is feedback in the current singing. In practical applications, feedback can refer to the phenomenon that occurs when the sound in the amplification system is captured by the microphone again and amplified in an infinite loop.
[0070] As an example, song performance quality helps distinguish and determine the parts of the user's singing of a target song that require automatic switching to karaoke mode. Feedback can affect song performance quality. To obtain accurate song performance quality, the audio processing device can acquire the user's current singing voice of the target song, then analyze the current singing voice according to a preset feedback detection method, determine the feedback detection result corresponding to the current singing voice, and determine the song performance quality based on the feedback detection result. For example, if the feedback detection result indicates that there is feedback in the current singing voice, the audio processing device can determine that the song performance quality of the current singing voice is lower than the preset standard quality, and thus determine that the original sound mode meets the first preset switching condition, and the audio processing device needs to switch to karaoke mode; if the feedback detection result indicates that there is no feedback in the current singing voice, the audio processing device can determine that the song performance quality of the current singing voice is higher than or equal to the preset standard quality, and the audio processing device does not need to switch to karaoke mode, that is, the audio processing device is still in original sound mode.
[0071] In this embodiment, by obtaining the feedback detection result corresponding to the current singing voice and determining the singing quality based on the feedback detection result, it is possible to accurately analyze whether there is feedback in the current singing voice and obtain accurate singing quality. This facilitates subsequent accurate judgment on whether to switch from the original mode to the artificial intelligence mode based on the singing quality, realizes accurate switching of the karaoke mode, avoids the singing quality affecting the karaoke effect, and improves the singing effect of karaoke.
[0072] In an exemplary embodiment, the method further includes: obtaining the current singing voice of the user singing the target song; obtaining the singing score of the current singing voice; and determining the singing quality of the current singing voice based on the singing score.
[0073] Among them, the singing score can refer to information that characterizes the user's singing performance of the target song. In practical applications, the singing score can be obtained by the scoring model processing the current singing voice.
[0074] As an example, singing scores can also affect the quality of a song's performance. An audio processing device can acquire the current vocal quality of a user singing a target song, then input this vocal quality into a pre-trained scoring model to determine the singing score and, based on the score, the overall quality of the song. For instance, the audio processing device can use the scoring model to acquire the singing score of each phrase (such as the current phrase or the current vocal quality) during the user's performance of the target song in real time. When the singing score is lower than a preset singing score threshold, the audio processing device can determine that the current vocal quality is lower than a preset standard quality, thus determining that the original sound mode meets the first preset switching condition, and the audio playback device needs to switch to karaoke mode. When the singing score is higher than or equal to the preset singing score threshold, the audio playback device can determine that the current vocal quality is higher than or equal to the preset standard quality, and the audio processing device does not need to switch to karaoke mode; that is, the audio processing device remains in original sound mode.
[0075] In this embodiment, by obtaining the singing score corresponding to the current singing voice and determining the singing quality of the song based on the singing score, the singing score of the current singing voice can be accurately analyzed to obtain the accurate singing quality of the song. This facilitates the subsequent accurate judgment on whether to switch to karaoke mode based on the singing quality of the song, avoids the singing quality of the song affecting the karaoke effect, and improves the karaoke singing effect.
[0076] In an exemplary embodiment, the method further includes: acquiring the current singing voice of a user singing a target song, acquired by the audio acquisition device in artificial intelligence mode, as the current singing voice of the target song; switching the artificial intelligence mode to the original sound mode when the artificial intelligence mode meets a second preset switching condition; wherein the artificial intelligence mode meeting the second preset switching condition includes: the singing quality of the current singing voice is higher than or equal to a preset standard quality, or the song segment corresponding to the current singing voice is not a specified song segment in the target song, or the audio processing device responds to the second mode switching operation triggered by the user; playing the current singing voice of the user singing the target song, and acquiring and playing the subsequent singing voice of the user singing the target song, acquired by the audio acquisition device in original sound mode.
[0077] The second preset switching condition can refer to information used to determine whether to switch the artificial intelligence mode to the original sound mode. In practical applications, the second preset switching condition may include, but is not limited to, the singing quality of the song being sung is higher than or equal to the preset standard quality, or the song segment corresponding to the current singing is not a specified song segment in the target song, or the audio processing device responding to the second mode switching operation triggered by the user.
[0078] As an example, in AI mode, the audio processing device plays an AI-generated voice that simulates the user singing the target song. To ensure a good karaoke experience, the audio processing device can acquire the current vocal recording of the user singing the target song, captured by the audio acquisition device in AI mode. As the current vocal recording of the target song in AI mode, the audio processing device can combine the singing quality, song segments, and user actions corresponding to the current vocal recording to determine whether to play the actual human voice of the user singing the target song, captured by the audio acquisition device.
[0079] In AI mode, the audio processing device combines the current singing quality of the user's performance of the target song, the song segment, and user actions to determine in real time whether the AI mode meets the second preset switching condition. When the AI mode meets the second preset switching condition, the audio processing device can switch the audio acquisition device from AI mode to original sound mode. In original sound mode, the audio processing device can play the real human voice of the user singing the target song, which has been captured by the audio acquisition device. In practical applications, if the singing quality of the current performance is higher than or equal to the preset standard quality, it means that playing the current performance will not affect the karaoke effect, and the audio processing device can determine that the AI mode meets the second preset switching condition. The user can also specify a song segment in the target song. If the song segment corresponding to the current performance is not the specified song segment in the target song, the audio processing device can determine that the AI mode meets the second preset switching condition. The user can also control the audio processing device to perform a second mode switching operation. When the audio processing device responds to the user-triggered second mode switching operation, the audio processing device can determine that the AI mode meets the second preset switching condition. In practical applications, when the artificial intelligence mode does not meet the second preset switching condition, the audio processing device does not need to switch the audio acquisition device from artificial intelligence mode to original sound mode. The audio acquisition device remains in artificial intelligence mode, and the audio processing device can play the currently sung song acquired by the audio acquisition device.
[0080] In the original sound mode, the audio acquisition device can capture the user's real vocals singing the target song as the current singing voice in the original sound mode. The audio processing device can play the current singing voice at this time. The audio processing device can also play the subsequent singing voices captured by the audio acquisition device in the original sound mode. At this time, the content played by the audio processing device can include the user's real vocals in the original sound mode (such as the current singing voice captured by the audio acquisition device in the original sound mode) and the accompaniment audio corresponding to the real vocals in the target song.
[0081] In one exemplary embodiment, the audio processing device and the audio acquisition device can also be integrated into the same audio device. For specific limitations on the karaoke method implemented using this audio device, please refer to the limitations on the karaoke method above, which will not be repeated here.
[0082] In an exemplary embodiment, a karaoke method is provided, which is illustrated by taking the application of the method to an audio acquisition device as an example, and includes the following steps:
[0083] In original audio mode, the current singing progress of the user singing the target song is captured as the current singing voice of the target song.
[0084] Among them, the original sound mode can refer to the mode used to play the user's real voice. In practical applications, the singing voice played by the audio processing device in the original sound mode is the user's real voice or the real voice after voice processing.
[0085] In this context, "user" refers to the singer. In practical applications, "user" can include any person singing a song through an audio capture device.
[0086] The target song can refer to the song sung by the user during karaoke. In practical applications, the target song can include any song in a preset song database.
[0087] The current progress can refer to the current singing progress of the target song at the current moment. In practical applications, the current progress can be determined by the timestamp of the song being sung by the user at the current moment.
[0088] As an example, in original sound mode, the audio acquisition device can capture the current singing progress of the user singing the target song as the current singing voice of the target song. The audio acquisition device can send the current singing voice to the audio processing device, so that the audio processing device can combine the singing quality corresponding to the current singing voice, the song segment, and the operation performed by the user to determine whether to play the current singing voice.
[0089] The current singing voice is sent to the audio processing device; the audio processing device is used to switch the audio acquisition device from the original voice mode to the artificial intelligence mode when the original voice mode meets the first preset switching conditions; wherein the original voice mode meets the first preset switching conditions, including: the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing voice is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0090] The first preset switching condition can refer to information used to determine whether to switch from the original sound mode to the artificial intelligence mode. In practical applications, the first preset switching condition may include, but is not limited to, the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the current singing is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0091] Among them, the artificial intelligence mode can refer to a mode that uses artificial intelligence technology to simulate the user's singing voice and replace the user's real voice with it. In practical applications, the artificial intelligence mode can be regarded as an AI mode.
[0092] As an example, the audio acquisition device can send the current singing voice to the audio processing device. The audio processing device, based on the singing quality of the user's current singing voice, the song segment, and the user's actions, determines in real time whether the original sound mode meets the first preset switching condition. When the original sound mode meets the first preset switching condition, the audio processing device can switch the audio acquisition device from the original sound mode to the artificial intelligence mode. In the artificial intelligence mode, the audio processing device can use artificial intelligence technology to simulate the user's singing voice, replacing the user's real voice, and then play it. In practical applications, if the singing quality of the current singing voice is lower than the preset standard quality, it means that playing the current singing voice will affect the karaoke experience, and the audio processing device can determine that the original sound mode meets the first preset switching condition. The user can also specify a song segment in the target song. If the song segment corresponding to the current singing voice is the specified song segment in the target song, the audio processing device can determine that the original sound mode meets the first preset switching condition. The user can also control the audio processing device to perform a first mode switching operation. When the audio processing device responds to the user-triggered first mode switching operation, the audio processing device can determine that the original sound mode meets the first preset switching condition. In practical applications, when the original sound mode does not meet the first preset switching condition, the audio processing device does not need to switch the audio acquisition device from the original sound mode to the artificial intelligence mode. The audio acquisition device remains in the original sound mode, and the audio processing device can play the currently sung voice acquired by the audio acquisition device.
[0093] In artificial intelligence mode, the AI singing voice is generated by simulating the subsequent progress of the user singing the target song using artificial intelligence technology, and then sent to the audio processing device for playback.
[0094] The subsequent progress can refer to the progress after the current progress. In practical applications, the current progress can be determined by the timestamp of the song sung by the user at the current moment, and the subsequent progress can be determined by the timestamp after the timestamp corresponding to the current progress.
[0095] Among them, AI singing voice can refer to the singing voice generated by using artificial intelligence technology to simulate a user singing a target song. In practical applications, the timbre of AI singing voice can be the same as that of the user's timbre.
[0096] As an example, in AI mode, the audio acquisition device can use AI technology to simulate the user's singing of the target song at a later stage to obtain an AI singing voice. The AI mode uses AI technology to simulate the user's singing voice to replace the user's real voice and play it. Therefore, the audio processing device can play the AI singing voice. At this time, the content played by the audio processing device can include the AI singing voice and the corresponding accompaniment audio in the target song.
[0097] In the aforementioned karaoke method, the user's current singing voice is captured in real time in the original audio mode. Audio processing equipment is used to detect whether the original audio mode meets preset switching conditions based on the current singing voice and user operations. When the original audio mode meets the preset switching conditions, it is switched to artificial intelligence mode. In artificial intelligence mode, an AI-generated voice simulating the user's singing of the target song replaces the user's subsequent singing, achieving intelligent karaoke. Based on the user's singing voice and user operations in the original audio mode, it can flexibly select either the user's original singing voice or the AI-generated voice as the output voice for karaoke, avoiding the influence of factors such as song performance quality on the output voice and karaoke effect, thereby improving the karaoke performance.
[0098] In one exemplary embodiment, an AI-generated vocal track is obtained by simulating a user singing a subsequent segment of a target song using artificial intelligence technology, and then sent to an audio processing device for playback. This includes: determining a timestamp corresponding to the subsequent segment of the target song based on the timestamp corresponding to the current segment of the song; simulating a user singing a segment after the timestamp corresponding to the subsequent segment of the target song using artificial intelligence technology; and sending the AI-generated vocal track to the audio processing device for playback. The audio processing device is used to play the AI-generated vocal track starting from the timestamp corresponding to the subsequent segment of the target song.
[0099] The timestamp corresponding to the current singing can be information that represents the position of the song segment sung by the user at the current moment in the target song.
[0100] The timestamp corresponding to the subsequent progress of the target song can refer to information representing the position of the song segment corresponding to the subsequent progress of the target song within the target song.
[0101] As an example, the audio acquisition device can first obtain the timestamp corresponding to the current singing of the target song. Then, based on the timestamp of the current singing, it determines the timestamps corresponding to the subsequent stages of the target song. The server can then use artificial intelligence technology to simulate the user singing the song after the timestamps corresponding to the subsequent stages, obtaining an AI-generated voice. This AI-generated voice is then sent to the audio processing device for playback, allowing the audio processing device to start playing the AI-generated voice from the timestamps corresponding to the subsequent stages of the target song. For example, the timestamp of the song segment corresponding to the current singing of the target song can be represented as 1 minute and 10 seconds. The audio acquisition device can determine the timestamps corresponding to the subsequent stages of the target song, such as 1 minute and 15 seconds. The audio acquisition device can then use artificial intelligence technology to simulate the user singing the song after the timestamps corresponding to the subsequent stages, obtaining an AI-generated voice. This AI-generated voice (along with the timestamps of the subsequent stages of the target song) is sent to the audio processing device, which can then start playing the AI-generated voice from the timestamps corresponding to the subsequent stages of the target song.
[0102] In this embodiment, by determining the timestamp corresponding to the current singing voice and the timestamp corresponding to the subsequent progress of the target song, the audio processing device can accurately start playing the AI singing voice from the timestamp corresponding to the subsequent progress of the target song. This avoids overlap or excessive intervals between the user's real voice and the AI singing voice due to mode switching during karaoke, thereby improving the naturalness of karaoke and optimizing the singing effect.
[0103] In some embodiments, the TV karaoke application can set an AI voice switching mode (such as an automatic switching mode): (1) When the user sings the target song, if the singing score is lower than the preset singing score threshold, the next line will automatically switch to AI voice (i.e., switch to artificial intelligence mode), and when the singing score is higher than or equal to the preset singing score threshold, it will automatically switch back to the original voice mode; (2) In the target segment (such as the chorus), the original voice mode will be switched to AI voice mode (i.e., artificial intelligence mode). The mobile karaoke application can set a segment mode: When the user sings the target song, it will detect whether there is feedback in the current singing voice. When feedback occurs, the mobile phone can turn off the voice acquisition, and the TV, as an audio processing device, can switch to artificial intelligence mode and can switch back to the original voice mode at any time.
[0104] As an example, such as Figure 3The diagram illustrates a process for optimizing the karaoke experience using a mobile phone microphone using AI technology. The mobile phone and TV can connect to the same local area network (LAN) and establish a UDP communication connection. The mobile phone's microphone hardware receives the vocals, which then pass through the hardware driver layer, the Android Audio Framework layer, and finally to the mobile karaoke application. The application performs a series of vocal processing steps, including feedback suppression and sound effect compensation, before sending the data via UDP to the TV's karaoke sound receiver (e.g., a karaoke sound receiver). The karaoke sound receiver then hands the sound to the Mic Service middleware, which writes the sound data to the system process and sends it to the HAL driver for final output. Specifically, after the user starts singing a target song using the mobile phone microphone, the mobile phone's AI inference module is activated. Based on the automatic switching mode or segment mode, it infers the AI vocals for the entire song or a segment of AI vocals in real time. The TV can also switch to AI vocal mode; that is, after switching to AI vocal mode, it no longer receives the mobile phone's vocals but directly uses the AI vocals generated on the TV to replace the original singing voice.
[0105] For the automatic switching mode, when the TV karaoke app detects that the singing score of the current line is too low, it sends a mode switching command to the mobile karaoke app, requesting an AI mode switch (i.e., switching to artificial intelligence mode). After receiving the mode switching command, the mobile karaoke app turns off the voice capture and voice processing functions, switches to AI inference mode (such as artificial intelligence mode), and obtains AI voice data (such as AI singing voice) corresponding to the singing progress according to the timestamp of the next line (such as singing progress), and sends it to the TV. The TV can play the AI singing voice and the corresponding accompaniment audio. When the TV karaoke app detects that the singing score of the current singing voice has returned to normal (such as being higher than or equal to the preset singing score threshold), the TV sends the mode switching command to the mobile app again, requesting a switch to original sound mode. After receiving the mode switching command, the mobile karaoke app restores the original sound mode, restarts the voice capture and voice processing functions, and sends the processed current singing voice (such as the current singing voice) to the TV.
[0106] For the segment mode, when the TV karaoke app plays a specified segment (such as a designated segment within a target song), it sends a mode switching command to the mobile app, switching to AI mode (such as artificial intelligence mode). After the specified segment finishes playing, it sends a mode switching command back to original audio mode. This can be understood as the mobile app also having a mode switching switch. When the user experiences feedback or a low singing score, they can manually switch to AI mode (such as artificial intelligence mode) by operating their phone. When there is no feedback during the current singing or the singing score returns to normal (such as being higher than or equal to a preset singing score threshold), the user can manually switch back to original audio mode by operating their phone.
[0107] In this embodiment, AI technology is used to generate AI voices in real time on the mobile phone. When AI mode is enabled, the voice acquisition function is disabled, and the AI voice is used to replace the original voice and sent to the TV. The playback logic on the TV remains unchanged. The AI voice can be inferred in real time using AI technology. Based on the real-time singing situation, the system can intelligently switch between the original voice and the AI voice, ultimately achieving the goal of optimizing the mobile phone microphone karaoke experience using AI technology.
[0108] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0109] Based on the same inventive concept, this application also provides a karaoke device for implementing the karaoke method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more karaoke device embodiments provided below can be found in the limitations of the karaoke method described above, and will not be repeated here.
[0110] In one exemplary embodiment, such as Figure 4 As shown, a karaoke device is provided, applied to an audio processing device, including: an acquisition module 402, a switching module 404, and a playback module 406, wherein:
[0111] The acquisition module 402 is used to acquire the current progress of the singing of the target song by the user, which is acquired by the audio acquisition device in the original sound mode, as the current singing voice of the target song.
[0112] The switching module 404 is used to switch the audio acquisition device from the original sound mode to the artificial intelligence mode when the original sound mode meets the first preset switching conditions; wherein the original sound mode meets the first preset switching conditions, including: the singing quality of the song being sung is lower than the preset standard quality, or the song segment corresponding to the song being sung is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0113] The playback module 406 is used to acquire the AI singing voice obtained by the audio acquisition device in the AI mode using AI technology to simulate the user singing the target song at a later stage, and to play the AI singing voice; wherein the later stage is the stage after the current stage.
[0114] In one exemplary embodiment, the acquisition module 402 is further configured to send a mode switching instruction to the audio acquisition device in response to switching the original sound mode to the artificial intelligence mode; the mode switching instruction is used to control the audio acquisition device to send the artificial intelligence singing voice to the audio processing device; receive the artificial intelligence singing voice sent by the audio acquisition device, and play the artificial intelligence singing voice according to the singing progress of the target song.
[0115] In one exemplary embodiment, the above-described apparatus further includes a howling detection module, which is specifically used to acquire the current singing voice of the user singing the target song; acquire the howling detection result corresponding to the current singing voice, and determine the song singing quality of the current singing voice based on the howling detection result; the howling detection result is used to characterize whether howling exists in the current singing voice.
[0116] In one exemplary embodiment, the apparatus further includes a scoring module, which is specifically used to obtain the current singing voice of the user singing the target song; obtain a singing score of the current singing voice; and determine the song singing quality of the current singing voice based on the singing score.
[0117] In one exemplary embodiment, the device further includes a second switching module, which is specifically configured to acquire the current singing voice of a user singing a target song, acquired by the audio acquisition device in the artificial intelligence mode, as the current singing voice of the target song; and, when the artificial intelligence mode meets a second preset switching condition, switch the artificial intelligence mode to the original sound mode; wherein the artificial intelligence mode meets the second preset switching condition including: the singing quality of the current singing voice is higher than or equal to a preset standard quality, or the song segment corresponding to the current singing voice is not a specified song segment in the target song, or the audio processing device responds to the second mode switching operation triggered by the user; play the current singing voice of the user singing the target song, and acquire and play the subsequent singing voice of the user singing the target song, acquired by the audio acquisition device in the original sound mode.
[0118] In one exemplary embodiment, the above-described apparatus further includes an integration module specifically configured to integrate the audio processing device and the audio acquisition device into a single audio device.
[0119] The modules in the aforementioned karaoke device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0120] In one exemplary embodiment, a karaoke device is provided, applied to an audio acquisition device, comprising: an acquisition module, a determination module, and a transmission module, wherein:
[0121] The acquisition module is used to acquire the current singing progress of the user singing the target song in the original audio mode as the current singing voice of the target song.
[0122] The determination module is used to send the current singing voice to the audio processing device; the audio processing device is used to switch the audio acquisition device from the original sound mode to the artificial intelligence mode when the original sound mode meets the first preset switching condition; wherein the original sound mode meets the first preset switching condition includes: the singing quality of the current singing voice is lower than the preset standard quality, or the song segment corresponding to the current singing voice is a specified song segment in the target song, or the audio processing device responds to the first mode switching operation triggered by the user.
[0123] The sending module is used in the artificial intelligence mode to simulate the user's singing of the target song at a subsequent stage using artificial intelligence technology to obtain an artificial intelligence singing voice, and to send the artificial intelligence singing voice to an audio processing device for playback; the subsequent stage is the stage after the current stage.
[0124] In one exemplary embodiment, the sending module is further configured to determine the timestamp corresponding to the subsequent progress of the target song based on the timestamp corresponding to the current singing of the target song; use artificial intelligence technology to simulate the user singing the song after the timestamp corresponding to the subsequent progress of the target song to obtain the artificial intelligence singing voice; send the artificial intelligence singing voice to the audio processing device for playback; the audio processing device is configured to play the artificial intelligence singing voice starting from the timestamp corresponding to the subsequent progress of the target song.
[0125] The modules in the aforementioned karaoke device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0126] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a karaoke method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0127] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0128] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0129] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0130] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0134] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A karaoke method, characterized in that, Applied to an audio processing device, the method includes: The current singing voice of a user singing a target song, captured by an audio acquisition device in original mode, is used as the current singing voice of the target song. Obtain the feedback detection result corresponding to the current singing voice, and determine the singing quality of the current singing voice based on the feedback detection result; the feedback detection result is used to characterize whether there is feedback in the current singing voice; or, obtain the singing score of the current singing voice, and determine the singing quality of the current singing voice based on the singing score; When the original sound mode meets the first preset switching condition, the audio acquisition device is switched from the original sound mode to the artificial intelligence mode; wherein, the first preset switching condition refers to information used to determine whether to switch the original sound mode to the artificial intelligence mode; the original sound mode meeting the first preset switching condition includes: the singing quality of the song being sung is lower than a preset standard quality; the artificial intelligence mode refers to a mode that uses artificial intelligence technology to simulate the user's singing voice, replaces the user's real voice, and plays it. The audio acquisition device uses artificial intelligence technology in the artificial intelligence mode to simulate the user singing the target song at a subsequent stage, and plays the artificial intelligence singing voice; wherein the subsequent stage is the stage after the current stage. In the artificial intelligence mode, the audio acquisition device acquires the current singing progress of the user singing the target song as the current singing voice of the target song. When the artificial intelligence mode meets the second preset switching condition, the artificial intelligence mode is switched to the original sound mode; wherein, the second preset switching condition refers to information used to determine whether to switch the artificial intelligence mode to the original sound mode; the artificial intelligence mode meeting the second preset switching condition includes: the singing quality of the currently sung song is higher than or equal to a preset standard quality; Play the current singing voice of the user singing the target song, and acquire and play the subsequent singing voice of the user singing the target song in the original sound mode, which was acquired by the audio acquisition device.
2. The method according to claim 1, characterized in that, The acquisition of the AI singing voice obtained by the audio acquisition device in the AI mode using AI technology to simulate the user singing the target song at subsequent stages includes: In response to switching the original sound mode to the artificial intelligence mode, a mode switching command is sent to the audio acquisition device; the mode switching command is used to control the audio acquisition device to send the artificial intelligence singing voice to the audio processing device. Receive the AI-generated singing voice sent by the audio acquisition device.
3. The method according to claim 1, characterized in that, The audio processing device and the audio acquisition device are integrated into the same audio device.
4. A karaoke method, characterized in that, Applied to an audio acquisition device, the method includes: In the original audio mode, the current singing progress of the user singing the target song is captured as the current singing voice of the target song; The current singing voice is sent to an audio processing device; the audio processing device is used to switch the audio acquisition device from the original sound mode to the artificial intelligence mode when the original sound mode meets a first preset switching condition; wherein, the first preset switching condition refers to information used to determine whether to switch the original sound mode to the artificial intelligence mode; the original sound mode meeting the first preset switching condition includes: the singing quality of the current singing voice is lower than a preset standard quality; the singing quality of the current singing voice is determined by the audio processing device acquiring the feedback detection result corresponding to the current singing voice and determining it based on the feedback detection result; or, it is determined by the audio processing device acquiring the singing score of the current singing voice and determining it based on the singing score; the feedback detection result is used to characterize whether there is feedback in the current singing voice; the artificial intelligence mode refers to a mode that uses artificial intelligence technology to simulate the user's singing voice to replace the user's real voice and play it; In the AI mode, AI technology is used to simulate the user's singing of the target song at a subsequent stage to obtain an AI singing voice, which is then sent to an audio processing device for playback; the subsequent stage refers to the stage after the current stage. In the AI mode, the current singing progress of the user singing the target song is collected as the current singing progress of the target song; The system sends the current vocal performance of the user singing the target song to an audio processing device. The audio processing device, when the artificial intelligence mode meets a second preset switching condition, switches the artificial intelligence mode to the original sound mode, plays the current vocal performance of the user singing the target song, and acquires and plays the subsequent vocal performance of the user singing the target song, acquired by the audio acquisition device in the original sound mode. The second preset switching condition refers to information used to determine whether to switch the artificial intelligence mode to the original sound mode. Meeting the second preset switching condition includes: the vocal quality of the current performance is higher than or equal to a preset standard quality.
5. The method according to claim 4, characterized in that, The process of using artificial intelligence technology to simulate the user's singing of the target song to obtain an AI-generated singing voice, and then sending the AI-generated singing voice to an audio processing device for playback, includes: Based on the timestamp corresponding to the current singing of the target song, determine the timestamp corresponding to the subsequent progress of the target song; The AI singing voice is obtained by simulating the user's singing of the target song after the timestamp corresponding to the subsequent progress of the song using artificial intelligence technology. The AI-generated singing voice is sent to the audio processing device for playback; the audio processing device is used to play the AI-generated singing voice starting from the timestamp corresponding to the subsequent progress of the target song.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for intelligently adjusting sound effects
CN109905806A
Singing synthesis method and device
CN111354332A
Microphone and audio system
CN206993333U