Song processing method, device, apparatus, and computer-readable storage medium

CN113703882BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110251029.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-08
Publication Date
2026-09-25
Estimated Expiration
2041-03-08

AI Technical Summary

Technical Problem

[0004]相关技术中,K歌应用提供了合唱功能,也即提供一个删除了部分原唱音频的视频文件,用户可以基于该视频文件,对删除的部分进行演唱,进而生成一个合唱音频文件;但通过上述方法实现合唱功能,用户只能够对固定的角色片段进行演唱,合唱方式单一、用户的自主选择性低

Benefits of technology

[0023]本申请实施例提供一种计算机设备,包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113703882B_ABST
    Figure CN113703882B_ABST
Patent Text Reader

Abstract

The application provides a song processing method and device, equipment and computer readable storage medium, and relates to the technical field of artificial intelligence; the method comprises the following steps: in response to a chorus instruction of a target object for a target song, a chorus performer selection interface including at least two original chorus performers of the target song is presented; based on the chorus performer selection interface, in response to a first selection operation for the original chorus performers, at least two chorus performers singing the target song are determined, and the at least two chorus performers include the target object and at least one target original chorus performer; in response to a chorus recording instruction triggered based on the at least two chorus performers, a song recording interface corresponding to the target song is presented, and first song data of the target object is recorded; and in response to a recording end instruction for the target song, a chorus media file including the first song data and second song data of the target original chorus performer is generated. Through the application, the selection of chorus performers can be realized, and the autonomous selection of users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a song processing method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] Artificial intelligence (AI) is a comprehensive discipline encompassing both hardware and software technologies. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning, with speech technology being a key technology in the field of artificial intelligence commands. Enabling computers to hear, see, speak, and feel represents the future direction of human-computer interaction.

[0003] As people's quality of life improves, karaoke apps are gradually becoming a part of people's daily lives, and their functions are becoming increasingly diverse.

[0004] In related technologies, karaoke applications provide a duet function, which is to provide a video file with some of the original audio deleted. Users can sing the deleted parts based on the video file to generate a duet audio file. However, by implementing the duet function in this way, users can only sing fixed parts, resulting in a limited variety of duet options and low user autonomy. Summary of the Invention

[0005] This application provides a song processing method, apparatus, device, and computer-readable storage medium, which enables the selection of chorus members and enhances the user's autonomy and choice.

[0006] The technical solution of this application embodiment is implemented as follows: This application provides a song processing method, including: In response to a chorus instruction from the target object for a target song, a chorus selection interface is presented, including at least two original chorus members of the target song. Based on the chorus selection interface, in response to a first selection operation for the original chorus, at least two chorus members are determined to sing the target song, the at least two chorus members including the target object and at least one target original chorus member; In response to a chorus recording instruction triggered by the at least two choruses, a song recording interface corresponding to the target song is presented, and the first song data of the target object is recorded; In response to a recording end command for the target song, a choral media file is generated that includes the first song data and the second song data of the target original chorus singers.

[0007] In the above scheme, recording the first song data of the target object includes: Extract the original song data of the target chorus from the original song data corresponding to the target song; During the recording of the target object's song data, the instrumental music of the target song is played, and During the singing portion of the target original chorus, play the song data of the target original chorus; Based on the playing accompaniment music and the song data of the target original chorus, the first song data of the target object is recorded.

[0008] In the above scheme, before presenting the chorus selection interface including at least two original choruses of the target song, the following steps are also included: Present a song search interface and display search function items within the song search interface; Received the input song information; In response to a search command for the song information triggered based on the search function item, songs that match the song information are presented; In response to a song selection operation triggered based on a song that matches the song information, the song corresponding to the selection operation is selected as the target song.

[0009] This application provides a song processing device, including: The presentation module is used to respond to the target object's chorus instruction for the target song and present a chorus selection interface including at least two original chorus singers of the target song. A determination module is configured to determine, based on the chorus selection interface and in response to a first selection operation for the original chorus, at least two chorus members who will sing the target song, wherein the at least two chorus members include the target object and at least one target original chorus member. The recording module is used to respond to a chorus recording instruction triggered based on the at least two choruses, present a song recording interface corresponding to the target song, and record the first song data of the target object; The generation module is used to generate a choral media file including the first song data and the second song data of the target original chorus singers in response to the recording end command for the target song.

[0010] In the above scheme, the determining module is further used to take the original chorus corresponding to the first selection operation as the target original chorus. The target original chorus and the target object are identified as at least two chorus members who sing the target song.

[0011] In the above scheme, the determining module is further configured to receive a chorus invitation instruction for the target song when the number of original choruses other than the target original chorus is at least two. In response to the chorus invitation instruction, a chorus invitation message is sent to invite at least one user object to be a chorus member of the target song.

[0012] In the above scheme, the determining module is further configured to determine the original chorus corresponding to the second selection operation in response to the second selection operation for the original chorus; Obtain the remaining original chorus members besides the target original chorus members and the original chorus members corresponding to the second selection operation; Send a chorus invitation message carrying the remaining original chorus members to invite at least one user object to sing the part corresponding to the remaining original chorus members.

[0013] In the above scheme, the determining module is further configured to receive and present acceptance invitation information of at least one user object, the acceptance invitation information being used to indicate the singing part selected by the at least one user object; The song content, excluding the singing part corresponding to the original target chorus and the singing part selected by at least one user object, shall be used as the target singing part to be sung by the target object. The recording module is also used to output prompt information during the recording of the target object's song data, so as to prompt the target object to sing the target singing part; Based on the prompt information, record the first song data of the target object.

[0014] In the above scheme, the determining module is also used to obtain the recorded song data of the user object when a user object accepts the chorus invitation; Based on the first song data, the user object's song data, and the target original chorus's second song data, a chorus media file corresponding to the target song is synthesized.

[0015] In the above scheme, the generation module is further used to extract the second song data of the target original chorus from the original audio file of the target song; Based on the first song data and the second song data, a chorus media file corresponding to the target song is synthesized.

[0016] In the above scheme, the generation module is further used to obtain the image file corresponding to the target song; Based on the first song data, the second song data, and the image file, video encoding is performed to obtain a chorus video file corresponding to the target song.

[0017] In the above scheme, the generation module is further used to extract the second song data of the target original chorus from the original song data corresponding to the target song; When it is determined that there is target song data in the original song data other than the first song data and the second song data, the timbre of the target original chorus singer is used to generate third song data corresponding to the target song data. Based on the first song data, the second song data, and the third song data, a chorus media file corresponding to the target song is synthesized.

[0018] In the above scheme, the generation module is further configured to present an invitation prompt when it detects that the chorus media file contains an unsung portion, so as to prompt at least one user object to sing the unsung portion.

[0019] In the above scheme, the generation module is also used to present the lyrics of the target song; In response to a lyrics selection operation triggered based on the lyrics of the target song, the selected lyrics are used as the target singing part to be sung by the target object; The recording module is also used to output prompt information during the recording of the target object's song data, so as to prompt the target object to sing the target singing part; Based on the prompt information, record the first song data of the target object.

[0020] In the above scheme, the generation module is also used to present an editing interface corresponding to the choral media file; In response to an editing operation triggered by the editing interface, one of the following parameters of the choral media file is adjusted: vocal volume, accompaniment volume, reverb mode, and equalization status.

[0021] In the above scheme, the recording module is further used to extract the original song data of the target chorus from the original song data corresponding to the target song; During the recording of the target object's song data, the instrumental music of the target song is played, and During the singing portion of the target original chorus, play the song data of the target original chorus; Based on the playing accompaniment music and the song data of the target original chorus, the first song data of the target object is recorded.

[0022] In the above solution, the presentation module is also used to present a song search interface and to present search function items in the song search interface; Received the input song information; In response to a search command for the song information triggered based on the search function item, songs that match the song information are presented; In response to a song selection operation triggered based on a song that matches the song information, the song corresponding to the selection operation is selected as the target song.

[0023] This application provides a computer device, including: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the song processing method provided in the embodiments of this application.

[0024] This application provides a computer-readable storage medium storing executable instructions for inducing a processor to execute and implement the song processing method provided in this application.

[0025] By applying the above embodiments, a chorus selection interface is presented, including at least two original choruses of the target song; based on the chorus selection interface, in response to a first selection operation for the original choruses, at least two choruses to sing the target song are determined; in response to a chorus recording instruction triggered based on the at least two choruses, a song recording interface corresponding to the target song is presented, and first song data of the target object is recorded; in response to a recording end instruction for the target song, a chorus media file including the first song data and second song data of the target original choruses is generated; thus, when the number of original choruses of the target song is at least two, the target object can independently select choruses to sing the target song with through the chorus selection interface, improving the user's autonomy and choice. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the interface for the choral recording process provided by the relevant technology; Figure 2 This is an optional architecture diagram of the song processing system 100 provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of the computer device 500 provided in the embodiments of this application; Figure 4 This is a flowchart illustrating the song processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the interface for the song search process provided in an embodiment of this application; Figure 6 This is a schematic diagram of a singer's homepage provided in an embodiment of this application; Figure 7This is a schematic diagram of the details page of the target song provided in the embodiments of this application; Figure 8 This is a schematic diagram of the chorus selection interface provided in an embodiment of this application; Figure 9 This is a schematic diagram of the song recording interface provided in an embodiment of this application; Figure 10 This is a schematic diagram of the singing section selection interface provided in an embodiment of this application; Figure 11 This is a schematic diagram of the interface for displaying invitation prompt information provided in an embodiment of this application; Figure 12 This is a schematic diagram of the editing interface provided in an embodiment of this application; Figure 13 This is a flowchart illustrating the audio material preparation stage provided in the embodiments of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0029] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0031] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0032] 1) Client: An application that runs on a terminal and provides various services, such as a video client or a music client.

[0033] 2) In response, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0034] Figure 1 This is a schematic diagram of the interface for the choral recording process provided by the relevant technology. See [link / reference]. Figure 1 First, when a user triggers the duet function, such as clicking the "Celebrity Duet" option, the terminal displays a song selection interface 101, allowing the user to choose the song they want to duet with. After the user selects a song, a corresponding duet recording interface 102 is displayed. During the recording process, the terminal plays a video file in which some of the original audio has been removed. The removed portion is the part the user needs to sing. The user's song data and video data are recorded based on this video. After recording, a duet video 103 is synthesized based on the video file, the recorded song data, and the video data. The user can then publish the duet video.

[0035] In the process of implementing the embodiments of this application, the applicant discovered that the chorus function in the related technology relies on video files uploaded by the platform. Some of the original audio has been deleted from the video files, and users can only select fixed character segments to sing. In this process, users cannot choose the part they want to sing, nor can they choose who to sing with, and the user's autonomy is very low.

[0036] Based on this, embodiments of this application provide a song processing method, apparatus, device, and computer-readable storage medium, which can enhance the user's autonomy and choice.

[0037] See Figure 2 , Figure 2 This is an optional architecture diagram of the song processing system 100 provided in the embodiments of this application. In order to support an exemplary application, the terminal (terminal 400-1 and terminal 400-2 are shown as examples) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0038] Server 200 is used to store song data for at least one original chorus member for each song; The terminal is configured to, in response to a chorus instruction from the target object for a target song, present a chorus selection interface including at least two original choruses of the target song; based on the chorus selection interface, in response to a first selection operation for the original choruses, determine at least two choruses to sing the target song, the at least two choruses including the target object and at least one target original chorus; and in response to a chorus recording instruction triggered based on the at least two choruses, send a request to the server 200 to obtain the second song data of the target original chorus. Server 200 is used to locate the second song data of the target original chorus singers and return it to the terminal; The terminal is used to present the song recording interface for the corresponding target song and record the first song data of the target object; in response to the recording end command for the target song, it generates a chorus media file based on the first song data and the second song data.

[0039] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals may be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, in-vehicle devices, smart TVs, etc., but are not limited to these.

[0040] See Figure 3 , Figure 3 This is a schematic diagram of the structure of the computer device 500 provided in the embodiments of this application. In practical applications, the computer device 500 can be... Figure 1 The terminal (such as 400-1) or server 200 in the middle, with computer equipment as Figure 2 Taking the terminal shown as an example, the computer device implementing the song processing method of the present application will be described. Figure 3 The computer device 500 shown includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the computer device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general labeled all buses as Bus System 540.

[0041] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0042] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0043] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.

[0044] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0045] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0046] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 553 is used to enable the presentation of information (e.g., user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., display screen, speaker, etc.) associated with user interface 530. The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.

[0047] In some embodiments, the song processing apparatus provided in this application can be implemented in software. Figure 3 A song processing device 555 stored in memory 550 is shown. It can be software in the form of programs and plug-ins, including the following software modules: presentation module 5551, determination module 5552, recording module 5553 and generation module 5554. These modules are logical and can therefore be arbitrarily combined or further split according to the functions they implement.

[0048] The functions of each module will be explained below.

[0049] In other embodiments, the song processing device provided in this application can be implemented in hardware. As an example, the song processing device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the song processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0050] The song processing method provided in this application will be described in conjunction with exemplary applications and implementations of the terminals provided in the embodiments of this application.

[0051] See Figure 4 , Figure 4 This is a flowchart illustrating the song processing method provided in the embodiments of this application, which will be combined with... Figure 4 The steps shown are explained.

[0052] Step 401: In response to the target object's chorus instruction for the target song, the terminal presents a chorus selection interface that includes at least two original chorus members of the target song.

[0053] In practice, the terminal is equipped with a client application, such as a karaoke app client, a music client, or an instant messaging client. This client provides a karaoke function, allowing users to record their own singing and generate corresponding audio files. The client also offers a duet function, allowing users to sing a song with the original singer and generate a corresponding media file.

[0054] Here, a chorus instruction for a target song can be triggered via the client. After receiving the chorus instruction, the terminal presents at least two original chorus members for the target song through the client. That is, the number of original chorus members for the target song is at least two. For example, the target song can be a duet or a song performed by a group.

[0055] In practical applications, the target audience needs to first select the song they want to sing along to before triggering a sing-along command for that song. In some embodiments, the terminal can present a song recommendation interface with recommended songs, allowing the target audience to select the song they want to sing along to from the recommendations. In another embodiment, the terminal can present a song search interface, allowing the target audience to independently search for the song they want to sing based on the search results and then select the target song.

[0056] In some embodiments, when the target object selects a target song based on the song search interface, the terminal presents the song search interface and displays a search function item in the song search interface; receives the input song information; in response to a search instruction for the song information triggered based on the search function item, presents songs that match the song information; in response to a song selection operation triggered based on the song that matches the song information, selects the song corresponding to the selection operation as the target song.

[0057] Here, song information can be the name of the singer, such as the group name or the original chorus name, or the song title itself. The search function can be a combination of a search box and a search button, or it can be a voice search button. In practical applications, when the search function is a combination of a search box and a search button, the user inputs song information into the search box. After inputting, the user triggers the search button, and the terminal receives a search command for the song information triggered by the search button. When the search function is a voice search button, the user triggers the voice search button and speaks the song information. The terminal collects the user's voice information; for example, the user presses the voice search button and speaks the song information while pressing. When the user releases the voice search button, the terminal performs speech recognition based on the collected voice information to obtain the song information and triggers a search command for that song.

[0058] As an example, Figure 5 This is a schematic diagram of the interface for the song search process provided in an embodiment of this application. See also: Figure 5 The system presents a song search interface, including a search box 501 and a search button 502. The user enters song information, such as a group name, in the search box 501 and then clicks the search button 502, triggering a search command for the song information. The terminal then displays multiple songs 503 that match the song information. For example, if the user enters the group name, the system displays multiple songs performed by that group, allowing the user to select a song as the target song from the displayed songs.

[0059] In some embodiments, the terminal may also display the artist's homepage (including individuals and groups) and the songs performed by that artist on that homepage, allowing users to select a target song based on the artist's homepage. In practice, the terminal may present multiple recommended artists for the user to choose from. After the user selects an artist, the terminal displays that artist's homepage; alternatively, the user may enter an artist's name on the aforementioned song search interface, and the terminal will display the homepage entry for the artist matching that name. Upon receiving a trigger operation, this homepage entry will display the corresponding artist's homepage.

[0060] As an example, see Figure 5 When the entered song information is a group name, the group's homepage entry 504 is also displayed. After the target clicks on the homepage entry 504, the group's artist homepage is displayed. Figure 6 This is a schematic diagram of a singer's homepage provided in an embodiment of this application. See also... Figure 6 The system displays the artist's homepage, which shows the artist's detailed information, including the 601 songs performed by the group. Users can select their target song based on the artist's homepage.

[0061] In some embodiments, after the target object selects the target song to sing along to, a sing-along instruction for the target song can be automatically triggered, for example, see [link to relevant documentation]. Figure 5 Based on the multiple songs presented, the target user clicks on one of the songs, and the clicked song is designated as the target song. This automatically triggers a duet instruction for the target song, and then presents a duet selection interface for the target song. In other embodiments, after the target user selects the target song to duet with, it is necessary to manually trigger the duet instruction for the target song. For example, after the user selects the target song, a duet function item is presented, and by triggering the duet function item, the duet instruction for the target song is triggered.

[0062] As an example, after the target audience selects the song they want to sing along to, a details page for the target song is displayed. Figure 7This is a schematic diagram of the details page of the target song provided in this application embodiment. See also: Figure 7 The details page displays detailed information about the target song, such as its name, artist, and a ranking of users who have sung the song (historical best solo performances). It also presents a duet function item 701 for the target song. When a user clicks on this duet function item 701, it triggers a duet command for the target song.

[0063] In some embodiments, after receiving an instruction for a target song, the terminal retrieves at least two original singers of the target song and presents a singer selection interface, displaying the original singers of the target song. When presenting at least two original singers of the target song, it can be either all or a subset of the original singers. For example, when the number of original singers of the target song is five, the user's preference for each singer can be determined based on user information, and the three most preferred singers can be presented. It should be noted that "original singers" refers to the performers of the original audio file of the target song.

[0064] In practice, when at least two original chorus members are presented on the chorus member selection interface, they can be presented in text form or in image form. There is no limitation on the presentation format of the original chorus members.

[0065] As an example, Figure 8 This is a schematic diagram of the chorus selection interface provided in the embodiment of this application. In the chorus selection interface, multiple original chorus singers 801 of the target song are presented. Here, the original chorus singers are presented in a combination of images and text, that is, the names and avatars of each original chorus singer are presented.

[0066] Step 402: Based on the chorus selection interface, in response to the first selection operation for the original chorus, determine at least two chorus members to sing the target song.

[0067] Among them, at least two choruses include the target object and at least one target original chorus. In actual implementation, the first selection operation can be used to select the target original chorus to sing with the target object, or it can be to select the role that the target object wants to sing, that is, the singing part corresponding to the original chorus selected by the target object is the part that the target object wants to sing. Then, among the original choruses that are not selected by the user, at least one original chorus is the target original chorus.

[0068] In practical applications, the first selection operation can be a click operation targeting the original chorus singers. For example, the original chorus singer corresponding to the click operation can be selected as the chosen original chorus singer. The first selection operation can also include a click operation targeting the original chorus singers and a click operation targeting the "confirm" function. That is, the user pre-selects the original chorus singers through a click operation. Here, the pre-selected original chorus singers can be presented in a way that distinguishes them from other original chorus singers. For example, the pre-selected original chorus singers can be highlighted. After clicking the "confirm" function, the pre-selected original chorus singers are confirmed as the final selected original chorus singers. In other words, the selected original chorus singers can be modified before clicking the "confirm" function. It should be noted that the triggering form of the first selection operation is not limited to the above two methods; other methods can also be used.

[0069] Taking the first selection operation, which includes clicking on the original chorus members and clicking on the selected function item, as an example, see [link to example]. Figure 8 In the chorus selection interface, not only are multiple original chorus singers 801 for the target song displayed, but also a confirmation function 802 is presented. Here, you can pre-select an original chorus singer by clicking on the name or avatar, such as selecting Qu xx. After the selection is completed, the pre-selected original chorus singers are displayed separately. At this time, you can click the confirmation function 802 to confirm the final selected original chorus singer.

[0070] In some embodiments, at least two choruses for singing the target song can be determined by: using the original chorus corresponding to the first selection operation as the target original chorus; and determining the target original chorus and the target object as at least two choruses for singing the target song.

[0071] In practice, when the first selection operation is used to select the target original chorus singer to sing with the target object, the original chorus singer corresponding to the first selection operation is obtained and used as the target original chorus singer. Therefore, the chorus singers performing the target song include the selected target original chorus singer and the target object. Here, the number of target original chorus singers can be one or more.

[0072] For example, if the original chorus of the target song includes original chorus A, original chorus B, and original chorus C, the target object can choose one or more of them as the target original chorus. If the original chorus corresponding to the first selection operation is original chorus A, then the at least two chorus singing the target song include at least the target object and original chorus A. If the original chorus corresponding to the first selection operation is original chorus A and original chorus B, then the at least two chorus singing the target song include at least the target object, original chorus A, and original chorus B.

[0073] In some embodiments, once a target original chorus is selected, the singing parts corresponding to the target original chorus are sung by the target original chorus; for the singing parts of original choruses other than the target original chorus, the target object can sing alone or the target object can sing together with other objects. These other objects can be other user objects or original choruses, such as the target original chorus object.

[0074] As an example, if the original chorus of the target song includes original chorus A, original chorus B, and original chorus C, and the target song consists of three sections, with original chorus A corresponding to the first section, original chorus B singing the second section, and original chorus C corresponding to the third section; based on the user's first selection operation, original chorus A corresponding to the first selection operation is selected as the target original chorus. Then, the first section is sung by original chorus A, and the second and third sections can both be sung by the target object, or the target object can sing only one of the second and third sections, and then the remaining section is sung by other objects.

[0075] In some embodiments, after the terminal determines the target original chorus and the target object as at least two chorus members to sing the target song, it may also receive a chorus invitation instruction for the target song when the number of original chorus members other than the target original chorus is at least two; in response to the chorus invitation instruction, it sends a chorus invitation message to invite at least one user object as a chorus member of the target song.

[0076] In practice, when there are at least two original chorus members besides the target original chorus member, the target object can invite one or more user objects to sing the parts sung by the original chorus members other than the target original chorus member. Here, the chorus invitation command can be triggered through a chorus invitation function item. For example, the terminal can display a chorus invitation function item, and the target object can click on the invitation function item to trigger the chorus invitation command.

[0077] In some embodiments, the chorus invitation information can be sent to a user object specified by the target object. That is, after receiving the chorus invitation instruction, the terminal presents a user object selection interface, which presents selectable user objects, such as user objects that have social relationships with the target object. The user can select one or more user objects as the recipient of the chorus invitation information based on the presented user objects, and the terminal sends the chorus invitation information to the recipient of the chorus invitation information.

[0078] In some embodiments, the chorus invitation information can be sent to any user of the current client. For example, the chorus invitation information can be presented on the client's recommendation interface so that all users can see the chorus invitation information when viewing the recommendation interface; or, the chorus invitation information can be pushed to the clients of all online users in the form of a system message.

[0079] It should be noted that when sending a chorus invitation, it can be sent based on the current client or by calling a third-party client. For example, if the current client is a karaoke application client, the terminal can call an instant messaging client to send the chorus invitation.

[0080] In practical applications, the chorus invitation message can carry the original chorus singers. These original chorus singers indicate the part the user wants to sing; that is, the part corresponding to the original chorus singer is the part the user wants to perform. There can be one or more original chorus singers. When only one original chorus singer is included, the invited user can only choose to sing the part corresponding to that singer. When multiple original chorus singers are included, the invited user can choose one singer's part to perform.

[0081] In some embodiments, before receiving a chorus invitation instruction for a target song, the terminal may also respond to a second selection operation for the original chorus singers, determine the original chorus singers corresponding to the second selection operation, obtain the remaining original chorus singers other than the target original chorus singer and the original chorus singers corresponding to the second selection operation, and accordingly, send chorus invitation information in the following manner: send chorus invitation information carrying the remaining original chorus singers to invite at least one user object to sing the singing part corresponding to the remaining original chorus singers.

[0082] In practice, before triggering the chorus invitation command for the target song, the target user can independently select the part they want to sing, which is the part sung by the original chorus corresponding to the second selection operation. This selected part is the part the target user wants to sing. Therefore, the rest of the song content, excluding the original chorus's part and the target part, needs to be sung by other user users. Based on this, a chorus invitation message can be generated to invite at least one user user to sing the remaining parts sung by the original chorus.

[0083] For example, if the original chorus of the target song includes original chorus A, original chorus B, and original chorus C, and the first selection operation corresponds to original chorus A and the second selection operation corresponds to original chorus B, then the singing part of original chorus C needs to be sung by other users. Therefore, a chorus invitation message can be generated based on original chorus C and sent to invite at least one user to sing the singing part corresponding to original chorus C.

[0084] In some embodiments, after sending the chorus invitation information, the terminal can also receive and present acceptance invitation information from at least one user object. The acceptance invitation information indicates the singing portion selected by the at least one user object. The song content excluding the singing portion corresponding to the original target chorus member and the singing portion selected by the at least one user object is taken as the target singing portion to be sung by the target object. Accordingly, the first song data of the target object can be recorded in the following manner: during the recording of the target object's song data, a prompt message is output to prompt the target object to sing the target singing portion; based on the prompt message, the first song data of the target object is recorded.

[0085] In practice, a chorus invitation can be sent first, allowing the recipient to select the part they wish to sing. The remaining part is then designated as the target part for the intended recipient. Here, the chorus invitation includes information about the original chorus members other than the target original chorus members, enabling the recipient to select the part to sing based on the invitation, such as the part corresponding to one or more original chorus members.

[0086] In practical applications, after a user selects the part they wish to sing, they can trigger an acceptance command for a duet invitation. This sends an acceptance message to the current terminal. Upon receiving the invitation, the terminal displays it to indicate that a user has accepted the duet invitation and informs the user of the selected part. The song content excluding the parts sung by the original duet singers and at least one part selected by the user constitutes the target part the user is to sing. During the recording process, based on this target part, prompts are output to guide the user in singing that part.

[0087] Step 403: In response to a chorus recording instruction triggered by at least two choruses, present the song recording interface for the corresponding target song and record the first song data of the target object.

[0088] In actual implementation, after receiving the chorus recording instruction, the terminal will present the corresponding target song recording interface. The target person can sing the target song based on the song recording interface, and the terminal will record the target person's first song data.

[0089] Here, to enhance the singing experience for the target audience, during the recording of the target audience's first song data, the terminal will display information such as the lyrics and pitch of the target song on the song recording interface to help the target audience sing the target song.

[0090] As an example, Figure 9 This is a schematic diagram of the song recording interface provided in an embodiment of this application. See also: Figure 9 In the song recording interface, the lyrics 901 and pitch 902 of the target song are displayed, and the target can sing based on the displayed lyrics 901 and pitch 902.

[0091] The terminal can also display multiple recording functions on the recording interface, such as pause, re-record, and complete, so that recording can be paused or resumed during the recording process.

[0092] In some embodiments, to prevent the target object from singing the wrong part during the chorus, the terminal can output prompt information during the recording of the target object's first song data to prompt the target object to sing the target singing part.

[0093] In practice, there are several ways to output prompts. For example, the song corresponding to the target singing part can be displayed differently; or a text reminder can be presented when the target singing part is reached.

[0094] In some embodiments, the first song data of the target object is recorded as follows: the song data of the target original chorus is extracted from the original song data corresponding to the target song; during the recording of the song data of the target object, the accompaniment music of the target song is played, and the song data of the target original chorus is played during the singing part of the target original chorus; the first song data of the target object is recorded based on the played accompaniment music and the song data of the target original chorus.

[0095] In practice, during the recording of the target subject's first song data, the accompaniment music and the original chorus's song data are played during the original chorus's singing part; during the target subject's intended singing part, only the accompaniment music is played. This allows the target subject to be prompted to sing the intended part through audio output, providing a better singing experience and giving them the feeling of singing a duet with the original chorus.

[0096] In some embodiments, after determining at least two choruses singing the target song, the terminal may also present the lyrics of the target song; in response to a lyrics selection operation triggered based on the lyrics of the target song, the selected lyrics are used as the target singing part to be sung by the target object; correspondingly, the terminal may also record the first song data of the target object in the following manner: during the recording of the song data of the target object, output prompt information to prompt the target object to sing the target singing part; based on the prompt information, record the first song data of the target object.

[0097] In practice, after determining at least two choruses for the target song, the target audience can further select the target singing part. Here, the selection is no longer based on the singing role, but on the lyrics. That is, the target audience can choose a part of the singing part corresponding to a certain original chorus, rather than necessarily choosing the entire singing part corresponding to that original chorus. In this way, the target audience has greater autonomy in selecting the target singing part.

[0098] For example, Figure 10 This is a schematic diagram of the singing section selection interface provided in an embodiment of this application. See also... Figure 10 The singing selection interface displays the lyrics of the target song. Each line of lyrics is preceded by a selection option 1001. Clicking this option allows you to select the corresponding lyrics. After making the selection, clicking the OK button 1002 confirms that the lyrics selected by the user are the target singing part that the user wants to sing.

[0099] In some embodiments, for song content that the target object has not selected and that does not belong to the original target chorus, at least one user object may be invited to sing.

[0100] Step 404: In response to the recording end command for the target song, generate a chorus media file including the first song data and the second song data of the target original chorus singers.

[0101] In practice, the recording end command for the target song can be triggered automatically after the entire song has been recorded, or it can be triggered by the user. For example, see [link to relevant documentation]. Figure 9 User click Figure 9 The completion icon 903 in the recording prompts the end-of-recording command.

[0102] In some embodiments, a chorus media file including first song data and second song data of the target original chorus singers can be generated by: extracting the second song data of the target original chorus singers from the original audio file of the target song; and synthesizing a chorus media file corresponding to the target song based on the first song data and the second song data.

[0103] In practice, the original audio file includes accompaniment data and song data for each original chorus member. It is necessary to extract the second song data for the target original chorus member from the original audio file. This can be achieved as follows: First, the original audio file of the target song is converted into a spectrogram. Image recognition is then performed on the spectrogram to determine the spectrograms for the corresponding vocal parts and the corresponding accompaniment music. The spectrograms for the corresponding vocal parts are then converted into audio files, generating the corresponding vocal audio file. Similarly, the spectrograms for the corresponding accompaniment music are converted into audio files, generating the corresponding accompaniment audio file. The song data within the vocal audio files is then separated to obtain individual audio track files for each original chorus member of the target song. The song data within these individual audio track files is the second song data. Finally, the first song data and the second song data are combined to obtain the corresponding choral media file for the target song.

[0104] In practical applications, accompaniment data can also be added, that is, the first song data, the second song data and the accompaniment music data can be combined to obtain the chorus media file of the corresponding target song.

[0105] In some embodiments, a chorus media file including first song data and second song data of the target original chorus singer can be generated in the following manner: when a user object accepts a chorus invitation, the recorded song data of the user object is obtained; based on the first song data, the user object's song data, and the second song data of the target original chorus singer, a chorus media file corresponding to the target song is synthesized.

[0106] In practice, when other users are invited to sing along, the generated chorus media file processes the first song data, the second song data, and the user's song data. Based on this, the first song data, the user's song data, and the target original chorus's second song data are combined to obtain the corresponding target song's chorus media file.

[0107] The song data of the user object is recorded by the user object through its own terminal and then sent to the current terminal.

[0108] In practical applications, user objects can record songs simultaneously with target objects, or they can record songs at any time before generating the chorus media file.

[0109] In some embodiments, a chorus media file including first song data and second song data of the target original chorus singers can be generated by: acquiring an image file of the corresponding target song; and performing video encoding based on the first song data, the second song data, and the image file to obtain a chorus video file of the corresponding target song.

[0110] In practice, the synthesized media file can be not only a synthesized audio file, but also a synthesized video file. The audio part of the synthesized video file is synthesized based on the first song data and the second song data, and the image part of the video file is obtained based on the image file of the target song. Then, based on the first song data, the second song data, and the image file, video encoding is performed to obtain the chorus video file of the corresponding target song.

[0111] Here, the image file corresponding to the target song can be either an image file or a video file of the target song.

[0112] In some embodiments, the image portion of the video file can also be synthesized based on the image file of the target song and the image file of the corresponding target object. For example, while recording the first song data of the target object, the video file of the target object can be recorded, and the image file of the target song can be synthesized with the video file of the target object to obtain the image portion of the chorus video file.

[0113] In some embodiments, a choral media file including first song data and second song data of the target original chorus singer can be generated in the following manner: extracting the second song data of the target original chorus singer from the original song data corresponding to the target song; when it is determined that there is target song data other than the first song data and the second song data in the original song data, using the timbre of the target original chorus singer to generate third song data corresponding to the target song data; and synthesizing a choral media file corresponding to the target song based on the first song data, the second song data, and the third song data.

[0114] In practice, for the target song data, artificial intelligence algorithms can be used to extract the timbre of the original target chorus singer based on the audio data of the original chorus singer, and then use the timbre of the original target chorus singer to generate the corresponding third song data. The third song data here is audibly sung by the original target chorus singer.

[0115] In some embodiments, after generating a chorus media file that includes first song data and second song data of the target original chorus singers, an invitation message may be presented when an unsung portion is detected in the chorus media file, prompting an invitation to at least one user to sing the unsung portion.

[0116] In practice, the presence of unsung parts in the chorus media file can be determined by comparing the original vocal parts sung by the target singers with the target user's chosen vocal parts. Alternatively, the vocal parts within the chorus media file can be analyzed to determine if unsung parts are included. If unsung parts are found, the target user can be prompted to invite at least one user to sing those parts.

[0117] As an example, Figure 11 This is a schematic diagram of the invitation prompt information presentation interface provided in an embodiment of this application. See also: Figure 11 The system displays an invitation message 1101 to inform the target user whether the chorus media file contains any unsung parts, and prompts the target user to invite at least one user to sing the unsung parts.

[0118] In practical applications, the target object can trigger a command to send a chorus invitation message based on this invitation prompt. After receiving the sending command, the terminal sends the chorus invitation message to invite at least one user object to sing the unsung parts. For example, see... Figure 11 When the user clicks the invitation button 1102, the terminal sends a chorus invitation message.

[0119] In some embodiments, after generating a chorus media file including first song data and second song data of the target original chorus singers, the terminal may also present an editing interface for the corresponding chorus media file; in response to an editing operation triggered based on the editing interface, one of the following parameters of the chorus media file may be adjusted: vocal volume, accompaniment volume, reverb mode, and equalization status.

[0120] In practice, after generating the choral media file, the target audience can also adjust the vocal volume, accompaniment volume, reverb mode, equalization status, etc. of the choral media file.

[0121] Figure 12 This is a schematic diagram of the editing interface provided in an embodiment of this application. See also: Figure 12 The editing interface is presented, which includes controls for adjusting the volume of the vocals (1201) and the volume of the accompaniment (1202). The target object can adjust the corresponding parameters based on these controls.

[0122] By applying the above embodiments, a chorus selection interface is presented, including at least two original choruses of the target song; based on the chorus selection interface, in response to a first selection operation for the original choruses, at least two choruses to sing the target song are determined; in response to a chorus recording instruction triggered based on the at least two choruses, a song recording interface corresponding to the target song is presented, and first song data of the target object is recorded; in response to a recording end instruction for the target song, a chorus media file including the first song data and second song data of the target original choruses is generated; thus, when the number of original choruses of the target song is at least two, the target object can independently select choruses to sing the target song with through the chorus selection interface, improving the user's autonomy and choice.

[0123] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0124] In practice, the terminal first presents a song search interface. The target user searches for and finds the song they want to sing along to. The original singers for this song must be at least two. Then, the user triggers a sing-along command for the target song. The terminal then presents a singer selection interface, including the original singers of the target song. The user selects at least one of the original singers as the target original singer for the song. After receiving the user's selection, a performance selection interface is presented. The user selects the desired performance part. Next, the user triggers a recording command. Upon receiving the recording command, the terminal records the target user's song data and generates a choral media file. After generating the choral media file, the terminal can also present an editing interface, allowing the target user to adjust vocal volume, accompaniment volume, reverb mode, and equalization settings.

[0125] In practical applications, after the terminal displays the song search interface, users can enter song information on the song search interface to search for songs they want to sing together. The song information here can be the singer (such as the group name or the original singer's name) or the song title.

[0126] As an example, see Figure 5 The system presents a song search interface, including a search box 501 and a search button 502. Users can input song information, such as a group name, into the search input box. Upon receiving a search command for the song information, the system displays multiple songs that match the information. For example, if a singer's name is entered, the system displays songs 503 sung by that singer.

[0127] See here. Figure 5 When the entered song information is a group name, a 504 error message appears indicating that group's homepage entry. Clicking this entry displays the artist's homepage for that group. (See also...) Figure 6 The system displays the artist's homepage, which shows the artist's detailed information, including the 601 songs performed by the group. Users can select their target song based on the artist's homepage.

[0128] After selecting a target song, users can either directly trigger a duet command for that song or view the song's details page, which they can then use to trigger the duet command.

[0129] For example, see Figure 7 The details page displays detailed information about the target song, such as its name, artist, and a ranking of users who have sung the song (historical best solo performances). It also presents a duet function item 701 for the target song. When a user clicks on this duet function item 701, it triggers a duet command for the target song.

[0130] In practical applications, after receiving a singing instruction for a target song, the terminal redirects to the chorus selection interface, for example, see... Figure 8 In the chorus selection interface, multiple original chorus singers 801 for the target song are presented. The names and avatars of each original chorus singer are displayed here. Users can pre-select original chorus singers by clicking on the name or avatar. After the selection is completed, the user can click the "Confirm" function 802 to confirm the final selected target original chorus singer.

[0131] After the user selects the target original chorus, the terminal redirects to the singing part selection interface. Here, the lyrics of the target song are displayed on the singing part selection interface. Based on the displayed lyrics, the user selects the target singing part to sing, that is, which lines of lyrics to sing.

[0132] For example, see Figure 10 The singing selection interface displays the lyrics of the target song. Each line of lyrics is preceded by a selection option 1001. Clicking this option allows you to select the corresponding lyrics. After making the selection, clicking the OK button 1002 confirms that the lyrics selected by the user are the target singing part that the user wants to sing.

[0133] Here, after selecting the target singing section, a chorus recording command for the target song can be automatically triggered. The terminal displays a chorus recording interface, records the user's singing data, and generates a chorus media file upon completion. See also Figure 9The lyrics of the target song are displayed in the chorus recording interface, and users can sing according to the displayed lyrics.

[0134] This could be a confirmation of completion after the entire song has been recorded, or a confirmation of completion after receiving a user-triggered end-of-recording command. For example, if the user clicks... Figure 10 The completion icon 1001 in the record triggers the end-of-recording command.

[0135] Based on the above description of the song processing method according to the embodiments of this application, the technical implementation of the song processing method according to the embodiments of this application will be described below. In actual implementation, the technical implementation of the song processing method according to the embodiments of this application includes three parts: audio material preparation stage, recording stage, and synthesis stage.

[0136] First, let's explain the audio material preparation stage.

[0137] Figure 13 This is a flowchart illustrating the audio material preparation stage provided in an embodiment of this application. See also... Figure 13 The audio material preparation stage includes: Step 1301: Convert the original audio file of the target song into a spectrogram.

[0138] Step 1302: Based on the spectrogram, generate the audio file corresponding to the vocal part and the audio file corresponding to the accompaniment music.

[0139] Here, a convolutional neural network is used to perform image recognition on the spectrogram to determine the spectrogram of the corresponding human voice part and the spectrogram of the corresponding accompaniment music. The spectrogram of the corresponding human voice part is converted into audio to generate an audio file of the corresponding human voice part; the spectrogram of the corresponding accompaniment music is converted into audio to generate an audio file of the corresponding accompaniment music.

[0140] Step 1303: Convert the audio file corresponding to the human voice part into a PCM format audio file.

[0141] Step 1304: Divide the PCM format audio file into several speech units according to the preset step size and preset cutting length.

[0142] The preset step size is less than the preset cutting length.

[0143] Step 1305: Extract speech features from each speech unit sequentially.

[0144] Here, speech features include: left and right channel balance, volume, duration, intensity, pitch, and interval.

[0145] Step 1306: Obtain the matching values ​​of speech feature parameters between speech units.

[0146] Step 1307: Determine whether the matching value is higher than the preset threshold. If yes, proceed to step 1008; otherwise, do not perform any processing.

[0147] Step 1308: Save the two speech units in the same audio file in sequence.

[0148] Step 1309: Separate all voice units within the same audio file into individual audio track files corresponding to each original chorus singer.

[0149] Step 1310: Upload multiple solo audio track files and accompaniment music audio files to the platform.

[0150] Next, the audio recording stage will be explained.

[0151] In practice, after uploading the audio files to the platform, each original chorus artist's individual audio track file needs to be tagged to identify the original chorus artist corresponding to that track. Here, after the user selects a target original chorus artist, the corresponding individual audio track file for that artist is filtered out and used when synthesizing the choral audio file.

[0152] Before recording audio, the recorder needs to be initialized; during audio recording, recording can be paused and resumed; when recording is finally completed, the buffered data obtained from the recorder is stored in a PCM file to obtain the user's song data.

[0153] Finally, the synthesis stage will be explained.

[0154] Here, a choral audio file is synthesized based on the user's recorded song data, the accompaniment music audio file, and the individual audio track files of the original target chorus singers. After synthesizing the choral audio file, the user can choose to save it locally or upload it directly. The synthesized choral audio file saved locally can also be uploaded at any time.

[0155] In some embodiments, this application can also generate a chorus video file. Here, the synthesis method of the audio part in the chorus video file is the same as the synthesis method of the chorus audio file described above. The image part in the synthesized video file can be obtained by acquiring the video file of the corresponding target song and recording the user's video file, and then synthesizing the image in the target song's video file with the image in the user's video file; or it can be obtained by acquiring the image of the target song and the user's image, and then synthesizing the image of the target song with the user's image.

[0156] In some embodiments, the target song in this application can also be a song from a live concert, that is, the chorus function can be implemented based on a song from a live concert.

[0157] By applying the above embodiments, users can independently select the original chorus members to sing with, as well as the parts to be sung, thus enhancing user autonomy and choice.

[0158] The following description continues to illustrate the exemplary structure of the song processing device 555 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 3 As shown, the software modules stored in the song processing device 555 in the memory 550 may include: The presentation module is used to respond to the target object's chorus instruction for the target song and present a chorus selection interface including at least two original chorus singers of the target song. A determination module is configured to determine, based on the chorus selection interface and in response to a first selection operation for the original chorus, at least two chorus members who will sing the target song, wherein the at least two chorus members include the target object and at least one target original chorus member. The recording module is used to respond to a chorus recording instruction triggered based on the at least two choruses, present a song recording interface corresponding to the target song, and record the first song data of the target object; The generation module is used to generate a choral media file including the first song data and the second song data of the target original chorus singers in response to the recording end command for the target song.

[0159] In some instances, the determining module is further configured to use the original chorus corresponding to the first selection operation as the target original chorus. The target original chorus and the target object are identified as at least two chorus members who sing the target song.

[0160] In some instances, the determining module is also configured to receive a chorus invitation instruction for the target song when the number of original choruses other than the target original chorus is at least two. In response to the chorus invitation instruction, a chorus invitation message is sent to invite at least one user object to be a chorus member of the target song.

[0161] In some instances, the determining module is further configured to, in response to a second selection operation for the original chorus, determine the original chorus corresponding to the second selection operation; Obtain the remaining original chorus members besides the target original chorus members and the original chorus members corresponding to the second selection operation; Send a chorus invitation message carrying the remaining original chorus members to invite at least one user object to sing the part corresponding to the remaining original chorus members.

[0162] In some instances, the determining module is further configured to receive and present acceptance invitation information for at least one user object, the acceptance invitation information being used to indicate the singing portion selected by the at least one user object; The song content, excluding the singing part corresponding to the original target chorus and the singing part selected by at least one user object, shall be used as the target singing part to be sung by the target object. The recording module is also used to output prompt information during the recording of the target object's song data, so as to prompt the target object to sing the target singing part; Based on the prompt information, record the first song data of the target object.

[0163] In some instances, the determining module is also used to obtain the recorded song data of the user object when a user object accepts a chorus invitation; Based on the first song data, the user object's song data, and the target original chorus's second song data, a chorus media file corresponding to the target song is synthesized.

[0164] In some instances, the generation module is also used to extract second song data of the target original chorus from the original audio file of the target song; Based on the first song data and the second song data, a chorus media file corresponding to the target song is synthesized.

[0165] In some instances, the generation module is also used to obtain an image file corresponding to the target song; Based on the first song data, the second song data, and the image file, video encoding is performed to obtain a chorus video file corresponding to the target song.

[0166] In some instances, the generation module is also used to extract the second song data of the target original chorus from the original song data corresponding to the target song; When it is determined that there is target song data in the original song data other than the first song data and the second song data, the timbre of the target original chorus singer is used to generate third song data corresponding to the target song data. Based on the first song data, the second song data, and the third song data, a chorus media file corresponding to the target song is synthesized.

[0167] In some instances, the generation module is also configured to present an invitation message when it detects that the chorus media file contains an unsung portion, prompting at least one user object to sing the unsung portion.

[0168] In some instances, the generation module is also used to present the lyrics of the target song; In response to a lyrics selection operation triggered based on the lyrics of the target song, the selected lyrics are used as the target singing part to be sung by the target object; The recording module is also used to output prompt information during the recording of the target object's song data, so as to prompt the target object to sing the target singing part; Based on the prompt information, record the first song data of the target object.

[0169] In some instances, the generation module is also used to present an editing interface corresponding to the choral media file; In response to an editing operation triggered by the editing interface, one of the following parameters of the choral media file is adjusted: vocal volume, accompaniment volume, reverb mode, and equalization status.

[0170] In some instances, the recording module is also used to extract the original song data of the target chorus from the original song data corresponding to the target song; During the recording of the target object's song data, the instrumental music of the target song is played, and During the singing portion of the target original chorus, play the song data of the target original chorus; Based on the playing accompaniment music and the song data of the target original chorus, the first song data of the target object is recorded.

[0171] In some instances, the presentation module is also used to present a song search interface and to present search function items in the song search interface; Received the input song information; In response to a search command for the song information triggered based on the search function item, songs that match the song information are presented; In response to a song selection operation triggered based on a song that matches the song information, the song corresponding to the selection operation is selected as the target song.

[0172] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the song processing method described above in this application.

[0173] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to perform the method provided in this application, for example... Figure 4 The method shown.

[0174] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0175] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0176] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0177] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0178] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A song processing method, characterized in that, The method includes: In response to a chorus instruction from the target object for a target song, a chorus selection interface is presented, which includes at least two original chorus singers of the target song, the original chorus singers being the performers of the original audio file of the target song. Based on the chorus selection interface, in response to a first selection operation for the original chorus, at least two chorus members are determined to sing the target song, the at least two chorus members including the target object and at least one target original chorus member; Present the lyrics of the target song; in response to a lyrics selection operation triggered based on the lyrics of the target song, use the selected lyrics as the target singing part to be sung by the target object; In response to a chorus recording instruction triggered by the at least two choruses, a song recording interface corresponding to the target song is presented. During the recording of the target object's song data, a prompt message is output to prompt the target object to sing the target singing part. Based on the prompt message, the first song data of the target object is recorded. In response to a recording end command for the target song, extract the second song data of the target original chorus from the original audio file of the target song; Based on the first song data and the second song data, a chorus media file corresponding to the target song is synthesized.

2. The method as described in claim 1, characterized in that, The determination of at least two choruses to sing the target song includes: The original chorus corresponding to the first selection operation is taken as the target original chorus. The target original chorus and the target object are identified as at least two chorus members who sing the target song.

3. The method as described in claim 2, characterized in that, After determining that the target original chorus members and the target object are at least two chorus members singing the target song, the method further includes: When there are at least two original chorus members other than the target original chorus member, a chorus invitation instruction for the target song is received. In response to the chorus invitation instruction, a chorus invitation message is sent to invite at least one user object to be a chorus member of the target song.

4. The method as described in claim 3, characterized in that, Before receiving the chorus invitation instruction for the target song, the method further includes: In response to a second selection operation for the original chorus, the original chorus corresponding to the second selection operation is determined; Obtain the remaining original chorus members besides the target original chorus members and the original chorus members corresponding to the second selection operation; The sending of the chorus invitation information includes: Send a chorus invitation message carrying the remaining original chorus members to invite at least one user object to sing the part corresponding to the remaining original chorus members.

5. The method as described in claim 3, characterized in that, After sending the chorus invitation information, the method further includes: Receive and present an invitation acceptance message for at least one user object, the invitation acceptance message being used to indicate the singing part selected by the at least one user object; The song content, excluding the singing part corresponding to the original target chorus and the singing part selected by at least one user object, shall be used as the target singing part to be sung by the target object. The first song data recorded for the target object includes: During the recording of the target object's song data, a prompt message is output to prompt the target object to sing the target singing part; Based on the prompt information, record the first song data of the target object.

6. The method as described in claim 3, characterized in that, The step of synthesizing a chorus media file corresponding to the target song based on the first song data and the second song data includes: When a user object accepts a chorus invitation, the recorded song data of that user object is obtained; Based on the first song data, the user object's song data, and the target original chorus's second song data, a chorus media file corresponding to the target song is synthesized.

7. The method as described in claim 1, characterized in that, The step of synthesizing a chorus media file corresponding to the target song based on the first song data and the second song data includes: Obtain the image file corresponding to the target song; Based on the first song data, the second song data, and the image file, video encoding is performed to obtain a chorus video file corresponding to the target song.

8. The method as described in claim 1, characterized in that, The step of synthesizing a chorus media file corresponding to the target song based on the first song data and the second song data includes: When it is determined that there is target song data in the original song data other than the first song data and the second song data, the timbre of the target original chorus singer is used to generate third song data corresponding to the target song data. Based on the first song data, the second song data, and the third song data, a chorus media file corresponding to the target song is synthesized.

9. The method as described in claim 1, characterized in that, After synthesizing the chorus media file corresponding to the target song based on the first song data and the second song data, the method further includes: When the chorus media file is detected to contain an unsung portion, an invitation message is displayed to prompt at least one user to sing the unsung portion.

10. The method as described in claim 1, characterized in that, After synthesizing the chorus media file corresponding to the target song based on the first song data and the second song data, the method further includes: Presents the editing interface corresponding to the choral media file; In response to an editing operation triggered by the editing interface, one of the following parameters of the choral media file is adjusted: vocal volume, accompaniment volume, reverb mode, and equalization status.

11. A song processing device, characterized in that, The device includes: The presentation module is used to respond to the target object's chorus instruction for the target song and present a chorus selection interface including at least two original chorus singers of the target song, wherein the original chorus singers are the singers of the original audio file of the target song. A determination module is configured to determine, based on the chorus selection interface and in response to a first selection operation for the original chorus, at least two chorus members who will sing the target song, wherein the at least two chorus members include the target object and at least one target original chorus member. A generation module is used to present the lyrics of the target song; in response to a lyrics selection operation triggered based on the lyrics of the target song, the selected lyrics are used as the target singing part to be sung by the target object; The recording module is configured to respond to a chorus recording instruction triggered by the at least two choruses, present a song recording interface corresponding to the target song, output prompt information during the recording of the target object's song data to prompt the target object to sing the target singing part, and record the first song data of the target object based on the prompt information. The generation module is further configured to, in response to a recording end command for the target song, extract the second song data of the target original chorus from the original audio file of the target song; and synthesize a chorus media file corresponding to the target song based on the first song data and the second song data.

12. The apparatus according to claim 11, characterized in that, The determining module is further configured to use the original chorus corresponding to the first selection operation as the target original chorus. The target original chorus and the target object are identified as at least two chorus members who sing the target song.

13. The apparatus according to claim 12, characterized in that, The determining module is further configured to receive a chorus invitation instruction for the target song when the number of original choruses other than the target original chorus is at least two. In response to the chorus invitation instruction, a chorus invitation message is sent to invite at least one user object to be a chorus member of the target song.

14. A computer device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the song processing method according to any one of claims 1 to 10.

15. A computer-readable storage medium, characterized in that, It stores executable instructions for implementing the song processing method according to any one of claims 1 to 10 when executed by a processor.

16. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the song processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • K song chorus recording method

    CN108109652A

  • Chorus file generation method, device and equipment and computer readable storage medium

    CN112130727A