Method for processing virtual concert, processing apparatus, electronic device, and computer program
The method for processing virtual concerts addresses the limitation of existing technologies by allowing users to create and share virtual concerts of specific singers, using voice conversion to maintain the singer's voice quality, thereby enhancing emotional connection and entertainment options.
Patent Information
- Application Number
- JP2024515164
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-22
- Filing Date
- 2022-09-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-28
AI Technical Summary
Existing technologies cannot produce or hold virtual concerts for specific singers, limiting user interaction and entertainment options.
A method for processing virtual concerts that allows users to create and hold virtual concerts of a target singer by mimicking their songs, using voice conversion technology to maintain the singer's voice quality.
Enables users to produce and share virtual concerts of specific singers, enhancing emotional connection, entertainment options, and information diversification, while simplifying the sharing process and improving interaction efficiency.
Smart Images

Figure 0007696498000009 
Figure 0007696498000010 
Figure 0007696498000011
Abstract
Description
Technical Field
[0001] This application claims the priority of the Chinese patent application with the application number 202111386719.X and the filing date of November 22, 2021, and all the content of the Chinese patent application is incorporated into this application for reference.
[0002] This application relates to computer technology and voice technology, and particularly to a method for processing virtual concerts, Processing devices, Electronic machines apparatus and computer programs to .
Background Art
[0003] With the maturity of voice technology, much exploration and research have been carried out on the development and application of voice technology. In the music field, imitating the singing of professional and charismatic singers has become the exploration goal. For example, after recording a song, users can add echo or perform various personalized sound modifications ("voice conversion") to enjoyably participate in activities such as recording, publishing, and sharing the song even if they cannot sing.
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in related technologies, it can only provide users with the above-mentioned simple and random singing, and cannot provide the production or holding of a virtual concert of a specific singer.
Means for Solving the Problems
[0005] The method for processing virtual concerts provided in the embodiments of this application, process devices, electronic machines apparatus and computer programs to can enable users to produce or hold a virtual concert of a target singer.
[0006] The technical solution of the embodiment of the present application is realized as follows.
[0007] The method for processing a virtual concert provided in the embodiment of the present application is a method for processing a virtual concert executed by an electronic device, including: receiving a concert production instruction for a target singer in response to a trigger operation of a current target person on a presented concert entrance; creating a concert hall for mimicking and singing the songs of the target singer in the virtual concert in response to the concert production instruction; collecting singing content in which the current target person mimics and sings the songs of the target singer, and playing the singing content through the concert hall. do. In response to the trigger operation of the current target person with respect to the presented concert entrance, the step of receiving the concert production instruction for the target singer includes presenting a singer selection interface including at least one candidate singer in response to the trigger operation with respect to the presented concert entrance, and receiving the concert production instruction for the target singer in response to a selection operation for the target singer among the at least one candidate singer. The received concert production instruction is the concert production instruction for the target singer who has the production qualification for the current target person to produce the virtual concert. The singing content is 、 used for playing on the terminal of each target person in the concert hall.
[0008] The apparatus for processing a virtual concert provided in the embodiment of the present application includes: a command receiving module that receives a concert production instruction for a target singer in response to a trigger operation of a current target person on a presented concert entrance; a room creating module that creates a concert hall for mimicking and singing the songs of the target singer in the virtual concert in response to the concert production instruction; and a singing playback module that collects singing content in which the current target person mimics and sings the songs of the target singer, and plays the singing content through the concert hall. do. The instruction receiving module presents a singer selection interface including at least one candidate singer in response to the trigger operation with respect to the presented concert entrance, and receives the concert production instruction for the target singer in response to a selection operation for the target singer among the at least one candidate singer. The received concert production instruction is the concert production instruction for the target singer who has the production qualification for the current target person to produce the virtual concert, The singing content is 、 used for playing on the terminal of each target person in the concert hall.
[0009] The electronic device provided in the embodiment of the present application includes: a memory that stores computer-executable instructions; and a processor that, when executing the computer-executable instructions stored in the memory, realizes the method for processing a virtual concert provided in the embodiment of the present application.
[0010] The computer-readable storage medium provided in the embodiments of the present application stores executable instructions that, when executed by a processor, implement the method for processing a virtual concert provided in the embodiments of the present application.
[0011] The computer program provided in the embodiments of the present application is The method for processing a virtual concert provided in the embodiments of the present application to cause the computer Implement to do 。
Advantages of the Invention
[0012] The embodiments of the present application have the following beneficial effects. That is, according to the embodiments of the present application, the current target person creates a concert room for the target singer through the concert entrance, sings the songs of the target singer in the concert room, and allows the target person in the concert room to watch online, thereby realizing the reproduction of the concert of the target singer. Such a performance method can better convey the emotions of the target singer, provide more entertainment options for users, and meet the increasing demand for information diversification of users. In addition, since the created concert room corresponds to the target singer, the target person who enters this concert room can continuously enjoy many songs of the target singer, and the current target person can continuously share the songs of the target singer, thereby improving the sharing efficiency of the songs to a specific target person. Furthermore, compared with the point-to-point song sharing method in the related art, the user does not need to repeatedly execute the song sharing operation. When the song to be shared is a plurality of songs of a specific singer, the process of sharing the plurality of songs is simplified, and the man-machine interaction efficiency can be improved.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Embodiments for Carrying Out the Invention
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in combination with the drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0015] In the following description, the description related to "some embodiments" means that they are subsets of all possible embodiments. However, "some embodiments" may be the same subset or different subsets of all possible embodiments, and may also be combined with each other if they do not conflict.
[0016] In the following description, the description related to the terms "first / second..." only distinguishes similar objects and does not indicate a specific order for the objects. "First / second..." can be rearranged in a specific order or sequence before and after if permitted, and the embodiments of the present application described herein can be implemented in an order other than the order illustrated or described herein.
[0017] Unless otherwise defined, all technical and scientific terms used in the text have the same meaning as commonly understood by those skilled in the art. The terms used in the text are only for the purpose of explaining the embodiments of the present application and are not intended to limit the present application.
[0018] Before explaining the embodiments of the present application in detail, the names and terms related to the embodiments of the present application will be explained. The following interpretations apply to the names and terms related to the embodiments of the present application.
[0019] 1) A client is an application program that operates on a terminal to provide various services. For example, an instant messaging client, a video playback client, a live streaming client, a learning client, a singing client, etc.
[0020] 2) "Respond to ···" means to respond to the conditions or states on which the operations to be executed depend. When the dependent conditions or states are satisfied, one or more operations may be executed in real time or may be executed with a predetermined delay. Unless otherwise specified, the execution order of the multiple operations to be executed is not limited.
[0021] 3) Voice conversion generally refers to the technology of changing the voice quality of speech. This technology can convert the voice quality of speech from speaker A to speaker B. Here, speaker A is the person who uttered this speech, and is generally referred to as the source speaker. Speaker B is the speaker with the converted target voice quality, and is generally referred to as the target speaker. Currently, Voice The conversion technology can be divided into three types: one-to-one (only possible to convert the voice of a specific person to the voice of another specific person), many-to-one (able to convert the voices of any person to the voice of a specific person), and many-to-many (able to convert the voices of any person to the voices of any other person).
[0022] 4) A phoneme refers to the smallest speech unit classified based on the natural attributes of speech.
[0023] 5) Phonetic Posterior Grams (PPG) is a matrix whose size is the number of speech frames × the number of phonemes, and is used to represent the probability of phonemes that may be uttered for each phoneme frame in a speech segment.
[0024] 6) Naturalness is one of the evaluation metrics commonly used in speech synthesis tasks or voice conversion tasks, and it is used to determine whether the voice sounds natural as if a human is speaking.
[0025] 7) Similarity is one of the evaluation metrics commonly used in voice conversion tasks, and it is used to determine whether the voice sounds similar to that of the target speaker.
[0026] 8) Spectrum is the data in the frequency domain obtained by performing a Fourier transform on a speech signal. Generally, a speech signal is formed by the superposition of multiple sine waves, but the spectrum can more clearly depict the waveform composition of the speech signal. If the frequency is discretized and displayed, the spectrum is a one-dimensional vector quantity (only in the frequency domain).
[0027] 9) A spectrogram refers to a spectrogram obtained by dividing speech into frames (which may include in-frame signal processing steps such as applying a window function), then performing a Fourier transform on the signal of each frame to obtain the spectrum, and then overlapping them in the time domain. The spectrogram can reflect the time-varying changes of the superimposed sine waves in the speech signal in the time domain. A Mel spectrogram, also abbreviated as Mel diagram or Mel spectrogram, is a spectrogram obtained by filtering the spectrum using a designed filter based on the spectrogram. Compared with a general spectrogram, it has fewer frequency dimensions and focuses on the speech signals in the low-frequency band where the human auditory system is more sensitive. Generally, the Mel diagram is easier to extract / separate information from the speech signal and is also easier to modify the speech.
[0028] Please refer to FIG. 1. FIG. 1 is a schematic architecture diagram of a virtual concert processing system 100 provided in an embodiment of the present application. To support exemplary application cases, terminals (exemplarily shown as terminal 400-1 and terminal 400-2) are connected to server 200 via network 300. Network 300 may be a wide area communication network, a local communication network, or a combination of both, and uses a wireless link to realize data transmission.
[0029] In actual operation, the terminal may be various user terminals such as a smartphone, a tablet computer, a notebook computer, etc., or a desktop computer, a television receiver, or any combination of two or more of these data processing devices. Server 200 may be a single server arranged alone to support various tasks, may be arranged as a group of servers, or may be a cloud server, etc.
[0030] In actual operation, clients such as an instant message client, a video playback client, a live delivery client, a learning client, a singing client, etc. are installed on the terminal. When the user (the current target person) starts a client on the terminal to perform singing practice or produce a virtual concert, the terminal receives a concert production instruction of the target singer based on the presented concert entrance. Then, in response to the concert production instruction, the terminal sends a production request to server 200 to request the production of a concert room for imitating and singing the target singer's song. Server 200 produces a concert room for imitating and singing the target singer's song based on the production request, and returns it to the terminal for display. When the current user sings the target singer's song in the concert room, the terminal collects the singing content of the target singer's song imitated and sung by the current target person, and sends the collected singing content to server 200. Server 200 distributes the received singing content to the terminals of each target person who enters the concert room, and each terminal plays the singing content through the concert room.
[0031] Please refer to FIG. 2. FIG. 2 is a schematic structural diagram of the electronic device 500 provided in the embodiment of the present application. In actual operation, the electronic device 500 may be the terminal or the server 200 in FIG. 1. Taking the case where the electronic device is the terminal shown in FIG. 1 as an example, the electronic device for implementing the virtual concert processing method of the embodiment of the present application will be described. The electronic device 500 shown in FIG. 2 includes at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. Each unit in the electronic device 500 is connected together via a bus system 540. It should be noted that the bus system 540 is used for connection communication between these units. In addition to the data bus, the bus system 540 includes a power bus, a control bus and a status signal bus. However, for the sake of clarity in the description, in FIG. 2, all various buses are denoted as the bus system 540.
[0032] The processor 510 may be an integrated circuit chip with the ability to process signals. For example, it may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Here, the general-purpose processor may be a microprocessor or any ordinary processor, etc.
[0033] The user interface 530 includes one or more output devices 531 capable of presenting media content. The output devices 531 include one or more speakers and / or one or more visual displays. The user interface 530 further includes one or more input devices 532. The input devices 532 include user interface members for the user to input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.
[0034] Memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid memory, hard disk drives, optical disk drives, and the like. Memory 550 may include one or more storage devices physically remote from processor 510.
[0035] Memory 550 may include volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory may be ROM (Read Only Memory), and the volatile memory may be RAM (Random Access Memory). Memory 550 described in the embodiments of the present application may be any suitable type of memory.
[0036] In some embodiments, memory 550 can store data and support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof. This will be illustrated by way of example below.
[0037] Operating system 551 includes system programs configured to process various basic system services and execute hardware-related tasks, such as a framework layer, a core library layer, a drive layer, etc., to implement various basic tasks and process hardware-based tasks.
[0038] Network communication module 552 is configured to reach other computer devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include Bluetooth (registered trademark), WiFi (registered trademark), USB (Universal Serial Bus), and the like.
[0039] The hint module 553 is configured to be able to present information via one or more output devices 531 (such as a display, a microphone, etc.) associated with the user interface 530 (for example, a user interface for operating peripheral devices and displaying content and information).
[0040] The input processing module 554 is configured to detect one or more inputs or interactions of one or more users from one of the one or more input devices 532 and translate the detected inputs or interactions.
[0041] In some embodiments, the processing device of the virtual concert provided in the embodiments of the present application can be implemented in a software manner. FIG. 2 shows the processing device 555 of the virtual concert stored in the memory 550. The processing device 555 may be software in the form of programs and plugins, etc., including a command receiving module 5551, a room creation module 5552, and a singing playback module 5553, which are software modules. Since these modules are logical modules, any combination or division is possible according to the functions to be realized. The functions of each module will be described later.
[0042] In some other embodiments, the virtual concert processing apparatus provided in the embodiments of the present application can be implemented in a hardware manner. As an example, the virtual concert processing apparatus provided in the embodiments of the present application is a processor using a hardware decoding processor method and is programmed to execute the virtual concert processing method provided in the embodiments of the present application. For example, for a processor using a hardware decoding method, one or more ASICs (Application Specific Integrated Circuits), DSPs, PLDs (Programmable Logic Devices), CPLDs (Complex Programmable Logic Devices), FPGAs (Field-Programmable Gate Arrays) or other electronic devices can be adopted.
[0043] In some embodiments, a terminal or a server 200 can implement the virtual concert processing method provided in the embodiments of the present application by operating a computer program. For example, the computer program may be a native program or a software module in an operating system. It may also be a native APP (Application) program, that is, a program that can operate only after being installed in an operating system, such as a live streaming APP or an instant messaging APP. It may also be an applet that can operate only after being downloaded in a browser environment. It may also be an applet that can be incorporated into any APP. That is, the above computer program may be an application program, a module, or a plugin in any form.
[0044] Next, the processing method of the virtual concert provided in the embodiments of the present application will be described in combination with the drawings. The processing method of the virtual concert provided in the embodiments of the present application may be executed independently by the terminal in FIG. 1, or may be executed jointly by the terminal and the server 200 in FIG. 1. Hereinafter, the case where the terminal in FIG. 1 independently executes the processing method of the virtual concert provided in the embodiments of the present application will be described as an example. Please refer to FIG. 3. FIG. 3 is a flowchart overview diagram of the processing method of the virtual concert provided in the embodiments of the present application. It will be described in connection with the steps shown in FIG. 3.
[0045] Note that the method shown in FIG. 3 can be executed by various computer programs operating on the terminal, and is not limited to the above-mentioned client, and may be the above-mentioned operating system 551, software module, and script. Therefore, the client should not be regarded as limiting the embodiments of the present application.
[0046] Step 101: The terminal presents a concert entrance.
[0047] In actual operation, clients such as an instant messaging client, a video playback client, a live streaming client, a learning client, and a singing client are installed on the terminal. The user can listen to songs, sing, or hold a concert corresponding to the target singer through the client on the terminal. In actual operation, when the terminal presents a song practice interface and presents a concert entrance for creating a virtual concert on the song practice interface, the creation and holding of the concert are realized based on the concert entrance.
[0048] The concert corresponding to the above target singer is essentially a virtual concert produced and held by a user (who is not the same person as the target singer). A so-called virtual concert refers to a concert for singing in imitation of or mimicking the target singer. Based on the produced virtual concert, the user can imitate the songs sung by a specific singer. Here, the virtual concert usually corresponds to a singer, such as a virtual concert of singer A or a virtual concert of singer B. Taking the virtual concert of singer A as an example, when a user produces or holds a virtual concert of singer A, it means that the user creates a concert hall for the purpose of imitating and singing the songs of singer A. That is, the user creates a concert hall where they can imitate the voice quality of the original singer and sing the songs of the original singer. For example, if the user creates a concert hall where they can imitate the voice quality of original singer A and sing song B of original singer A, and imitate and sing the songs of singer A in the created concert hall, the purpose of holding a concert of singer A can be achieved. Especially when the singer to be imitated is a deceased singer, in the real world, it is impossible to hold a concert of the deceased singer again, but the concert of the deceased singer can be reproduced by such a method of holding a virtual concert, and such a performance method has the effect of better conveying the emotions of the singer. In this way, the produced concert hall corresponds to the target singer, and the person who enters the concert hall can continuously enjoy multiple songs of the target singer. The current person can continuously share the songs of the target singer that they have imitated and sung, improving the sharing efficiency of the songs to a specific person. Also, compared with the method of sharing songs point-to-point in related technologies, since the user does not need to repeatedly perform the song sharing operation, when the songs to be shared are multiple songs of a specific singer, the process of sharing the multiple songs can be simplified, and the efficiency of man-machine interaction is increased. Also, compared with the simple random singing in related technologies, the singing interaction method becomes richer, contributing to increasing the stickiness and retention rate of users.
[0049] In some embodiments, the terminal can present a concert entrance to the current target user's song practice interface in the following manner. That is, present a song practice entrance for practicing a song on the song practice interface, receive a practice command for the song of the target singer based on the song practice entrance, in response to the song practice command, collect the practice voice of the current target user singing the song of the target singer, and if it is determined based on the practice voice that the current target user has the production qualification to produce a concert of the target singer, present a concert entrance associated with the target singer on the current target user's song practice interface.
[0050] In actual operation, in order to provide a realistic auditory enjoyment, it is necessary to ensure that the singing level of the current target user singing the song of the target singer is comparable to that of the target singer himself / herself. Therefore, if the user wants to produce a virtual concert of the target singer, it is necessary to practice singing the song of the target singer to improve the user's imitation ability for the song of the target singer. Only when the practice result shows that the current target user has the production qualification to produce a concert of the target singer (for example, when the current target user sings the song of the target singer, the voice, vocal quality, etc. are very close to or the same as the original), present a concert entrance associated with the target singer on the current target user's song practice interface so that the concert of the target singer can be produced through the concert entrance. Of course, in actual operation, the concert holding qualification conditions can be lowered or cancelled to lower the hurdle of virtual concert production and realize an environment where everyone enjoys singing together as a "concert for all".
[0051] Here, the production qualification of the current target person for the concert of the target singer will be described. In actual operation, the terminal acquires the last practice song that the user practiced singing for the song of the target singer, and compares the practice song with the original voice of the target singer in at least one singing feature (such as voice quality). If the similarity reaches the similarity threshold, it is determined that the current target person has the production qualification for the concert of the target singer. In some embodiments, the terminal acquires a plurality (at least two) of practice songs that the user practiced singing for the song of the target singer within a certain recent period, compares each practice song with the original voice of the target singer in at least one singing feature (such as voice quality), acquires the similarity corresponding to each practice song, averages the similarities of the at least two acquired practice songs to obtain an average similarity, and if the average similarity reaches the similarity threshold, it may be determined that the current target person has the production qualification for the concert of the target singer.
[0052] In some embodiments, based on the song practice entry, the terminal receives a song practice command for the target singer in the following manner. That is, in response to a trigger operation on the song practice entry, a singer selection interface including at least one candidate singer is presented, in response to a selection operation for the target singer among the at least one candidate singer, at least one candidate song corresponding to the target singer is presented, in response to a selection operation for the target song among the at least one candidate song, a voice recording entry for singing the target song is presented, and in response to a trigger operation on the voice recording entry, a song practice command for the target song of the target singer is received.
[0053] Please refer to FIG. 4. FIG. 4 is a schematic diagram showing the display of the concert entrance provided in the embodiment of the present application. First, a song practice entrance 401 for practice is presented on the song practice interface. When the user triggers (e.g., clicks, double-clicks, swipes, etc.) the song practice entrance 401, the terminal responds to the trigger operation by presenting a singer selection interface 402 and presenting a plurality of selectable candidate singers on the singer selection interface 402. When the user selects a target singer from among them, the terminal responds to the selection operation by presenting a plurality of candidate songs for the practice corresponding to the target singer. When the user selects a target song, the terminal responds to the selection operation by presenting a voice recording entrance 403. When the user triggers the voice recording entrance 403, the terminal responds to the trigger operation by receiving a song practice command for the target song, and in response to the song practice command, collecting the practice voice of the current target person singing the song of the target singer, and determining whether the current target person has the qualification to produce a concert of the target singer based on the practice voice. If it is determined that the current target person has the qualification to produce a concert of the target singer, a concert entrance 404 is presented on the song practice interface.
[0054] In some embodiments, the number of target songs may be plural (two or more). For example, please refer to FIG. 5. FIG. 5 is a schematic diagram of the singing song selection provided in the embodiment of the present application. For a plurality of candidate songs for practice corresponding to the presented target singer, a selection key for triggering each candidate song is associated. When the user triggers any selection key (for example, three selection keys), the terminal first receives the user's trigger operation for the selection key (three selection keys) associated with the candidate song (three songs) that the user wants to practice, and in response to the decision command for the selected selection key, receives a selection operation for the target song. At this time, the target song is the candidate song (three songs) corresponding to the selected selection key (three selection keys), and an audio recording entrance is presented in response to the selection operation. The terminal receives a song practice command for the target song (three songs) in response to the trigger operation for the audio recording entrance, and sequentially collects the practice voices (practice voices corresponding to three songs) of the current target person singing the songs of the target singer in response to the song practice command. Then, based on the practice voice, it is determined whether the current target person has the production qualification to produce a concert of the target singer. If it is determined that the current target person has the production qualification to produce a concert of the target singer, a concert entrance is presented on the song practice interface. In this way, by selecting and practicing a plurality of songs at one time, the practice efficiency of the songs can be improved.
[0055] In some embodiments, before presenting a concert entrance associated with the target singer on the song practice interface of the current target person, it may be determined whether the current target person has the production qualification to produce a concert of the target singer in the following manner. That is, a practice score obtained by grading the practice voice is presented. If the practice score reaches the target score, it is determined that the current target person has the production qualification to produce a concert of the target singer. If the practice score is lower than the target score, it is determined that the current target person does not have the production qualification to produce a concert of the target singer. In this case, a re-practice entrance is presented so that the current target person can practice the songs of the target singer again.
[0056] Here, the scoring of the practice audio of the target song will be described. When actually implemented, at least one of the singing parameters such as the pitch, rhythm, melody, intonation, lyrics, and emotion of the practice audio is obtained. Based on the singing time point, the singing parameters of the practice audio are compared with the singing parameters of the original audio of the target song to obtain the similarity, and based on the magnitude of the similarity and the mapping relationship between the magnitude of the similarity and the score, the score of the practice audio is determined.
[0057] Please refer to FIG. 6. FIG. 6 is a schematic diagram showing an overview of the practice result provided in the embodiment of the present application. By presenting the practice score in the practice result interface and determining whether the practice score reaches a predetermined target score (assuming the target score is 95 points with a full score of 100 points), it is determined whether the current target person has the production qualification to produce a concert of the target singer. In (1), since the practice score (98 points) reaches the predetermined target score (95 points), prompt information 601 for notifying that the current target person has the production qualification to produce a concert of the target singer is presented. In (2), since the practice score (80 points) is lower than the predetermined target score (95 points), prompt information 602 for notifying that the current target person does not have the production qualification to produce a concert of the target singer and a re-practice entry are presented. The current target person can re-practice the songs of the target singer through the re-practice entry. If the practice score increases to the target score through multiple practices by learning the singing skills, vocal quality, tone, etc. of the target singer, the current target person can obtain the production qualification to produce a concert of the target singer.
[0058] In some embodiments, before presenting the practice score of the practice audio, the terminal can determine the practice score of the practice audio in the following manner. That is, when the number of practiced songs is at least two, the practice scores corresponding to the practice audio of each song of the current target person are presented, the singing difficulty of each song is obtained, and the weight corresponding to the song is determined based on the singing difficulty. Based on the weight, the practice scores of the practice audio of each song are weighted-averaged to obtain the practice score of the practice audio of the songs practiced by the current target person.
[0059] The singing difficulty may be the level of the song or the difficulty coefficient. Usually, the higher the level of the song or the larger the difficulty coefficient, the higher the singing difficulty and the greater the corresponding weight. By comprehensively averaging the practice scores of multiple target songs practiced by the current subject using the weighted average method to calculate the final practice score, the true singing level of the current subject for the songs of the target singer can be accurately represented, an objective evaluation of the singing level of the current subject can be ensured, and the scientific nature and rationality of practice score acquisition can be improved.
[0060] In some embodiments, the practice score includes at least one of a voice quality score and an emotion score. Correspondingly, before presenting the practice score corresponding to the practice voice, the terminal can determine the practice score of the practice voice in the following way. That is, when the practice score includes a voice quality score, perform voice quality conversion on the practice voice to obtain a practice voice quality corresponding to the target singer, compare the practice voice quality with the original voice quality when the target singer sang the song, obtain the corresponding voice quality similarity, and determine the voice quality score based on the voice quality similarity. When the practice score includes an emotion score, perform emotion recognition on the practice voice, obtain the corresponding practice emotion degree, compare the practice emotion degree with the original emotion degree when the target singer sang the song, obtain the corresponding emotion similarity, and determine the emotion score based on the emotion similarity.
[0061] When performing voice quality conversion, convert the practice voice of the current subject to match the original voice quality of the target singer to obtain a practice voice quality relatively close to the original voice quality of the target singer. Note that even after performing voice quality conversion, the converted practice voice quality will not be exactly the same as the original voice quality of the original singer, but only relatively closer. Also, since the singing levels vary from user to user, the practice voice qualities obtained by converting the practice voices will not be the same if the users are different. Therefore, the voice quality similarity between the practice voice quality and the original voice quality also varies from user to user, resulting in differences in the voice quality scores.
[0062] In some embodiments, the terminal performs voice quality conversion on the practice voice in the following manner to obtain a practice voice quality corresponding to the target singer. That is, the phoneme recognition model performs phoneme recognition on the practice voice to obtain the corresponding phoneme sequence. Loudness recognition is performed on the practice voice to obtain the corresponding loudness characteristics. Melody recognition is performed on the practice voice to obtain a sine excitation signal representing the melody. The phoneme sequence, loudness characteristics, and sine excitation signal are combined and processed by a sound wave synthesizer to obtain a practice voice quality corresponding to the target singer.
[0063] As shown in FIG. 18, the phoneme recognition module, also called a PPG extractor, is part of an automatic speech recognition (ASR) model. The function of the ASR model is to convert speech into text. In essence, it first converts speech into a phoneme sequence consisting of multiple phonemes, and then converts the phoneme sequence into text. The function of the PPG extractor is to first convert speech into a phoneme sequence and is used to extract information unrelated to voice quality, such as text content information, from the practice voice. Note that a phoneme refers to the smallest speech unit classified based on the natural attributes of speech.
[0064] In actual operation, as shown in FIG. 19, considering that the practice voice is actually a chaotic waveform signal in the time domain, in order to make it easier to analyze, the practice voice in the time domain is converted into the frequency domain by fast Fourier transform to obtain the voice spectrum corresponding to the voice data. Based on the obtained voice spectrum, the degree of difference between the voice spectra corresponding to adjacent sampling windows is obtained. Further, based on the obtained plurality of degrees of difference, the energy spectrum corresponding to each sampling window is specified, and finally, a spectrogram (for example, a Mel spectrogram) corresponding to the practice voice may be obtained. Then, downsampling processing in the downsampling layer is performed on the spectrogram corresponding to the practice voice. The downsampling layer has a two-dimensional convolutional structure and downsamples the input spectrogram at a two-fold time scale to obtain downsampling characteristics. Then, the downsampling characteristics are input to an encoder (an integrated encoder or a transformer encoder) for encoding processing to obtain corresponding encoded characteristics. Then, by inputting the encoded characteristics into a decoder for decoding processing, the phoneme sequence of the practice voice is predicted. Here, the decoder may be a CTC decoder. The decoder includes one fully connected layer, and the decoding process is as follows. That is, based on the encoded characteristics, the phoneme with the highest probability is screened for each frame of the practice voice, and the corresponding phonemes with the highest probability for each frame of the screened practice voice are used to form a time-series phoneme sequence. In the time-series phoneme sequence, adjacent identical phonemes are integrated to obtain a phoneme sequence.
[0065] The voice loudness characteristic is the time series of loudness for each frame of the practice voice in the practice voice, that is, the corresponding maximum amplitude for each frame of the practice voice obtained by performing a short-time Fourier transform on the practice voice. Voice loudness refers to the intensity of sound, and loudness is the degree of the intensity of sound judged by the human ear's sensation, that is, the voice loudness. Based on this, the practice voices can be arranged in a series from weak to strong. The sine excitation signal is calculated using the fundamental frequency of the sound (F0, the fundamental frequency of each frame of the sound is equal to the pitch of each frame of the sound) and is used to represent the melody of the voice. Generally, melody is a series with structure and rhythm formed by adding artistic concepts to several musical sounds, composed of a certain pitch, duration, and volume, and progresses through a single melodic part with logical factors. Melody is formed by the organic combination of many basic musical elements, such as key, rhythm, beat, intensity, timbre, expression method / way, etc. The purpose of the sound wave synthesizer is to synthesize three features unrelated to the speaker's timbre, namely the phoneme series of the practice voice, the voice loudness characteristic, and the sine excitation signal, into the sound wave of the singing voice sung using the timbre of the target singer (that is, the practice timbre corresponding to the above-mentioned target singer).
[0066] In actual operation, the sound wave of the singing voice sung using the timbre of the target singer (that is, the practice timbre corresponding to the above-mentioned target singer) synthesized from the above-mentioned user's practice voice can be provided to the user so that the user can enjoy or share it. Based on the obtained practice timbre corresponding to the target singer, the user can understand the voice conversion effect, thereby identify which singing parts have room for improvement, and learn the singing skills, timbre, tone, etc. of the target singer (original singer), so as to continuously improve their own singing technical level step by step, make the singing skills and singing methods closer and closer to the original singer, and finally achieve the goal of raising the practice score until obtaining the production qualification to produce the target singer's concert.
[0067] In some embodiments, before presenting the practice score corresponding to the practice voice, the terminal can determine the practice score of the practice voice in the following manner. That is, the practice voice is transmitted to the terminals of other subjects, and the terminals of other subjects are caused to obtain a manual score corresponding to the input practice voice based on the scoring entry corresponding to the practice voice. Then, the manual score returned from other terminals is received, and the practice score corresponding to the practice voice is determined based on the manual score.
[0068] Here, the practice voice to be scored is cast into a voting pool corresponding to the target singer, and the practice voice is pushed to the terminals of other subjects. Other subjects score the practice voice of the current subject through the scoring entry presented on other terminals. Please refer to FIG. 7. FIG. 7 is a schematic diagram of the scoring of the practice voice provided in the embodiment of the present application. A scoring entry for scoring the practice voice of singing the song of the target singer is presented on the user scoring interface, and a manual score is obtained by scoring the practice voice to be scored through the scoring entry, and the manual score returned from the terminals of other subjects is used as the practice score corresponding to the practice voice.
[0069] In actual operation, when determining the manual score, the attributes of each subject participating in the manual score (such as identity, level, etc.) can be considered, and appropriate scoring weights can be determined based on the attributes of each subject. For example, the identities of the subjects participating in the manual score include professional musicians, the media, the general public, etc., and the corresponding manual score weights are different when the identities of the subjects are different. Also, for example, the singing levels of the subjects participating in the manual score are at levels from 0 to 5, and the corresponding manual score weights may also be different depending on the differences in the levels of the subjects. After obtaining the scores for the practice voices of each subject, each score is weighted and averaged based on the weights of each subject to obtain the practice score of the practice voice. By doing so, the obtained practice score can accurately represent the true singing level of the current subject for the song of the target singer, ensure an objective evaluation of the singing level of the current subject, and improve the scientificity and rationality of obtaining the practice score.
[0070] In some embodiments, the terminal can obtain machine scoring corresponding to the practice voice, and when the machine scoring reaches the scoring threshold, send the practice voice to the terminals of other subjects. Correspondingly, the terminal can determine the practice score of the practice voice based on manual scoring by averaging the machine scoring and the manual scoring to obtain the practice score corresponding to the practice voice.
[0071] Here, first, the practice voice is machine-scored by an artificial intelligence method to obtain the corresponding machine scoring. When the machine scoring reaches a predetermined scoring threshold (for example, with a full score of 100 points and the scoring threshold being 80 points), the practice voice can be cast into the voting pool corresponding to the target singer, and the practice voice can be pushed to the terminals of other subjects. By having other subjects score the practice voice of the current subject through the scoring entry presented on the terminal, the manual scoring corresponding to the practice voice can be obtained. Then, the machine scoring and the manual scoring are combined to obtain the practice score corresponding to the practice voice. For example, the machine scoring and the manual scoring are averaged to obtain the practice score corresponding to the practice voice. By doing so, the accuracy of the practice score obtained by combining the machine scoring and the manual scoring is increased. The practice score with high accuracy can accurately represent the true singing level of the current subject for the target singer's song, ensuring an objective evaluation of the current subject's singing level, and improving the scientificity and rationality of practice score acquisition.
[0072] In some embodiments, before presenting the concert entrance associated with the target singer on the song practice interface corresponding to the current target person, the terminal may determine whether the current target person has the production qualification to produce the concert of the target singer in the following manner. That is, present the song practice ranking corresponding to the practice song of the current target person. If the song practice ranking is before the target ranking, it is determined that the current target person has the production qualification to produce the concert of the target singer. By doing so, only users with higher rankings are eligible to produce or hold the virtual concert of the target singer, ensuring that all users who produce or hold the virtual concert have a high singing level and guaranteeing the quality of the concert.
[0073] In actual operation, on the song practice interface, the song practice ranking corresponding to the song practiced by the current target person, which is determined based on the practice voice of the practiced song, may be presented. The song practice ranking is determined based on the practice score of the practice voice. For example, a descending song practice ranking is determined according to the order from the highest to the lowest practice score of the users who practiced the target singer. For example, please refer to FIG. 8. FIG. 8 is a schematic diagram of the song practice ranking provided in the embodiment of the present application. When there are multiple users who practiced song B of singer A, a descending song practice ranking is presented, and it is determined that the current target person has the production qualification to produce the concert of singer A only when the song practice ranking of the current target person is before the target ranking (for example, the 4th place). That is, the top three users all have the production qualification to produce the concert of singer A. When the song practice ranking of the current target person is the target ranking (the 4th place) or after the target ranking, it is determined that the current target person does not have the production qualification to produce the concert of the target singer A. In addition, a playback entrance may be presented on the song practice interface so that the practice voice when the corresponding user practiced song B can be played through the playback entrance.
[0074] In some embodiments, when the number of songs practiced by the current target person is at least two, the terminal presents the total score of all the songs sung by the current target person and a detailed entry for checking details, and in response to a trigger operation on the detailed entry, presents a detailed page and presents the practice score corresponding to each song on the detailed page.
[0075] The detailed page may be displayed in the form of a pop-up window or in the form of a sub-interface independent of the song practice interface. In the embodiments of the present application, the display form of the detailed page is not particularly limited.
[0076] Please refer to FIG. 9. FIG. 9 is a schematic diagram of the song practice ranking provided in the embodiments of the present application. When the number of songs practiced by each target person is multiple, a descending song practice ranking is presented, and at the same time, the total score of all the songs sung by each target person and a detailed entry for checking details are presented. For example, when the current target person triggers (such as clicks, double-clicks, swipes, etc.) the detailed entry 901 of the user 1 the terminal presents the detailed page 902 in the form of a pop-up window in response to the trigger operation. The detailed page 902 presents all the songs practiced by user 1, such as song 1, song 2, song 3, song 4, and the practice score corresponding to each song. By doing so, the user can enjoy or share the songs sung by each target person and the singing level among them, and thus can more comprehensively recognize their own singing level and the direction for improvement, contribute to continuously improving their own singing level step by step, get closer to the original singer in singing skills and singing methods, and finally achieve the purpose of raising the practice score until obtaining the production qualification for producing the concert of the target singer.
[0077] Step 102: Receive a concert production instruction about the target singer based on the concert entry.
[0078] In actual operation, when it is determined that the current target person has the production qualification to produce a concert of the target singer, only when presenting the concert entrance associated with the target singer, if the current target person triggers (for example, clicks, double-clicks, swipes, etc.) the concert entrance, the terminal can receive a concert production instruction for the target singer in response to the trigger operation and create a concert room for virtual singing for the songs of the target singer based on the concert production instruction. When the concert entrance is always presented on the song practice interface regardless of whether the current target person has the production qualification to produce a concert of the target singer, the terminal needs to first determine whether the current target person has the production qualification to produce a concert of the target singer in response to a trigger operation on the concert entrance. Only when the current target person has the production qualification to produce a concert of the target singer, can it receive a concert production instruction corresponding to the target singer. On the other hand, when the current target person does not have the production qualification to produce a concert of the target singer, even if the concert entrance is triggered, a concert production instruction for the target singer cannot be triggered.
[0079] In some embodiments, the terminal can receive a concert production instruction for the target singer based on the concert entrance in the following way. That is, in response to a trigger operation on the concert entrance, a singer selection interface including at least one candidate singer is presented, and in response to a selection operation of the target singer among at least one candidate singer, when it is determined that the current target person has the production qualification to produce a concert of the target singer, a concert production instruction for the target singer is received.
[0080] Please refer to FIG. 10. FIG. 10 is a schematic diagram of the trigger of the concert production instruction provided in the embodiment of the present application. The concert entrance 1001 is a common entrance for producing concerts of each singer. When the current target person triggers the concert entrance 1001, the terminal responds to the trigger operation and presents the singer selection interface. soundsPrompt and present at least one candidate singer that can be selected by the current target person on the singer selection interface. When the current target person selects the target singer 1002 from among them, the terminal, in response to the selection operation, determines whether the current target person has the qualification to produce a concert of the target singer 1002 and presents a prompt for notifying whether the production qualification is available. When the current target person has the qualification to produce a concert of the target singer 1002 , the terminal presents a prompt indicating that the production qualification is available and receives a concert production instruction for the target singer 1002 . On the other hand, when the current target person does not have the qualification to produce a concert of the target singer 1002 , a prompt indicating that the production qualification is not available is presented. Even if the concert entrance is triggered for the time being, the concert production instruction for the target singer 1002 cannot be triggered. By doing so, only users with the qualification to produce a concert of the target singer 1002 can produce a virtual concert of the target singer 1002 , so the quality of the concert is guaranteed.
[0081] In some embodiments, the terminal can receive a concert production instruction for the target singer based on the concert entrance in the following manner. That is, in response to a trigger operation on the concert entrance, a singer selection interface including at least one candidate singer for whom the current target person has the qualification to produce a concert is presented, and in response to a selection operation on the target singer among at least one candidate singer, a concert production instruction for the target singer is received.
[0082] In actual operation, the current target person may have the production qualifications for concerts of multiple singers, for example, having the production qualifications for the concerts of both singer A and singer B at the same time. In such a case, the concert entrance is a common entrance for producing the concerts of all singers with production qualifications. The terminal of the current target person can produce the concert of singer A or the concert of singer B through the concert entrance, and the current target person can select the concert of the target singer that he / she wants to hold this time from among them.
[0083] Please refer to FIG. 11. FIG. 11 is a schematic diagram of the trigger for the concert production instruction provided in the embodiment of the present application. When the current target person triggers the concert entrance 1101, the terminal presents a singer selection interface in response to the trigger operation, and presents candidate singers 1102 and 1103 that can be selected by the current target person in the singer selection interface. The current target person has the production qualifications for producing the concerts of candidate singer 1102 and candidate singer 1103 at the same time. When the current target person selects candidate singer 1103, the terminal, in response to the selection operation, takes candidate singer 1103 as the target singer and receives a concert production instruction for the target singer (i.e., candidate singer 1103).
[0084] In some embodiments, when the number of concert entrances is at least one, a singer is associated with the concert entrance, and there is a corresponding relationship between the singer associated with the concert entrance. The terminal receives a concert production instruction corresponding to the target singer based on the concert entrance in the following manner. That is, in response to a trigger operation on the concert entrance associated with the target singer, a concert production instruction corresponding to the target singer is received.
[0085] Here, the number of concert entrances presented on the song practice interface may be one or more (i.e., two or more). Each concert entrance has a concert to produce production qualificationA corresponding singer equipped with it is associated, and there is a one-to-one correspondence between the singer associated with the concert entrance. As shown in FIG. 12, FIG. 12 is a trigger overview diagram of the concert production instruction provided in the embodiment of the present application. In the related area of the song practice entrance 1201 called "Practice Start", two concert entrances, namely concert entrance 1202 and concert entrance 1203, are presented. Concert entrance 1202 is associated with singer A, and concert entrance 1203 is associated with singer B. That is, the current target person simultaneously has the production qualification to produce the concerts of singer A and singer B. Concert entrance 1202 is used to produce the concert of singer A, and concert entrance 1203 is used to produce the concert of singer B. The current target person can select the concert entrance corresponding to the concert of the target singer that he / she wants to hold this time from them. For example, when the user currently selects concert entrance 1203, the terminal responds to the trigger operation, takes the candidate singer B as the target singer, and receives a concert production instruction for the target singer (i.e., candidate singer B).
[0086] In some embodiments, the terminal can receive a concert production instruction for the target singer based on the concert entrance in the following way. That is, when the target singer is associated with the concert entrance, in response to the trigger operation for the concert entrance, prompt information for reminding whether to apply for the production of the concert corresponding to the target singer is presented. When a decision operation on the prompt information is received, a concert production instruction for the target singer is received.
[0087] Here, the fact that the target singer is associated with the concert entrance indicates that the current target person already has the production qualification to produce the concert of the target singer. When the current target person triggers the concert entrance, the terminal presents prompt information for reminding whether to apply for the production of the concert corresponding to the target singer in response to the trigger operation. The current target person can decide whether to produce the concert corresponding to the target singer based on the prompt information. For example, if the current target person decides to produce the concert corresponding to the target singer, a decision operation can be triggered by triggering the corresponding decision button. When the terminal receives the decision operation, it can receive a concert production instruction corresponding to the target singer. On the other hand, if the current target person decides not to produce the concert corresponding to the target singer, a cancellation operation can be triggered by triggering the corresponding cancellation button. When the terminal receives the cancellation operation, it does not receive a concert production instruction regarding the target singer. At this time, a song practice interface can be presented with a song practice entrance. The current target person can practice the songs of the target singer or other singers through the song practice entrance, continuously improve their singing skill level step by step, get closer and closer to the original singer in singing skills and singing methods, and achieve the goal of raising the practice score until obtaining the production qualification to produce the concert of the target singer.
[0088] In some embodiments, the terminal can realize receiving a concert production instruction corresponding to the target singer when receiving a decision operation on the prompt information in the following way. That is, when receiving a decision operation on the prompt information, an application interface for applying to produce the concert of the target singer is presented, and an editing entrance for editing concert-related information is presented on the application interface. The concert information edited based on the editing entrance is received, and in response to the decision operation on the concert information, a concert production instruction regarding the target singer is received.
[0089] Please refer to FIG. 13. FIG. 13 is a schematic diagram of the trigger for the concert production command provided in the embodiment of the present application. In response to a trigger operation for the concert entrance 1301, the terminal presents prompt information 1302 saying "Congratulations, your étude ranks first with singer A. Do you want to select the application for the virtual concert of singer A?", an immediate production button 1303 for immediately producing a concert room, and a cancel button 1304. When the user triggers the immediate production button 1303, the terminal receives a decision operation for the prompt information, and in response to the decision operation, presents an application interface 1305 and a decision button 1306 corresponding to the concert information. The application interface presents an editing entry for editing concert information related to the concert to be produced, such as the user name, the song scheduled to be sung, guest performers, concert time, whether it is paid, etc. In response to a trigger operation for the decision button 1306, the terminal receives a decision operation for the concert information, and in response to the decision operation, receives a concert production command for singer A. for the holding In addition, through the editing entry, it is also possible to edit promotional information related to the concert, such as concert introduction, concert start time, etc. The terminal generates a promotional poster or promotional applet with the promotional information in response to a decision operation for the promotional information, and shares the promotional poster and promotional applet to the terminals of other target persons, so as to widely promote and recommend the concert corresponding to the target singer held by the current target person. Thereby, allowing the terminals of other target persons to enter the concert room produced by the current target person, attracting more users to watch the online virtual concert produced by the current target person online, enabling the produced virtual concert to reach more people, and ultimately guiding more users to practice the songs of the target singer or other singers, improving the user retention rate.
[0090]
[0091] reservation In some embodiments, the concert room is produced reservationIt is also possible. The terminal can receive a concert production instruction corresponding to the target singer based on the concert entrance in the following manner. That is, a reservation entrance for reserving the production of a concert room is presented, and in response to a trigger operation on the reservation entrance, a reservation interface for reserving the production of a concert of the target singer is presented. At the same time, an editing entrance for editing the reservation information of the concert is presented on the reservation interface, and concert reservation information including at least the concert start time edited based on the editing entrance is received. In response to a decision operation on the concert reservation information, a concert production instruction for the target singer is received.
[0092] Please refer to FIG. 14. FIG. 14 is a trigger schematic diagram of the concert production instruction provided in the embodiment of the present application. In response to a trigger operation on the concert entrance 1401, the terminal presents the prompt information 1402 "Congratulations, your practice song ranks first with singer A. Do you select to apply for the virtual concert of singer A?" and at the same time presents a reservation entrance 1403 for reserving the production of a concert room. In response to a trigger operation on the reservation entrance 1403, a reservation interface 1404 of the concert room is presented. Information such as concert introduction, concert start time, concert duration, or more other information can be set on the reservation interface. The concert start time may be determined based on the time selected by the reservation time selection key, or may be determined based on the time recommended by the system. After the setting is completed, when the current target person triggers the reservation decision button 1405 "Produce", a decision operation on the concert reservation information is received, and in response to the decision operation, a concert production instruction for singer A is received.
[0093] Step 103: In response to the concert production instruction, produce a concert room for imitating and singing the songs of the target singer.
[0094] A concert room refers to a network live broadcast program opened by the current target person, which enables the current target person to imitate the target singer and sing the songs of the target singer. That is, the current target person sings the songs of the target singer in the concert room in the position of the anchor, and the singing content is distributed in real time so that the audience can watch it. The audience watches the singing content live-streamed by the current target person through the concert interface displayed on the web page or the concert room displayed by the client. That is, any user who enters the concert room or browses the concert interface on the live-streaming web page can watch the singing content of the current target person singing the songs of the target singer in the concert room. In actual operation, the concert room can be produced immediately or reserved for production. In the case of immediate production, as shown in Figure 13, the terminal responds to the concert production command, generates a production request, and sends it to the server (i.e., the background server of the client). The server produces the corresponding concert room based on the production request and returns the room ID of the concert room to the terminal. The terminal accesses and presents the produced concert room based on the room ID. In the case of reserved production, as shown in Figure 14, the terminal responds to the concert production command, generates a production request including concert reservation information, and sends it to the server. The server produces the corresponding concert room based on the production request and returns the room ID of the concert room to the terminal. When the live-streaming start time arrives, the terminal accesses and presents the produced concert room based on the room ID.
[0095] In actual operation, when a concert hall is created, the terminal of the current target person shares the room ID, concert information, or concert reservation information of the concert hall with the terminals of other target persons, so as to widely publicize and recommend the concert corresponding to the target singer held by the current target person. As a result, the terminals of other target persons are allowed to enter the concert hall created by the current target person based on the room ID, attracting more users to watch the online virtual concert created by the current target person online, enabling the created virtual concert to reach more people, and ultimately guiding more users to practice the songs of the target singer or other singers, thereby improving the user retention rate.
[0096] Step 104: Collect the singing content in which the current target person imitates and sings the songs of the target singer, and play the singing content through the concert hall.
[0097] The singing content is provided so that the terminal corresponding to the target person in the concert hall can play it through the concert hall. The singing content includes the audio content of singing the songs of the target singer, and the audio content can be obtained in the following way. That is, collect the singing voice of the current target person singing the songs of the target singer, perform voice conversion on the singing voice to obtain the converted voice with the voice quality of the target singer corresponding to the singing voice, and use the converted voice as the audio content in the singing content.
[0098] In actual operation, when holding a virtual concert, it is necessary to perform pseudo-real-time singing voice conversion using a voice conversion service. For example, when the current target person sings a song in the concert hall, a hardware microphone is used to collect the source voice stream of the singing in real time, and the collected source voice stream is transmitted to the voice conversion service in the form of a queue. After the voice conversion service performs voice conversion (such as voice quality conversion) on the source voice stream, the converted target voice stream is also output to the virtual microphone in the concert hall at a constant speed in the form of a queue. The target voice stream is played back in a live broadcast manner in the concert hall through the virtual microphone to achieve the purpose of playing back the singing content.
[0099] For example, if the current target person is holding a virtual concert of singer A and imitating the singing of singer A's song, the terminal collects the singing voice (source voice stream) of the current target person singing the song, performs voice quality conversion on the singing voice, obtains the converted voice (target voice stream) corresponding to the voice quality of singer A, and plays back the converted voice through the concert hall. By doing so, what other users hear is a voice that is relatively close to or almost the same as the voice quality of singer A, and the reproduction of the target singer's concert is realized.
[0100] In addition to singing voices (sounds), the singing content may also include screen content. In FIG. 13 or FIG. 14, when the current target person sings a song of the target singer in the concert hall, the related singing content is played through the concert hall. For example, in addition to the singing voice of the current target person, a virtual stage, virtual audience, virtual background, etc. are further presented. A virtual figure corresponding to the target singer may be presented on the virtual stage, or the original figure of the current target person or a virtual figure corresponding to the current target person may be presented. The virtual audience represents other target persons who have entered the concert hall and are watching the concert, and can be displayed in the form of a virtual figure. The virtual background may be a screen related to the currently sung song, for example, a singing screen (MV screen or real concert screen) where the target singer sang the current song in the past, or a real screen where the current target person is currently singing.
[0101] In some embodiments, while the terminal is playing the singing content through the concert hall, the concert hall may present interaction information of other target persons with respect to the singing content. As shown in FIG. 13 or Figure 14 In addition to playing the related singing content through the concert hall, interaction information of other target persons who have entered the concert hall with respect to the current singing content, such as publicly displayed bullet-point comments and "likes", can be presented. This enriches the concert playback content, while at the same time better conveying the emotions towards the target singer, providing more entertainment options for the user, and meeting the increasing demand of the user for information diversification.
[0102] Note that for the user information related to the embodiments of the present application, such as the practice voice of the current target person, concert-related information (such as concert ID, singing content, etc.) and other relevant data such as the interaction information of other target persons, when the embodiments of the present application are applied in a specific product or technology, it is necessary to obtain the permission or consent of the user. Furthermore, regarding the collection, use, and processing of relevant data, the relevant laws, regulations, and standards of the relevant country or region must be complied with.
[0103] Next, an application example in an actual application scenario of the embodiments of the present application will be described. Please refer to FIG. 15. FIG. 15 is a schematic diagram of the singing sound change provided in the embodiments of the present application. In the related art, after recording a song, the user can add echo or perform various personalized sound change (voice conversion) processes, and can enjoyably participate in the recording, publication, sharing, etc. of the song even if they cannot sing. However, the sound change function in the related art only supports four sound change functions: original, electronic sound, metal, and harmony. The function is fixed and the sound change effect is also limited, so the sound change has to be carried out directly. Since algorithm verification and user verification cannot be performed afterwards, the sound change effect cannot be recognized and cannot be continuously optimized. Moreover, in the above-mentioned sound change function, the user can simply perform random singing, but cannot produce or hold a virtual concert of a specific singer. Also, the related technology is a voice conversion technology based on CycleGAN (Cycle Generative Adversarial Networks). CycleGAN includes two generators and two discriminators. In the scenario of voice conversion, the two generators are respectively responsible for the conversion from speaker A to speaker B and the conversion from speaker B to speaker A. And the two discriminators are respectively responsible for determining whether the voice is the voice of speaker A and whether the voice is the voice of speaker B. By connecting the two generators cyclically and connecting the corresponding discriminators, adversarial training can be performed. However, this network architecture can only perform one-to-one voice conversion and cannot convert the voice of any speaker to a specific speaker.
[0104] Therefore, the method for processing a virtual concert provided in the embodiments of the present application can produce and hold a virtual concert for a specific target singer based on the one-to-many voice conversion technology, and realize the reproduction of the target singer's concert. Such a performance method contributes to better conveying the user's emotions towards the target singer, provides more entertainment options for the user, and can meet the increasing diverse information requirements of the user.
[0105] Please refer to FIG. 16. FIG. 16 is a schematic flowchart of the method for processing a virtual concert provided in the embodiments of the present application. The method for processing a virtual concert provided in the embodiments of the present application includes the following steps.
[0106] Step 201: The terminal presents a song practice entry on the song practice interface.
[0107] Step 202: In response to a trigger operation on the song practice entry, a singer selection interface including at least one candidate singer is presented.
[0108] Step 203: In response to a selection operation on the target singer among at least one candidate singer, at least one candidate song corresponding to the target singer is presented.
[0109] Step 204: In response to a selection operation on the target song among at least one candidate song, a voice recording entry for singing the target song is presented.
[0110] Step 205: In response to a trigger operation on the voice recording entry, a song practice instruction for the target song of the target singer is received.
[0111] Step 206: In response to the song practice instruction, practice voice of the current subject practicing the song of the target singer is collected.
[0112] Of course, if the current user stops practicing halfway, they exit from the song practice interface.
[0113] Step 207: Present machine scoring corresponding to the practice voice.
[0114] Step 208: Determine whether the machine scoring has reached the scoring threshold.
[0115] Here, the current subject can judge by themselves how much room for improvement there is in the voice quality score and the emotion score based on the practice voice quality (i.e., the voice after conversion) of each practice voice. By practicing multiple times the imitation of the singing skills, emotional richness, breath control, key change, etc. of the original target singer, the machine scoring such as the voice quality score and the emotion score can be improved. If the machine scoring has reached the scoring threshold (for example, setting the scoring threshold to 80 points out of 100 full marks), step 209 is executed. If the machine scoring has not reached the scoring threshold, step 205 is executed.
[0116] Step 209: Cast the practice voice into the voting pool corresponding to the target singer for manual scoring.
[0117] Here, the practice voice to be scored is cast into the voting pool corresponding to the target singer, and the practice voice is pushed to the terminals of other subjects. Other subjects score the practice voice of the current subject through the scoring entrance presented on other terminals, and return the obtained manual scoring so that it can be displayed on the terminal of the current subject.
[0118] Step 210: Present the manual scoring corresponding to the practice voice.
[0119] Here, the manual scoring may also be evaluated from two aspects: voice quality similarity and emotion similarity.
[0120] Step 211: Perform an averaging process on the machine scoring and the manual scoring to obtain the practice score corresponding to the practice voice and the song practice ranking corresponding to the practice song of the current subject.
[0121] The practice score corresponding to the practice voice is (machine scoring (voice quality score, emotion score) + human scoring (voice quality score, emotion score)) / 4. Taking Song B as an example, in the corresponding machine scoring, the voice quality score = 80 points, the emotion score = 75 points, in the human scoring, the voice quality score = 78 points, the emotion score = 70 points, and the practice score of this song is (80 + 75 + 78 + 70) / 4 = 75.75 points.
[0122] Here, when there are multiple people who have practiced the songs of the target singer, a descending song practice ranking is determined in the order from the highest to the lowest of the practice scores of the users who have practiced the target singer, and the song practice ranking of the current target person in the song practice ranking is determined.
[0123] Step 212: Determine whether the song practice ranking is ahead of the target ranking.
[0124] For example, when there are multiple users who have practiced the songs of Singer A, a descending song practice ranking is determined according to the practice scores of each user. Assuming that only the top 3 users have the production qualification to produce a concert of Singer A, it is determined whether the song practice ranking of the current target person is among the top 3 (that is, whether it is ahead of the 4th place) based on the practice score of the current target person. If the song practice ranking of the current target person is ahead of the 4th place, step 213 is executed. Otherwise, step 201 is executed.
[0125] Step 213: Present a concert entrance for producing a concert of the target singer.
[0126] In actual operation, the concert entrance and the song practice entrance may be the same entrance or may not be the same entrance. When both are the same entrance, if the current target person has the production qualification to produce a concert, display information for indicating that the current target person has the production qualification to produce a concert is presented in the relevant area of the song practice entrance (for example, indicated by a "red dot" at the song practice entrance).
[0127] Step 214: Present prompt information for reminding whether to apply for the production of the concert of the target singer in response to the trigger operation for the concert entrance.
[0128] Step 215: When a decision operation for the prompt information is received, receive a concert production instruction for the target singer.
[0129] Here, the current target person can decide whether to produce a concert corresponding to the target singer based on the prompt information. If the current target person decides to produce a concert corresponding to the target singer, the decision operation can be triggered by triggering the corresponding decision button. When the terminal receives the decision operation, it can receive a concert production instruction corresponding to the target singer. On the other hand, if the current target person decides not to produce a concert corresponding to the target singer, the cancellation operation can be triggered by triggering the corresponding cancellation button. When the terminal receives the cancellation operation, it will not receive a concert production instruction corresponding to the target singer. At this time, a song practice entrance can be presented on the song practice interface, and the current target person can practice the songs of the target singer or the songs of other singers through the song practice entrance.
[0130] Step 216: In response to the concert production instruction, create a concert room for imitating and singing the songs of the target singer.
[0131] The concert room is used for the current target person to imitate the target singer and sing the songs of the target singer. Any user who enters the concert room can view the singing content of the current target person singing the songs of the target singer in the concert room.
[0132] Step 217: Collect the singing content corresponding to the current target person's imitation singing of the songs of the target singer, and play the singing content through the concert room.
[0133] Now, please refer to FIG. 17. FIG. 17 is a processing flowchart of the virtual concert provided in the embodiment of the present application. To hold a virtual concert, it is necessary to perform pseudo real-time singing voice conversion by using the voice conversion service of voice processing software. For example, when the current target person sings a song in the concert hall, a hardware microphone is used to collect the source voice stream of the singing in real time, and the collected source voice stream is transmitted to the voice conversion service in the form of a queue. After the voice conversion service performs voice conversion on the source voice stream, the converted target voice stream is also output to the virtual microphone in the concert hall at a constant speed in the form of a queue. The target voice stream is played in the concert hall in a live broadcast manner through the virtual microphone to achieve the purpose of playing the singing content.
[0134] Next, machine scoring will be described. When the user's practice is completed, the terminal loads the voice conversion service, performs voice quality conversion on the practice voice collected by the voice conversion technology, converts the collected practice voice into a voice quality similar to that of the original target singer, and obtains the practice voice quality corresponding to the target singer. Then, the practice voice quality is compared with the original voice quality of the target singer to obtain the corresponding voice quality similarity, and a voice quality score is determined based on the voice quality similarity. At the same time, the emotional degree of the practice voice is identified to obtain the corresponding practice emotional degree, the practice emotional degree is compared with the original emotional degree of the target singer to obtain the corresponding emotional similarity, and an emotional score is determined based on the emotional similarity. The voice quality score and the emotional score are used for machine scoring.
[0135] Please refer to FIG. 18. FIG. 18 is a schematic diagram of voice quality conversion provided in the embodiment of the present application. When performing voice quality conversion on the practice voice, the phoneme identification model is used to perform phoneme identification on the practice voice to obtain the corresponding phoneme sequence. The sound loudness of the practice voice is identified to obtain the corresponding sound loudness characteristics. The melody of the practice voice is recognized to obtain a sine excitation signal representing the melody. The phoneme sequence, the sound loudness characteristics, and the sine excitation signal are combined and processed by a sound wave synthesizer to obtain the practice voice quality corresponding to the target singer.
[0136] The phoneme recognition module, also called the PPG extractor, is part of the ASR model. The function of the ASR model is to convert speech into text. In essence, it first converts speech into a phoneme sequence and then converts the phoneme sequence into text. The function of the PPG extractor is to first convert speech into a phoneme sequence and is used to extract information irrelevant to voice quality, such as text content information, from the training speech.
[0137] Please refer to FIG. 19. FIG. 19 is a structural schematic diagram of the phoneme recognition model provided in the embodiment of the present application. Before performing voice quality recognition, considering that the training speech is actually a chaotic waveform signal in the time domain, in order to make it easier to analyze, the training speech in the time domain is converted into the frequency domain by fast Fourier transform to obtain the speech spectrum corresponding to the speech data. Based on the obtained speech spectrum, the degree of difference between the speech spectra corresponding to adjacent sampling windows is obtained. Furthermore, based on the obtained multiple degrees of difference, the energy spectrum corresponding to each sampling window is specified, and finally, a spectrogram (such as a Mel spectrogram) corresponding to the training speech may be obtained. Then, downsampling processing is performed on the spectrogram corresponding to the training speech in the downsampling layer. The downsampling layer has a two-dimensional convolutional structure and downsamples the input spectrogram at a time scale of 2 to obtain downsampling characteristics. Then, the downsampling characteristics are input into an encoder (an integrated encoder or a Transformer encoder) for encoding processing to obtain corresponding encoded characteristics. Then, by inputting the encoded characteristics into a decoder for decoding processing, the phoneme sequence of the training speech is predicted. Here, the decoder may be a CTC decoder. The decoder includes a single fully connected layer, and the decoding process is as follows. That is, based on the encoded characteristics, the phoneme with the maximum probability for each frame of the training speech is screened, and the corresponding phonemes with the maximum probability for each frame of the screened training speech are used to form a time-series phoneme sequence. In the time-series phoneme sequence, adjacent identical phonemes are integrated to obtain a phoneme sequence.
[0138] When obtaining the spectrogram of the practice voice, the practice voice is divided into frames, and the signal of each frame is Fourier-transformed to obtain the spectrum. Then, they are superimposed in the time domain to obtain the spectrogram. The spectrogram can reflect the time-varying changes of the sine waves superimposed in the voice signal in the time domain. Alternatively, based on the spectrogram, a set spectrogram filter is used to filter the spectrum to obtain the mel spectrogram. Compared with the general spectrogram, the number of frequency dimensions is smaller, and it focuses on the voice signals in the low-frequency band to which the human auditory sense is sensitive. Generally, the Mel diagram is easier to extract / separate information than the voice signal, and it is also easier to modify the voice.
[0139] When training the phoneme recognition model, training is performed using a large number of voice-text training samples. The CTC loss as shown in the following formula can be used as the loss function for training.
Number
Number
[0140] The voice loudness characteristic is the time series of the loudness of the practice voice for each frame in the practice voice, that is, the maximum amplitude corresponding to the practice voice of each frame obtained by performing a short-time Fourier transform on the practice voice. The sine excitation signal is calculated using the fundamental frequency of the sound (F0, the fundamental frequency of each frame of the sound is equal to the pitch of each frame of the sound).
[0141] The purpose of the sound wave synthesizer is to synthesize three features independent of the speaker's voice quality, namely the phoneme sequence of the practice voice, the sound loudness characteristic, and the sine excitation signal, and use the voice quality of the target singer to obtain the sound wave of the sung voice (i.e., the practice voice quality corresponding to the above-mentioned target singer). Please refer to Figure 20. Figure 20 is a schematic structural diagram of the sound wave synthesizer provided in the embodiment of the present application. The sound wave synthesizer includes a plurality of upsampling blocks and downsampling blocks. In order to convert the practice voice into the practice voice quality (i.e., sound wave) corresponding to the target singer, four upsampling blocks are applied to the above-mentioned obtained phoneme sequence, and upsampling processing is sequentially performed with coefficients of 4, 4, 4, and 5. Four downsampling blocks are respectively applied to perform downsampling processing on the above-mentioned sound loudness characteristic and sine excitation signal with coefficients of 4, 4, 4, and 5. The features obtained by the processing are combined to obtain the practice voice quality corresponding to the target singer. As shown in Figure 21, Figure 21 is provided in the embodiment of the present application down It is a schematic structural diagram of the sampling block. The obtained phoneme sequence is down Input into the sampling block, down Through sampling, a plurality of layers of activation functions, and convolution processing, the corresponding down Sampling characteristics are obtained. As shown in Figure 22, Figure 22 is provided in the embodiment of the present application up It is a schematic structural diagram of the sampling block. The obtained sound loudness characteristic and sine excitation signal are up Input into the sampling block, up Through sampling, a plurality of layers of activation functions, convolution processing, and the processing of the Feature-wise Linear Modulation (FiLM) module, the corresponding up Sampling characteristics are obtained. The Feature-wise Linear Modulation (FiLM) module is used for feature affine. The information of the sine excitation signal and the sound loudness characteristic is combined with the phoneme sequence, thereby generating the scaling vector and shift vector given to the input. As shown in Figure 23, Figure 23 is a schematic overview diagram of the characteristic linear modulation module provided in the embodiment of the present application. The FiLM module corresponds to the upIt has the same number of convolutional channels as the sampling block.
[0142] When training the sound wave synthesizer, a self-repair training method can be adopted. That is, the singing voices of a large number of target speakers are used as training voices. From these voices, a phoneme sequence, a pitch loudness characteristic, and a sine excitation signal are separated and used as the input of the sound wave synthesizer, and the voice itself is used as the predicted output of the sound wave synthesizer for training. The target loss function for training is as follows.
Number
Number
Number
Number
[0143] Once the practice voice quality of the practice voice is obtained by the above method, the practice voice quality and the original voice quality are compared, and an appropriate voice quality score is determined based on the comparison result.
[0144] When determining the voice quality score, voice quality comparison may be performed based on a speaker identification model. The structure of the speaker identification model is as shown in FIG. 24, and FIG. 24 is a structural schematic diagram of the speaker identification model provided in the embodiments of the present application. The task trained by this model is a multi-class classification task, and speaker classification training is performed using six fully connected layers. The training source voice is data with a large number of speakers labeled, the training target is the one-hot encoding of speaker classification, and the loss function uses the cross-entropy loss of the following formula.
Equation
Equation
[0145] By the above method, the current target person can produce or hold a virtual concert corresponding to the target singer. When the current target person sings the song of the target singer in the concert hall, the relevant singing content is played through the concert hall. For example, in addition to the singing voice of the current target person, at least one of a virtual stage, virtual audience, and virtual background is presented. A virtual portrait corresponding to the target singer may be presented on the virtual stage, or the original portrait of the current target person or a virtual portrait corresponding to the current target person may be presented. The virtual audience represents other target persons who enter the concert hall and watch the concert, and can be displayed in the form of a virtual portrait. The virtual background may be a screen related to the currently sung song, such as a singing screen (MV screen or real concert screen) where the target singer sang the current song in the past, or a real screen where the current target person is singing. In addition, interaction information of other target persons who have entered the concert hall with the current singing content, such as publicly displayed bullet comments and "likes", can be presented. In this way, while enriching the concert playback content, the emotion towards the target singer can be better conveyed, more entertainment options can be provided to users, and the increasing demand for information diversification of users can be satisfied.
[0146] The virtual concert processing method provided in the embodiments of the present application can also be applied to game scenarios. For example, when a user or player is playing a game on a live streaming client, an interface for practicing the song of the current target person is presented, and a concert entrance is presented on the song practice interface. Based on the concert entrance, a concert production command for the target singer is received. In response to the concert production command, a concert room for mimicking and singing the song of the target singer is created. Singing content corresponding to the mimic singing of the song of the target singer by the current target person is collected, and the singing content is played back through the concert room so that the terminals corresponding to other players or users in the concert room can play back the singing content through the concert room.
[0147] Subsequently, an exemplary structure of the virtual concert processing apparatus 555 provided in the embodiments of the present application implemented as a software module will be described. In some embodiments, the software module in the virtual concert processing apparatus 555 stored in the memory 550 of FIG. 2 may include a command receiving module 5551 that receives a concert production command for the target singer based on the presented concert entrance, a room creating module 5552 that creates a concert room for mimicking and singing the song of the target singer in response to the concert production command, and a singing playback module 5553 that collects the singing content of the current target person mimicking and singing the song of the target singer and plays back the singing content through the concert room. The singing content is used to be played back on the terminals of each target person in the concert room.
[0148] In some embodiments, the apparatus further includes an entrance presentation module, which presents a music practice entrance on the music practice interface, receives a music practice command for the target singer based on the music practice entrance, collects practice audio of the current subject practicing the music of the target singer in response to the music practice command, and when it is determined based on the practice audio that the current subject has the qualification to produce a concert of the target singer, presents a concert entrance associated with the target singer on the music practice interface corresponding to the current subject.
[0149] In some embodiments, the entrance presentation module further presents a singer selection interface including at least one candidate singer in response to a trigger operation on the music practice entrance, presents at least one candidate song corresponding to the target singer in response to a selection operation on the target singer among the at least one candidate singer, presents an audio recording entrance for singing the target song in response to a selection operation on the target song among the at least one candidate song, and receives a music practice command for the target song of the target singer in response to a trigger operation on the audio recording entrance.
[0150] In some embodiments, the apparatus further includes a first qualification determination module, which presents a practice score corresponding to the practice audio, and when the practice score reaches a target score, determines that the current subject has the qualification to produce a concert of the target singer.
[0151] In some embodiments, the apparatus further includes a first score acquisition module, which, when the number of practiced songs is at least two, presents a practice score corresponding to the practice audio of each song of the current subject, obtains the singing difficulty of each song, determines a weight corresponding to the song based on the singing difficulty, and weighted-averages the practice scores of the practice audio of each song based on the weight to obtain the practice score of the practice audio.
[0152] In some embodiments, the practice score includes at least one of a voice quality score and an emotion score, and the device further includes a second score acquisition module. When the practice score includes a voice quality score, the second score acquisition module performs voice quality conversion on the practice voice to obtain a practice voice quality corresponding to the target singer, compares the practice voice quality with the original voice quality of the target singer to obtain a corresponding voice quality similarity, and determines the voice quality score based on the voice quality similarity. When the practice score includes the emotion score, the second score acquisition module performs emotion degree identification on the practice voice to obtain a corresponding practice emotion degree, compares the practice emotion degree with the original emotion degree when the target singer sings the song to obtain a corresponding emotion similarity, and determines the emotion score based on the emotion similarity.
[0153] In some embodiments, the second score acquisition module further performs phoneme identification on the practice voice by a phoneme identification model to obtain a phoneme sequence, performs sound loudness identification on the practice voice to obtain sound loudness characteristics, performs melody recognition on the practice voice to obtain a sine excitation signal representing the melody, and combines the phoneme sequence, the sound loudness characteristics, and the sine excitation signal with a sound wave synthesizer to obtain a practice voice quality corresponding to the target singer.
[0154] In some embodiments, the apparatus further includes a third score acquisition module. The third score acquisition module transmits the practice voice to the terminal of another subject, causes the terminal of the other subject to obtain a manual score corresponding to the input practice voice based on a scoring entry corresponding to the practice voice, receives the manual score returned from the other terminal, and determines a practice score corresponding to the practice voice based on the manual score.
[0155] In some embodiments, the third score acquisition module further acquires a machine score corresponding to the practice voice. When the machine score reaches a scoring threshold, the practice voice is transmitted to the terminal of another subject, and an averaging process is performed on the machine score and the human score to obtain a practice score corresponding to the practice voice.
[0156] In some embodiments, the apparatus further includes a second qualification determination module. The second qualification determination module presents a song practice ranking of the song corresponding to the current subject. When the song practice ranking is located before the target ranking, it is determined that the current subject has the production qualification to produce a concert of the target singer.
[0157] In some embodiments, the apparatus further includes a detailed check module. When the number of songs practiced is at least two, the detailed check module presents the total score of the current subject singing the at least two songs and a detailed entry for checking the score details of each song. In response to a trigger operation on the detailed entry, a detailed page is presented, and the practice score corresponding to each song is presented on the detailed page.
[0158] In some embodiments, the command receiving module further presents a singer selection interface including at least one candidate singer in response to a trigger operation on the concert entry. In response to a selection operation on the target singer among the at least one candidate singer, when it is determined that the current subject has the production qualification to produce a concert of the target singer, a concert production command corresponding to the target singer is received.
[0159] In some embodiments, the command receiving module further presents a singer selection interface including at least one candidate singer having the production qualification for the current target person to produce a concert in response to a trigger operation on the concert entrance, and receives a concert production command for the target singer in response to a selection operation on the target singer among the at least one candidate singer.
[0160] In some embodiments, when a target singer is associated with the concert entrance, the command receiving module further presents prompt information for reminding whether to apply for the production of a concert corresponding to the target singer in response to a trigger operation on the concert entrance, and receives a concert production command for the target singer when receiving a decision operation on the prompt information.
[0161] In some embodiments, when receiving a decision operation on the prompt information, the command receiving module further presents an application interface for applying for the production of the target singer's concert, presents an editing entry for editing the relevant information of the concert on the application interface, and receives a concert production command for the target singer in response to a decision operation on the concert information when receiving the concert information edited based on the editing entry.
[0162] In some embodiments, the command receiving module further presents a reservation entry for reserving the production of a concert hall while presenting the prompt information, and in response to a trigger operation on the reservation entry, presents a reservation interface for reserving the production of a concert by the target singer, presents an editing entry for editing the concert reservation information on the reservation interface, receives concert reservation information including at least the concert start time edited based on the editing entry, receives a concert production command corresponding to the target singer in response to a determination operation on the concert reservation information, the room production module further produces a concert hall for imitating and singing the songs of the target singer in response to the concert production command, and accesses and presents the concert hall when the concert start time is reached.
[0163] In some embodiments, the apparatus further includes a concert cancellation module, and when the concert cancellation module receives a cancellation operation on the prompt information, it presents a song practice entry for practicing the songs of the target singer or the songs of other singers on the song practice interface.
[0164] In some embodiments, when the number of the concert entries is at least one, a singer is associated with the concert entry, and there is a corresponding relationship between the singer associated with the concert entry and the concert entry. The command receiving module further receives a concert production command corresponding to the target singer in response to a trigger operation on the concert entry associated with the target singer.
[0165] In some embodiments, the apparatus further includes an interaction module, and the interaction module presents interaction information of other subjects on the singing content in the concert hall while the singing content is being played in the concert hall.
[0166] In some embodiments, the singing content includes audio content of singing a song of the target singer, and the singing playback module further collects singing voice of the current target person singing a song of the target singer, converts the voice quality of the singing voice to obtain converted voice of the voice quality of the target singer corresponding to the singing voice, and uses the converted voice as the audio content in the singing content.
[0167] Embodiments of the present application provide a computer program product or a computer program including computer instructions stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium and the processor executes the computer instructions, the computer device is caused to execute the processing method of the virtual concert in the embodiments of the present application.
[0168] Embodiments of the present application provide a computer-readable storage medium storing executable instructions that, when executed by a processor, cause the processor to execute the processing method of the virtual concert provided in the embodiments of the present application, for example, the method shown in FIG. 3.
[0169] In some embodiments, the computer-readable storage medium may be a memory such as FRAM (registered trademark) (Ferroelectric Random Access Memory), ROM (Read Only Memory), PROM (Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic memory, optical disk, or CD-ROM, or may be various devices including one or any combination of the above memories.
[0170] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any programming language (compiled or interpreted language, or declarative or procedural language), and may further be arranged in any form including as a stand-alone program or in the form of modules, components, subroutines, or other units suitable for use in a computer environment.
[0171] As an example, the executable instructions may correspond to files in a file system and may be stored as part of a file that stores other programs or data, for example, stored in one or more scripts in an HTML (Hyper Text Markup Language) document, stored in a single file dedicated to that program, or stored in multiple collaborative files (for example, files storing one or more modules, subprograms, code portions), but not necessarily so.
[0172] As an example, the executable instructions may be arranged to be executed on one computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected by a communication network.
[0173] The above are only examples of the present application and are not intended to limit the protection scope of the present application. Any changes, equivalent substitutions, improvements, etc. made within the spirit and scope of the present application are all included within the protection scope of the present application.
Claims
1. A method for processing a virtual concert executed by an electronic device, comprising: receiving a concert production instruction for a target singer in response to a trigger operation of a current target person with respect to a presented concert entrance; producing a concert hall for mimicking and singing the song of the target singer in the virtual concert in response to the concert production instruction; collecting singing content in which the current target person mimics and sings the song of the target singer, and playing the singing content through the concert hall; The step of receiving the concert production instruction for the target singer in response to the trigger operation of the current target person with respect to the presented concert entrance includes: presenting a singer selection interface including at least one candidate singer in response to the trigger operation with respect to the presented concert entrance; receiving the concert production instruction for the target singer in response to a selection operation for the target singer among the at least one candidate singer; The received concert production instruction is a concert production instruction for the target singer who has the production qualification for the current target person to produce the virtual concert; The singing content is used to be played on the terminal of each target person in the concert hall. A method for processing a virtual concert.
2. Before the step of receiving the concert production instruction for the target singer, further comprising: presenting a song practice entrance on the song practice interface of the current target person; receiving a song practice instruction for the target singer in response to a trigger operation of the current target person with respect to the song practice entrance. In response to the singing practice instruction, collecting practice audio of the current target person singing the song of the target singer; Based on the practice audio, if it is determined that the current target person has the production qualification to produce the virtual concert of the target singer, presenting the concert entrance associated with the target singer on the song practice interface; The method for processing a virtual concert according to claim 1.
3. The step of receiving the singing practice instruction for the target singer in response to the trigger operation of the current target person on the song practice entrance includes: Presenting a singer selection interface including at least one candidate singer in response to the trigger operation on the song practice entrance; Presenting at least one candidate song corresponding to the target singer in response to a selection operation on the target singer among the at least one candidate singer; Presenting an audio recording entrance for singing the target song in response to a selection operation on the target song among the at least one candidate song; Receiving the singing practice instruction for the target song of the target singer in response to the trigger operation of the current target person on the audio recording entrance. The method for processing a virtual concert according to claim 2.
4. Before the step of presenting the concert entrance associated with the target singer on the song practice interface, further including: Presenting a practice score obtained by grading the practice audio; Determining that the current target person has the production qualification to produce the virtual concert of the target singer when the practice score reaches the target score. The method for processing a virtual concert according to claim 2.
5. Before the step of presenting the practice score obtained by scoring the practice voice, further, When the number of songs practiced by the current subject is at least two, presenting a practice score corresponding to the practice voice of each song of the current subject; Obtaining the singing difficulty of each song and determining a weight corresponding to the song based on the singing difficulty; Based on the weight, weighted-averaging the practice scores of each practice voice to obtain the practice score of the practice voice, including: The method for processing a virtual concert according to claim 4.
6. The practice score includes at least one of a voice quality score and an emotion score. Before the step of presenting the practice score obtained by scoring the practice voice, further, When the practice score includes the voice quality score, performing voice quality conversion on the practice voice to obtain a practice voice quality corresponding to the target singer, comparing the practice voice quality with the original voice quality of the target singer singing the song to obtain a voice quality similarity, and determining the voice quality score based on the voice quality similarity; When the practice score includes the emotion score, performing emotion identification on the practice voice to obtain a practice emotion degree, comparing the practice emotion degree with the original emotion degree of the target singer singing the song to obtain an emotion similarity, and determining the emotion score based on the emotion similarity, including: The method for processing a virtual concert according to claim 4.
7. The step of performing voice quality conversion on the practice voice to obtain the practice voice quality corresponding to the target singer is: Performing phoneme identification on the practice voice by a phoneme identification model to obtain a phoneme sequence; Performing sound loudness identification on the practice voice to obtain a sound loudness characteristic; Performing melody recognition on the practice voice to obtain a sine excitation signal for representing the melody; combining, by a sound synthesizer, the phoneme sequence, the sound loudness characteristic, and the sine excitation signal to obtain the practice voice quality corresponding to the target singer; The method for processing a virtual concert according to claim 6.
8. Before the step of presenting the practice score obtained by scoring the practice voice, further comprising: transmitting the practice voice to a terminal of another target person, and causing the terminal of the other target person to obtain an artificial score corresponding to the input practice voice in response to a trigger operation of the current target person for the scoring entry of the practice voice; receiving the artificial score returned from the terminal of the other target person, and determining a practice score of the practice voice based on the artificial score. The method for processing a virtual concert according to claim 4.
9. The step of transmitting the practice voice to the terminal of the other target person includes: obtaining a machine score corresponding to the practice voice, and when the machine score reaches a scoring threshold value, transmitting the practice voice to the terminal of the other target person. The step of determining the practice score of the practice voice based on the artificial score includes: obtaining an average of the machine score and the artificial score to obtain the practice score of the practice voice. The method for processing a virtual concert according to claim 8.
10. Before the step of presenting the concert entry associated with the target singer on the song practice interface, further comprising: presenting a song practice ranking of the song corresponding to the current target person; when the song practice ranking is located before the target ranking, determining that the current target person has the production qualification to produce the virtual concert of the target singer. The method for processing a virtual concert according to claim 2.
11. When the number of the practiced songs is at least two, presenting a total score of the at least two songs sung by the current target person and a detailed entry for checking score details of each of the songs; responding to a trigger operation of the current target person on the detailed entry, presenting a detailed page and presenting practice scores of each of the songs on the detailed page, The method for processing a virtual concert according to claim 10.
12. The step of receiving the concert production instruction for the target singer in response to the trigger operation of the current target person on the presented concert entry is: presenting a singer selection interface including at least one candidate singer in response to the trigger operation on the presented concert entry; and in response to a selection operation on the target singer among the at least one candidate singer, receiving the concert production instruction for the target singer when it is determined that the current target person has the production qualification to produce the virtual concert of the target singer. The method for processing a virtual concert according to claim 1.
13. The step of receiving the concert production instruction for the target singer in response to the trigger operation of the current target person on the presented concert entry is: presenting a singer selection interface including at least one of the candidate singers among the candidate singers for whom the current target person has the production qualification to produce the virtual concert in response to the trigger operation on the presented concert entry; and in response to a selection operation on the target singer among the at least one candidate singer, receiving the concert production instruction for the target singer. The method for processing a virtual concert according to claim 1.
14. The step of receiving the concert production instruction for the target singer in response to the trigger operation of the current target person for the presented concert entrance is When the target singer is associated with the concert entrance, presenting prompt information for reminding to apply for the production of the virtual concert corresponding to the target singer in response to the trigger operation for the concert entrance; When receiving a decision operation for the prompt information, the step of receiving the concert production instruction for the target singer, including The method for processing a virtual concert according to claim 1.
15. When receiving the decision operation for the prompt information, the step of receiving the concert production instruction for the target singer is When receiving the decision operation for the prompt information, presenting an application interface for applying for the production of the virtual concert of the target singer and presenting an editing entrance for editing the related information of the virtual concert on the application interface; The step of receiving the concert information edited based on the editing entrance; When receiving a decision operation for the concert information, the step of receiving the concert production instruction for the target singer, including The method for processing a virtual concert according to claim 14.
16. When receiving the decision operation for the prompt information, the step of receiving the concert production instruction for the target singer is The step of presenting a reservation entrance for reserving the production of the concert room; In response to a trigger operation of the current target person on the reservation entry, presenting a reservation interface for reserving the production of the virtual concert of the target singer, and presenting an editing entry for editing the reservation information of the virtual concert on the reservation interface; Receiving concert reservation information including at least the concert start time, which is edited based on the editing entry; In response to a determination operation on the concert reservation information, receiving a concert production instruction for the target singer, including: In response to the concert production instruction, the step of producing a concert room for mimicking and singing the songs of the target singer is: In response to the concert production instruction, producing the concert room for mimicking and singing the songs of the target singer, and when the concert start time is reached, accessing and presenting the concert room, including the steps of: The method for processing a virtual concert according to claim 14.
17. When receiving a cancellation operation on the prompt information, presenting a song practice interface and presenting a song practice entry on the song practice interface, including: The song practice entry is used for singing practice of the songs of the target singer or other singers of the target singer. The method for processing a virtual concert according to claim 14.
18. When the number of the concert entries is at least one, a singer is associated with the concert entry, and there is a corresponding relationship between the singer associated with the concert entry and the concert entry. In response to the trigger operation of the current target person on the presented concert entry, the step of receiving the concert production instruction corresponding to the target singer is: In response to the trigger operation on the concert entrance associated with the target singer, receiving a concert production instruction corresponding to the target singer, including the step of The method for processing a virtual concert according to claim 1.
19. While the singing content is being played in the concert hall, presenting interaction information of other target persons in the concert hall with respect to the singing content of the current target person, including the step of The method for processing a virtual concert according to claim 1.
20. The singing content includes audio content that imitates and sings the songs of the target singer, and the step of collecting the singing content in which the current target person imitates and sings the songs of the target singer is Collecting the singing voice in which the current target person imitates and sings the songs of the target singer; and Converting the singing voice to obtain a converted voice with the voice quality of the target singer corresponding to the singing voice, and using the converted voice as the audio content, including the steps of The method for processing a virtual concert according to claim 1.
21. A command receiving module that receives a concert production instruction regarding a target singer in response to a trigger operation of a current target person on a presented concert entrance; A room production module that produces a concert hall for imitating and singing the songs of the target singer in a virtual concert in response to the concert production instruction; A singing playback module that collects singing content in which the current target person imitates and sings the songs of the target singer and plays back the singing content through the concert hall, including The command receiving module is In response to the trigger operation on the presented concert entrance, presenting a singer selection interface including at least one candidate singer; In response to a selection operation on the target singer among the at least one candidate singer, receive the concert production instruction for the target singer. The received concert production instruction is a concert production instruction for the target singer who has the production qualification for the virtual concert by the current target person. The singing content is used to be played on the terminal of each target person in the concert hall. A processing device for a virtual concert.
22. A memory for storing executable instructions. When the executable instructions stored in the memory are executed, a processor for realizing the virtual concert processing method according to any one of Claims 1 to 20, including. An electronic device.
23. To cause a computer to realize the virtual concert processing method according to any one of Claims 1 to 20. A computer program.
Citation Information
Patent Citations
Musical sound generator
JP1998161682A
Singing ability evaluation and singer selection system using the Internet and the selection method
JP2004500662A
Karaoke device which performs scoring display with scoring ranking by musical piece
JP2005099288A
Evaluation device, control method, and program
JP2007233078A
Karaoke scoring method and karaoke scoring system
JP2009237503A