Joint singing method and device, computer equipment and readable storage medium
By uniformly collecting and synthesizing audio and images on the first display terminal in the vehicle, generating synthetic singing clips and joint singing video clips, the problem of audio and image delay in in-vehicle online karaoke is solved, and the real-time and synchronization of multi-user joint music singing is improved.
Patent Information
- Application Number
- CN202510738429.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
In existing in-car online karaoke technology, audio or picture playback delays are large and audio and picture are out of sync, resulting in poor multi-user joint music singing effects.
By setting up a first display terminal in the vehicle to uniformly collect and synthesize audio and images, generate a synthesized singing clip, and combine the singing evaluation information to generate a joint singing video clip, which is sent to multiple display terminals for display. This avoids the independent collection and synthesis method of each terminal, and improves the synthesis efficiency and synchronization of audio and images.
It reduces the delay of audio and picture playback, improves the real-time performance and singing effect of multi-user joint singing, and improves the synchronization of multi-user joint music singing.
Smart Images

Figure CN120673732A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a joint singing method, apparatus, computer equipment, readable storage medium, and computer program product. Background Art
[0002] In the music software space, the emergence of in-car karaoke software offers users a rich entertainment experience while driving. Traditional in-car karaoke software primarily operates in a standalone mode, limiting users to a single screen and lacking interactivity with other passengers or online entertainment. With the increasing demand for in-car entertainment, online karaoke functionality has emerged.
[0003] In the current in-vehicle online karaoke technology, in terms of audio and image acquisition and synthesis, the audio of each user is independently collected through the screen terminals in each passenger position inside the vehicle. Each screen needs to synthesize the audio and image separately and play them separately. In addition, each screen terminal will refer to the time of the driver and co-driver screen terminals to perform delay compensation to achieve playback synchronization.
[0004] It can be seen that the independent data collection, data synthesis and synchronization mechanisms in the above technologies lead to problems such as large audio or picture playback delays and audio and picture asynchrony in the online karaoke process, resulting in poor singing effects when multiple users perform joint music performances. Summary of the Invention
[0005] Based on this, it is necessary to provide a joint singing method, device, computer equipment, computer-readable storage medium and computer program product that can improve the singing effect of multi-user joint music singing in response to the above technical problems.
[0006] In a first aspect, an embodiment of the present application provides a joint singing method, which is applied to a first display terminal disposed in a vehicle, wherein the vehicle is further provided with at least two second display terminals, wherein the first display terminal is communicatively connected to each of the at least two second display terminals; the method comprises:
[0007] In response to a joint singing request for a target song, obtaining a song video clip of the target song at a current playback progress; the joint singing request is for at least two users in the vehicle to sing the target song together;
[0008] Sending a song video clip of the target song at the current playing progress to each of the at least two second display terminals for display to prompt the current singing progress of the target song;
[0009] Collecting singing voices of each user for the current singing progress, generating a synthesized singing voice segment of the current singing progress based on the singing voices of each user, and sending the synthesized singing voice segment of the current singing progress to an audio playback device provided in the vehicle for playback;
[0010] Generate singing evaluation information of the synthesized singing segment at the current singing progress; the singing evaluation information is used to characterize the singing effect of each user at the current singing progress of the target song; combine the singing evaluation information of the synthesized singing segment at the current singing progress and the song video segment of the target song at the next playback progress of the current playback progress, and generate a joint singing video segment of the target song at the next playback progress;
[0011] The joint singing video clip of the target song at the next playback progress is sent to each second display terminal of the at least two second display terminals for display, so as to prompt each user of the singing effect of the current singing progress of the target song and the next singing progress of the target song.
[0012] In one embodiment, obtaining a song video clip of the target song at the current playback progress includes: sending a resource download request for the target song to a server; receiving video data of the target song sent by the server according to the resource download request; and obtaining a song video clip of the target song at the current playback progress from the video data of the target song.
[0013] In one embodiment, the generating of the singing evaluation information of the synthesized singing segment of the current singing progress includes: identifying the singing characteristics of each of the users singing the target song from the synthesized singing segment of the current singing progress; and generating the singing evaluation information based on the singing characteristics of each of the users singing the target song.
[0014] In one embodiment, the step of obtaining a song video clip of the target song at the current playback progress in response to a joint singing request for the target song includes: obtaining a song video clip of the target song at the current playback progress in response to a joint singing request for the target song sent by any of the second display terminals.
[0015] In one embodiment, in response to a joint singing request for a target song, obtaining a song video clip of the target song at the current playback progress includes: sending a joint singing request for the target song to each of the second display terminals, and receiving a joint singing confirmation message sent by any of the second display terminals in response to the joint singing request; and obtaining a song video clip of the target song at the current playback progress in response to the joint singing confirmation message.
[0016] In one embodiment, the obtaining of a song video clip of the target song at the current playback progress in response to a joint singing request for the target song includes: obtaining current working status information corresponding to the vehicle in response to the joint singing request for the target song; the current working status information is used to characterize whether the driving status of the vehicle is safe and whether the equipment in the vehicle is operating normally; when the current working status information indicates that the driving status of the vehicle is safe and the equipment in the vehicle is operating normally, obtaining a song video clip of the target song at the current playback progress.
[0017] In one embodiment, sending the joint singing video clip of the target song at the next playback progress to each second display terminal of the at least two second display terminals for display includes: rendering the joint singing video clip of the target song at the next playback progress to obtain a joint singing video clip including a song playback screen; sending the joint singing video clip including a song playback screen to each second display terminal.
[0018] In one embodiment, after generating a joint singing video clip of the target song at the next playback progress by combining the singing evaluation information of the synthesized singing segment at the current singing progress and the song video clip of the target song at the next playback progress at the current playback progress, the method further includes: exporting the joint singing video clips of the target song at each playback progress in response to the user's work creation instruction; editing the exported joint singing video clips of the target song at each playback progress according to the creation instructions of the work creation instruction to obtain a joint singing song work.
[0019] In a second aspect, an embodiment of the present application provides a joint singing method, which is applied to any one of at least two second display terminals provided in a vehicle, wherein the vehicle is also provided with a first display terminal, and the second display terminal is communicatively connected to the first display terminal; the method includes:
[0020] receiving a song video clip of a target song at a current playing progress sent by the first display terminal in response to a joint singing request for at least two users in the vehicle to sing the target song together;
[0021] Displaying a video clip of the target song at the current playing progress to prompt each user to sing the target song at the current progress, so as to prompt the user to sing according to the current progress;
[0022] Receive a joint singing video clip of the target song at the next playback progress of the current playback progress sent by the first display terminal; wherein the joint singing video clip of the target song at the next playback progress of the current playback progress includes singing evaluation information of the synthesized singing clip of the current singing progress and a song video clip of the target song at the next playback progress of the current playback progress; the synthesized singing clip of the current singing progress is generated by the singing voices of each of the users for the current singing progress;
[0023] Display a joint singing video clip of the target song at the next playback progress of the current playback progress to prompt the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
[0024] In a third aspect, the present application further provides a joint singing device, which is applied to a first display terminal provided in a vehicle, wherein the vehicle is further provided with at least two second display terminals, wherein the first display terminal is communicatively connected to each of the at least two second display terminals; the device comprises:
[0025] a first segment acquisition module configured to acquire a song video segment of a target song at a current playback progress in response to a joint singing request for a target song, wherein the joint singing request is for at least two users in the vehicle to sing the target song together;
[0026] A first segment sending module is used to send a song video segment of the target song at the current playing progress to each of the at least two second display terminals for display to prompt the current singing progress of the target song;
[0027] a synthesized singing voice generation module, configured to collect singing voices of the users for the current singing progress, generate a synthesized singing voice segment of the current singing progress based on the singing voices of the users, and send the synthesized singing voice segment of the current singing progress to an audio playback device provided in the vehicle for playback;
[0028] The second segment generation module is used to obtain singing evaluation information of the synthesized singing segment at the current singing progress; the singing evaluation information is used to characterize the singing effect of each user at the current singing progress of the target song; and combine the singing evaluation information of the synthesized singing segment at the current singing progress and the song video segment of the target song at the next playback progress of the current playback progress to generate a joint singing video segment of the target song at the next playback progress;
[0029] The second clip sending module is used to send the joint singing video clip of the target song at the next playback progress to each second display terminal of the at least two second display terminals for display, so as to prompt each user of the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
[0030] In a fourth aspect, the present application further provides a joint singing device, which is applied to any one of at least two second display terminals provided in a vehicle, wherein the vehicle is further provided with a first display terminal, and the second display terminal is communicatively connected to the first display terminal; the device comprises:
[0031] a first segment receiving module configured to receive a song video segment of a target song at a current playback progress sent by the first display terminal in response to a joint singing request for at least two users in the vehicle to sing the target song together;
[0032] A first segment display module is used to display a song video segment of the target song at the current playing progress to prompt each user to sing the current progress of the target song, so as to prompt the user to sing according to the current progress;
[0033] The second segment receiving module is configured to receive a joint singing video of the target song at the next playback progress of the current playback progress, sent by the first display terminal; wherein the joint singing video segment of the target song at the next playback progress of the current playback progress includes singing evaluation information of the synthesized singing segment of the current singing progress and a song video segment of the target song at the next playback progress of the current playback progress; the synthesized singing segment of the current singing progress is generated by the singing voices of each of the users for the current singing progress;
[0034] The second clip display module is used to display the joint singing video of the target song at the next playback progress of the current playback progress, so as to prompt the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
[0035] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first and second aspects are implemented.
[0036] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first and second aspects above.
[0037] In a seventh aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the first and second aspects above.
[0038] The above-mentioned joint singing method, device, computer equipment, storage medium and computer program product first respond to the joint singing request for the target song, obtain the song video clip of the target song at the current playback progress, and send it to each second display terminal to display the song video clip to prompt each user of the current singing progress of the target song, and then collect the singing voice of each user for the current singing progress, generate a synthetic singing voice clip based on the singing voice of each user, and send it to the audio playback device for playback, and at the same time, generate singing evaluation information of the synthetic singing voice clip for characterizing the singing effect, so as to combine the singing evaluation information with the song video clip of the next playback progress. The video clip is used to generate a joint singing video clip of the next playback progress, which is sent to each second display terminal for display to prompt each user of the next singing progress of the target song. Therefore, the method of each terminal independently collecting audio and independently synthesizing audio and pictures is avoided, and the first display terminal is used for unified collection, synthesis and distribution, which improves the accuracy of audio collection, improves the efficiency of audio and picture synthesis, avoids the traditional always-on synchronization method, reduces the delay of audio and picture playback, improves the real-time performance of joint singing, overcomes the problem of picture asynchrony in the process of multi-user joint singing, and thus improves the singing effect of multi-user joint music singing. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 11 is a flow chart of a joint singing method according to an embodiment;
[0041] Figure 2 is a flow chart of a joint singing method according to another embodiment;
[0042] Figure 3 is a flowchart of a joint singing method in yet another embodiment;
[0043] Figure 4 It is a schematic diagram of the principle of the existing joint singing technology;
[0044] Figure 5 A schematic diagram of the synchronization mechanism of an existing joint singing technology;
[0045] Figure 6 Schematic diagram of the principle of a joint singing method in one embodiment;
[0046] Figure 7 A product logic flow chart of a joint singing method in one embodiment;
[0047] Figure 8 A flowchart of a technical implementation of a joint singing method in one embodiment;
[0048] Figure 9 is a structural block diagram of a joint singing device in one embodiment;
[0049] Figure 10 A structural block diagram of a joint singing device in another embodiment;
[0050] Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0052] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0053] It should be noted that the joint singing method provided in the embodiments of the present application is applied in a vehicle, which is provided with a display terminal, i.e., a multimedia terminal with display functionality, comprising a first display terminal and a second display terminal. Generally, the first display terminal is used to capture audio, synthesize the audio and image, and transmit the synthesized audio and image to the second display terminal. The second display terminal is used to display the synthesized audio and image transmitted by the first display terminal and allow multiple users to perform joint singing. More specifically, the first display terminal can be a display terminal located at the corresponding position of the driver and front passenger seat, and the second display terminal can be a display terminal located in another rear seat. Multiple users in the other rear seats can perform joint singing through multiple second display terminals. For example, users in two rear seats can perform joint singing through the display terminals located in front of them.
[0054] In one embodiment, Figure 1 As shown, a joint singing method is provided, which is applied to a first display terminal provided in a vehicle. The vehicle is also provided with at least two second display terminals, and the first display terminal is communicatively connected to each of the at least two second display terminals. In this embodiment, the method includes the following steps:
[0055] S101, in response to a joint singing request for a target song, obtaining a song video clip of the target song at the current playing progress; the joint singing request is used for at least two users in a vehicle to sing the target song together.
[0056] Among them, the target song is the song to be sung jointly. Joint singing refers to multi-user online karaoke. The users are terminal users in the vehicle who have the need for online karaoke, such as the driver and co-driver passengers and the rear passengers. The joint singing request is used for at least two users in the vehicle to sing the target song together. For example, the driver and co-driver passengers initiate a joint singing request to the screen terminal of a rear passenger through the driver and co-driver integrated screen terminal to establish a data connection to realize the function of joint singing.
[0057] Here, a song video clip refers to video clip data that is continuously transmitted and played in real time in the form of a data stream. Naturally, in this application, a song video clip refers to a video clip data of a video resource such as music or music video associated with the target song. The song video clip currently playing refers to the song video clip used for song playback at the current moment, and this video clip corresponds to the image frame at the current moment. It should be noted that the current moment can be any moment during the joint singing process.
[0058] S102: Send a song video clip of the target song at the current playing progress to each of the at least two second display terminals for display to prompt the current singing progress of the target song.
[0059] Among them, the number of second display terminals is at least two, and each second display terminal is a terminal used by at least two users for joint singing. Each second display terminal is connected to the first display terminal respectively, and is used to receive the screen related to the synthesized target song sent by the first display terminal.
[0060] Specifically, after receiving the song video clip of the target song at the current playback progress, the second display terminal can display the song playback screen of the song video clip, that is, display the song playback screen of the target song at the current playback progress. The song playback screen is used to prompt each user to sing the current progress of the target song; for example, at the current moment of the current playback progress, the song playback screen of the target song records the lyrics and singing progress at the current moment, as well as the background picture of the song playback at the current moment. This information is used to prompt each user participating in the joint singing of the current singing progress, that is, the current singing progress.
[0061] S103, collecting the singing voices of each user according to the current singing progress, generating a synthetic singing voice segment of the current singing progress according to the singing voices of each user, and sending the synthetic singing voice segment of the current singing progress to the audio playback device set in the vehicle for playback.
[0062] Each user watches the song playback screen to learn the current progress of the target song and then sings. The first display terminal uses a recording device to record the audio of the singing voices of each participating user. The synthesized singing clip is audio data obtained by synthesizing the singing voices of each user. The audio playback device is used to play the synthesized singing clip. The audio playback device includes but is not limited to a car speaker. That is, the audio playback function of the embodiment of the application is uniformly played through the first display terminal.
[0063] Specifically, the first display terminal can also collect the singing voices of each user for the current singing progress, and synthesize the singing voices to generate a synthesized singing segment of the current singing progress, and can send the synthesized singing segment of the current singing progress to the audio playback device of the vehicle, and the audio playback device sends the synthesized singing segment of the current singing progress to the audio playback device set in the vehicle, so as to play the synthesized singing segment.
[0064] S104, generating singing evaluation information of the synthesized singing segment at the current singing progress; the singing evaluation information is used to characterize the singing effect of each user at the current singing progress of the target song; combining the singing evaluation information of the synthesized singing segment at the current singing progress and the song video segment of the target song at the next playback progress of the current playback progress, generating a joint singing video segment of the target song at the next playback progress.
[0065] The synthesized singing clip is audio data generated by synthesizing the singing voices of each user. Based on different evaluation criteria, the synthesized singing clip demonstrates a singing effect. This singing effect can be characterized by singing evaluation information, which can be qualitatively described (e.g., excellent, good, fair, poor) or quantitatively scored to intuitively reflect the overall quality of the singing. This singing evaluation information can also be used to generate a joint singing video clip of the target song at the next playback progress of the current playback progress, for display on the second display terminal.
[0066] Among them, the joint singing video clip is the video data displayed for the target song in the next playback progress. The joint singing video clip includes the singing evaluation information of the user singing the target song in the current playback progress, and also includes the song video clip of the target song in the next playback progress. The joint singing video clip corresponds to the image frame at the next moment.
[0067] Specifically, after obtaining the synthesized singing clip of the current singing progress, the singing effect of the synthesized singing clip can also be evaluated to obtain singing evaluation information. The singing evaluation information and the song video clip of the next playback progress can then be combined to generate a joint singing video clip of the target song in the next playback progress. At this time, the song playback screen of the corresponding joint singing video clip can record the lyrics and singing progress of the next playback progress, the background screen of the song playback in the next playback progress, and the singing evaluation information for the current playback progress.
[0068] Taking the first frame of the target song as the current playback progress as an example, after the first display terminal sends the song video clip of the first frame of the song to each second display terminal for display, it can also collect the singing voices of each user for the first frame song video clip, and synthesize the synthetic singing clip of the first frame. Afterwards, the above-mentioned singing voices can be evaluated to obtain the singing evaluation information corresponding to the first frame of the synthetic singing clip, and then combine the singing evaluation information of the first frame and the next playback progress, that is, the song video clip of the second frame, to generate a joint singing video clip of the target song in the second frame.
[0069] S105, sending the joint singing video clip of the target song at the next playback progress to each second display terminal of at least two second display terminals for display, so as to prompt each user of the singing effect of the current singing progress of the target song and the next singing progress of the target song.
[0070] Among them, the second display terminal is also used to display the playback screen corresponding to the joint singing video clip, which can be used to prompt each user of the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song. Specifically, after the first display terminal obtains the joint singing video clip of the next playback progress, it can also send the joint singing video clip to each second display terminal for display to show the song playback screen of the target song in the next playback progress. In addition to being used to prompt each user to sing the next singing progress of the target song, the song playback screen can also be used to prompt each user of the singing effect of the target song in the current singing progress.
[0071] In the above-mentioned joint singing method, first, in response to a joint singing request for a target song, a song video clip of the target song at the current playback progress is obtained, and sent to each second display terminal to display the song video clip, so as to prompt each user of the current singing progress of the target song. Then, the singing voice of each user at the current singing progress is collected, and a synthetic singing voice clip is generated based on the singing voice of each user, and sent to the audio playback device for playback. At the same time, the synthetic singing voice clip can also be generated to represent singing evaluation information of the singing effect, so as to combine the singing evaluation information with the song video clip of the next playback progress to generate a joint singing video clip of the next playback progress, and send it to each second display terminal for display, so as to prompt each user of the next singing progress of the target song. Therefore, the method of each terminal independently collecting audio and independently synthesizing audio and pictures is avoided, and the first display terminal is used for unified collection, synthesis and distribution, which improves the accuracy of audio collection, improves the efficiency of audio and picture synthesis, avoids the traditional always-synchronized method, reduces the delay of audio and picture playback, improves the real-time performance of joint singing, overcomes the problem of picture asynchrony during the joint singing of multiple users, and thus improves the singing effect of multi-user joint music singing.
[0072] In one embodiment, obtaining a song video clip of a target song at a current playback progress includes: sending a resource download request for the target song to a server; receiving video data of the target song sent by the server according to the resource download request; and obtaining a song video clip of the target song at a current playback progress from the video data of the target song.
[0073] Among them, the server is a server connected to the first display terminal for storing video data resources, for example, a background server connected to the driver and co-driver integrated screen terminal; the resource download request refers to a request for retrieving song resources from the server.
[0074] Exemplarily, the instruction of the joint singing request carries the identifier of the target song. The first display terminal can generate a resource download request for the target song through the identifier and send it to the server. The server responds to the resource download request, retrieves the song resource associated with the identifier and sends it to the first display terminal. The first display terminal can receive the song video clip of the target song at the starting moment sent by the server as the song video clip of the current playback progress, so that the first display terminal can realize the function of joint singing.
[0075] In this embodiment, a method is provided for retrieving song resources from a server using a resource download request. Compared with the user needing to manually upload the target song resources from other terminal devices, such as a mobile terminal, this method can improve the efficiency of obtaining K song resources and save space for resource storage on the first display terminal.
[0076] In one embodiment, generating a synthesized singing segment based on the singing voice of each user includes: obtaining the accompaniment of the target song; and generating a synthesized singing segment based on the singing voice of each user and the accompaniment of the target song.
[0077] Among them, the accompaniment of the target song is the accompaniment of the song. In the process of synthesizing the synthesized singing voice, in this embodiment, in addition to synthesizing the singing voice of each user, it is also necessary to synthesize it with the accompaniment. Finally, the synthesized singing voice clip is played in the audio playback device to improve the comprehensiveness of the audio synthesis.
[0078] In one embodiment, singing evaluation information of a synthesized singing segment of a current singing progress is generated, including: identifying singing characteristics of each user singing a target song from the synthesized singing segment of the current singing progress; and generating singing evaluation information based on the singing characteristics of each user singing the target song.
[0079] Among them, the singing features are used to characterize the characteristics of each user's singing of the target song, which refers to a set of features that can reflect the uniqueness of different users when singing the target song. These features may include pitch, rhythm control, timbre characteristics, vocal skills, etc., which are used to characterize each user's singing style and ability characteristics. Therefore, in this embodiment, the singing evaluation information can be generated after judging the effect of the user's singing of the target song based on the singing features.
[0080] For example, an audio analysis algorithm is used to compare each note in the synthesized vocal segment with the standard score to calculate the pitch deviation value as the intonation feature. A beat detection algorithm is used to determine the duration and rhythm of the notes in the synthesized vocal segment, and then compare them with the standard rhythm to obtain the rhythm feature. Acoustic feature extraction techniques, such as Mel-Frequency Cepstral Coefficients (MFCC), are used to extract the timbre features of the synthesized vocal segment. Then, based on preset evaluation criteria, weights are assigned to the intonation features, rhythm features, and timbre features. Each feature type is scored separately and the weighted calculation is used to obtain an overall score. The overall score can be used directly as performance evaluation information or can be used to evaluate the performance by assigning the overall score to an evaluation level.
[0081] In this embodiment, the singing features are first identified using algorithms such as audio analysis, and then singing evaluation information is generated based on preset evaluation criteria. The singing evaluation information can be used to prompt the user of the singing effect of the target song at the current singing progress. In this way, the user's perception of the singing quality is improved, thereby improving the effect of joint singing.
[0082] In one embodiment, in response to a joint singing request for a target song, a song video clip of the target song at the current playback progress is obtained, including: in response to a joint singing request for a target song sent by any second display terminal, a song video clip of the target song at the current playback progress is obtained.
[0083] In this embodiment, the second display terminal initiates a joint singing request, and the first display terminal responds to the joint singing request to obtain a song video clip of the target song at the current playback progress, so as to achieve diversity in opening the joint singing mode, thereby improving the singing effect of the joint singing.
[0084] In one embodiment, in response to a joint singing request for a target song, a song video clip of the target song at the current playback progress is obtained, including: sending a joint singing request for the target song to each second display terminal, and receiving a joint singing confirmation message sent by any second display terminal in response to the joint singing request; in response to the joint singing confirmation message, a song video clip of the target song at the current playback progress is obtained.
[0085] The joint singing confirmation information is confirmation information for the joint singing request, and is used to establish a data connection between the first display terminal and the second display terminal to transmit a song video clip of the target song at the current playing progress.
[0086] In this embodiment, a joint singing request is first sent to the second display terminal, and then a joint singing confirmation message is received from the second display terminal in response to the request. After receiving the confirmation message, the target song is matched according to the joint singing request, and finally a song video clip of the target song at the current playback progress is obtained. This embodiment implements the preparation process for multi-terminal joint singing through the technology of sending a request, receiving a confirmation, and then obtaining video data, and realizes the activation of the joint singing mode through the first display terminal to achieve diversity in the activation of the joint singing mode, thereby improving the singing effect of the joint singing.
[0087] It can be seen that the joint singing request can be initiated by the first display terminal or the second display terminal. Specifically, the joint singing request can be initiated by the driver and passenger integrated screen terminal or the rear terminal.
[0088] In one embodiment, in response to a joint singing request for a target song, a song video clip of the target song at the current playback progress is obtained, including: in response to the joint singing request for the target song, obtaining the current working status information corresponding to the vehicle; the current working status information is used to characterize whether the driving status of the vehicle is safe and whether the equipment in the vehicle is operating normally; when the current working status information indicates that the driving status of the vehicle is safe and the equipment in the vehicle is operating normally, a song video clip of the target song at the current playback progress is obtained.
[0089] Among them, the current working status information is used to characterize whether the driving status of the vehicle is safe and whether the equipment in the vehicle is operating normally. Among them, whether the driving status of the vehicle is safe can be whether the vehicle is driving dangerously, and whether the equipment in the vehicle is operating normally can include whether the equipment is faulty, whether the equipment communication link is abnormal, etc.
[0090] For example, before initiating a joint singing request to enter the joint singing mode, it is necessary to detect whether the current status allows karaoke. If it is allowed, that is, the driving state of the vehicle is safe and the equipment in the vehicle is operating normally, the rear screen host responds to the online karaoke request initiated by the designated host and notifies the designated host to create an online karaoke service. If it is not allowed, the user will be prompted.
[0091] In this embodiment, the current working status information corresponding to the vehicle is obtained by identifying the joint singing request, and the driving safety and equipment stability are identified based on the current working status information, thereby ensuring the safety of vehicle driving during the joint singing process.
[0092] In one embodiment, a joint singing video clip of the target song at the next playback progress is sent to each second display terminal of at least two second display terminals for display, including: rendering the joint singing video clip of the target song at the next playback progress to obtain a joint singing video clip including a song playback screen; and sending the joint singing video clip including a song playback screen to each second display terminal.
[0093] In this embodiment, off-screen rendering technology is used, and drawing is not performed directly on the second display terminal. Instead, graphics rendering is performed in an off-screen buffer in the memory of the first display terminal. The song playback screen in the target video data is processed in advance in the background, and a rendered joint singing video clip is generated, thereby reducing the rendering pressure of the second display terminal, improving rendering efficiency and quality, ensuring that each second display terminal can receive and display video data smoothly and with high quality, and improving the stability of multi-user joint singing.
[0094] In one embodiment, after generating a joint singing video clip of the target song at the next playback progress by combining the singing evaluation information of the synthesized singing clip at the current singing progress and the song video clip of the target song at the next playback progress of the current playback progress, the method also includes: exporting the joint singing video clips of the target song at each playback progress in response to the user's work creation instruction; editing the exported joint singing video clips of the target song at each playback progress according to the creation instructions of the work creation instruction to obtain a joint singing song work.
[0095] Among them, the joint singing song work is a work obtained by editing the joint singing video clips in accordance with the creation instructions of the work creation instructions.
[0096] In this embodiment, since the first display terminal is set as the terminal for audio and picture synthesis, compared with the traditional technology in which each terminal synthesizes separately, the synthesized audio and video resources can be exported more conveniently, thereby generating user original works based on the exported joint singing video clips, improving the efficiency of resource export and improving the interactive effect in the joint singing scenario.
[0097] In another embodiment, Figure 2 As shown, a joint singing method is provided, comprising the following steps:
[0098] S201, in response to a joint singing request for a target song, sending a resource download request for the target song to a server;
[0099] S202, receiving the video data of the target song sent by the server according to the resource download request, and obtaining a song video clip of the target song at the current playback progress from the video data;
[0100] S203, sending the song video clip at the current playing progress to each second display terminal of at least two second display terminals for display;
[0101] S204, collecting the singing voices and the accompaniment of the target song emitted by each user according to the current singing progress;
[0102] S205, generating a synthesized singing segment based on the singing voices of each user and the accompaniment of the target song, and sending the synthesized singing segment to an audio playback device provided in the vehicle for playback;
[0103] S206, identifying singing characteristics of each user singing the target song from the synthesized singing segments, and generating singing evaluation information corresponding to the synthesized singing segments;
[0104] S207, combining the singing evaluation information of the current singing progress and the song video clip of the target song at the next playing progress of the current playing progress, generating a joint singing video clip of the next playing progress;
[0105] S208: Send the joint singing video clip to each of the at least two second display terminals for display.
[0106] It should be noted that the specific limitations of the above steps can be found in the specific limitations of a joint singing method above, which will not be repeated here.
[0107] In one embodiment, Figure 3 As shown, a joint singing method is provided, which is applied to any one of at least two second display terminals provided in a vehicle, wherein a first display terminal is also provided in the vehicle, and the second display terminal is communicatively connected to the first display terminal. In this embodiment, the method includes the following steps:
[0108] S301, receiving a song video clip of a target song at a current playing progress sent by a first display terminal in response to a joint singing request; the joint singing request is for at least two users in a vehicle to sing the target song together.
[0109] The song video clip is the video data corresponding to the target song that matches the joint singing request. The joint singing request is for at least two users in the vehicle to sing the target song together. The target song is the song to be jointly sung. Joint singing refers to multi-user online karaoke, and users are terminal users in the vehicle who have the online karaoke request.
[0110] Among them, the song video clip with the current playback progress refers to the data in the video data used for song playback at the current moment, and the data corresponds to the image frame at the current moment. The current moment can be any moment in the joint singing process.
[0111] S302, showing a song video clip of the target song at the current playing progress to prompt each user to sing the current singing progress of the target song, so as to prompt the user to sing according to the current singing progress.
[0112] Among them, the playback screen of the song video clip can be used to prompt each user of the current singing progress of the target song. At the current moment, the song video clip of the target song records the lyrics and singing progress at the current moment, as well as the background picture of the song playing at the current moment. This information is used to prompt each user participating in the joint singing of the current singing progress at the current moment.
[0113] S303, receiving the joint singing video clip of the target song at the next playback progress of the current playback progress sent by the first display terminal; wherein, the joint singing video clip of the target song at the next playback progress of the current playback progress includes the singing evaluation information of the synthesized singing clip of the current singing progress and the song video clip of the target song at the next playback progress of the current playback progress; the synthesized singing clip of the current singing progress is generated by the singing voices of each user according to the current singing progress.
[0114] Among them, the joint singing video clip may include the singing evaluation information of the synthesized singing clip of the current singing progress, and the song video clip of the target song in the next playback progress. The singing evaluation information is generated according to the singing effect of the synthesized singing clip. The synthesized singing clip is synthesized according to the singing voice of each user. The singing voice is the singing voice of each user according to the singing progress obtained by the first display terminal.
[0115] S304, showing the joint singing video clip of the target song at the next playing progress of the current playing progress, to prompt the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
[0116] It should be noted that the specific limitations of the above steps can be found in the specific limitations of the first display terminal joint singing method applied in the vehicle above, which will not be repeated here.
[0117] In the existing technology, independent playback is achieved through multiple hosts. To achieve multi-screen linked playback, it is necessary to select the playback clock of a host as a reference, and other hosts follow the playback time of this clock. At the same time, the connection protocol is used to realize the communication interface of each host to achieve karaoke score synchronization, midi synchronization, sound effect synchronization, voice synchronization and other operation synchronization.
[0118] like Figure 4As shown, in the prior art, speakers 1, 4, 2, and 3 are provided in the front, rear, left, and right directions of the vehicle compartment, respectively. An integrated screen is provided for the driver and co-driver seats in the vehicle compartment, and screens are provided in front of the four rear seats, including rear screen 1, rear screen 2, rear screen 3, and rear screen 4. In the application scenario of "multi-screen karaoke", clock synchronization is required between the integrated screen for the driver and co-driver and each rear screen, as well as synchronization of various functions in the UI screen and synchronization of various functions in audio processing. Among them, the various functions in the UI screen include screen synthesis, scoring effects, MIDI display, lyrics display, special effects display, and off-screen video playback functions; the various functions in audio processing include audio synthesis, sound effects and scoring, as well as audio acquisition functions. Among them, MIDI refers to the rhythm data of lyrics, including pitch, tone, beat time, etc., which is used for scoring and recording alignment in karaoke projects.
[0119] like Figure 5 As shown, the technical implementation process of the existing solution includes the following steps: (1) starting all hosts and screens; (2) starting the Karaoke App on the integrated screen of the driver and passenger, and the user selects to start the "multi-screen linkage Karaoke" function mode and selects the song to be sung, so as to send a linkage Karaoke request to the rear host screen to start the "multi-screen linkage Karaoke" function mode for the song to be sung; (3) in response to the linkage Karaoke request of the integrated screen of the driver and passenger, the rear user confirms to turn on the multi-screen linkage Karaoke function through the rear host screen, and automatically establishes a communication link with the integrated screen of the driver and passenger; (4) The integrated screen of the driver and passenger starts to download resources and start playing, and at the same time, through the established communication link, "default" Notify the rear host to start downloading song resources and start playing; (5) The rear hosts obtain the currently playing song information, playlist, MIDI, lyrics information and other resources through the communication link, and start playing and perform audio collection, audio scoring, audio playback, and UI screen display; (6) The rear hosts poll the playback progress of the main and co-pilot integrated screens to achieve their own audio and video synchronization; Among them, online karaoke is that all hosts play a song at the same time for karaoke. In order to synchronize, the main and co-pilot integrated screens will be selected as the reference host, and other devices will refer to the song playback progress on the main and co-pilot integrated screens to play songs. A certain delay compensation will be given to the playback progress based on data transmission delay and system delay. (7) Later, it is necessary to synthesize the multiple audios collected by each rear host and the main and co-pilot integrated screen host to obtain the final karaoke work. Among them, the existing technology does not perform synthesis on the main and co-pilot integrated screens. The main and co-pilot integrated screens are only responsible for synchronizing the playback progress, playback resource information, playback status, etc., and the synthesis is performed independently by each host.
[0120] From the above steps, it can be seen that in the multi-screen linkage karaoke solution in the existing technology, the audio recorded by each host needs to be synchronized to all other hosts, and the sending and receiving processes are relatively time-consuming; audio synthesis requires a large amount of CPU resources, resulting in poor alignment of multiple synthesized voices; the later synthesis of user works is more difficult, and it is easy to have problems such as the voices of multiple people cannot be aligned, and the accompaniment and vocals are out of sync.
[0121] In addition, in terms of synchronization, such as Figure 5 As shown, in the prior art, the driver and passenger seats in the vehicle are provided with integrated driver and passenger screens, and the rear row is provided with screens, including rear screen 1, rear screen 2, rear screen 3, and rear screen 4. In the application scenario of "multi-screen linked karaoke", each rear screen separately polls the driver and passenger integrated screen host to obtain media clock information and karaoke information. Karaoke information includes playback progress, buffering status, song list, lyrics, MIDI, etc. A timed synchronization detector is set to determine whether the synchronization requirements are met based on the information obtained. If so, a detection cycle is set to continue the next round of detection; if the synchronization requirements are not met, the synchronization mechanism of each rear screen host is triggered to achieve alignment of the media clock information and karaoke information to complete the synchronization.
[0122] As can be seen, existing technologies use each host to independently play images, but select one host as a reference host to serve as the media clock (MediaClock). Generally, the front-end host is responsible for providing this clock, and other hosts actively poll the playback time, continuously correcting the playback time, lyrics, MIDI, etc. However, existing technologies also have many problems, the main problems are as follows:
[0123] (1) Each host needs to download audio and video resources, lyrics resources, and midi resources separately.
[0124] (2) There are differences in the karaoke scores, animations, and dry sounds (pure human voices without music) on each host. Specifically, the score display is not synchronized or delayed, the scoring animation is not synchronized or delayed, and the sound is not synchronized or delayed.
[0125] (3) The synchronization of clock, playlist, loading progress, pause / play, etc. of the rear host needs to be achieved through complex protocols.
[0126] (4) All hosts collect audio independently, and a chorus effect cannot be created. Specifically, the existing solution requires all hosts to collect audio separately. When singing a chorus song, each host needs to synthesize the audio separately. At this time, a single host needs to wait for multiple hosts to synchronize data. Due to factors such as performance and network, problems such as misaligned voices are prone to occur.
[0127] (5) Due to network and performance issues, audio and video synchronization issues may occur between different hosts. Specifically, each host connects to the network independently, downloads resources independently, and plays independently. However, due to network quality, device computing performance, buffering, and other reasons, the images and scores of each host may be inconsistent.
[0128] (6) Each host needs to independently perform scoring and audio acquisition, and cannot synthesize unified audio resources in real time. Specifically, network delays and instability may occur during the start and playback process, and buffering and synchronization data lag may occur. For example, during karaoke, if one device experiences buffering or lag, while the other devices can play normally, the device cannot record or score, and the screen may become stuck. Or if there is a serious network delay, if the sound of a certain device lags, the synthesized audio and lyrics and MIDI will not be aligned.
[0129] (7) It takes a long time for the back-row host to re-enter the karaoke after exiting it; it not only involves re-establishing the connection, but also the playback needs to be restarted. After the playback is restarted, the picture needs to "catch up" with other hosts. At the same time, the sound, scoring, and picture effects need to be synchronized from other devices, and buffering may also occur.
[0130] (8) The entire mechanism implementation process is relatively complex and requires a complex communication interface.
[0131] (9) It cannot support multiplayer PK in a friendly way. Specifically, the effect of multiplayer PK is not good, which is reflected in high latency, asynchrony between audio and video, and asynchrony between recorded chorus audio.
[0132] (10) Existing technologies still have various delay issues. For example, UI changes such as playlist operations (deletion, pinning, adding) on one host cannot be perceived by other users in real time. At the same time, the player's buffering status and seek (jump) cannot be quickly and accurately synchronized to other hosts. Therefore, complex protocols are required for interactive communication, but the maintenance cost is also very high.
[0133] In summary, the existing in-car karaoke method is based on multiple terminals that can independently realize the karaoke function. In the application scenario of "multi-screen linked karaoke", the existing technology has synchronization problems of audio real-time and picture consistency, karaoke resource download efficiency problems and delay problems.
[0134] Based on this, the embodiment of the present application provides a joint singing method, also known as a method for realizing multi-screen linkage karaoke on a car machine. Figures 6 to 8, the joint singing method is described in detail with a specific embodiment. It is worth noting that the following description is only an exemplary description and not a specific limitation of the application.
[0135] like Figure 6 As shown, an integrated screen is set at the driver and co-driver positions in the vehicle's cabin, and a screen terminal is set in the back row, including rear screen 1, rear screen 2, rear screen 3, and rear screen 4. The audio collection, audio processing and UI screen synthesis are completed through the integrated screen of the driver and co-driver, and then distributed to each rear screen, which solves the problem of multi-screen car computer. On the premise of allowing the front and rear screens to freely realize independent karaoke, multi-screen linkage karaoke can be realized, and karaoke videos, special effects, lyrics, midi, barrage, etc. can be freely synchronized to each screen. At the same time, the microphone of each screen can be used independently and online.
[0136] like Figure 7 As shown, on the product side, the joint singing method provided in the embodiment of the present application includes the following steps:
[0137] (1) Any rear passenger initiates an online karaoke request to the designated host (such as the integrated driver and passenger screen host, or any designated host that supports online karaoke service and is in working state) through the rear screen host in front of him or her. Among them, the car computer, such as the integrated driver and passenger screen, can usually keep the front and back end services working all the time, so the integrated driver and passenger screen host is often selected as the designated host. Due to the limitations of the actual environment, it is generally recommended to select the integrated driver and passenger screen to start the online karaoke service.
[0138] (2) Detect whether karaoke is allowed in the current state. If it is allowed, the rear screen host responds to the online karaoke request initiated by the designated host and notifies the designated host (such as the driver and co-driver integrated screen host) to create an online karaoke service. If it is not allowed, the user will be prompted. Situations where it is not allowed include dangerous driving, equipment failure, abnormal equipment communication link, etc.
[0139] (3) After the service is created, start the karaoke program (audio collector, scorer, lyrics, midi, MV player), etc.
[0140] (4) The karaoke program performs UI screen synthesis and audio synthesis, and synthesizes UI interface control buttons, such as pause button, volume adjustment button, etc., with song videos.
[0141] (5) Distribute the synthesized image to other screen hosts that have initiated online requests.
[0142] (6) If it is free karaoke (stand-alone karaoke), the audio collected by a single screen is directly output to the speaker. If it is online karaoke, the audio needs to be forwarded to the integrated driver and co-driver screens for synthesis, and then distributed to the speakers.
[0143] It can be seen that the advantages of the joint singing method provided in the embodiment of the present application are:
[0144] (1) The picture consistency is high, and there will be no audio and video synchronization problems; the solution to the synchronization problem is based on multiple technologies. The first choice is to implement a low-latency channel based on TinyAlas or hardware microphone on the Android platform, and then to achieve picture synthesis based on off-screen rendering. The communication link is based on WifiDisplay, HDMI Display, and MediaRouter to achieve picture and audio synchronization, and TRTC is responsible for command transmission.
[0145] (2) Only one host is needed to download MV, midi, lyrics, and playlist resources (the existing technology requires multiple hosts to download MV resources, original accompaniment, score files, lyrics, etc.).
[0146] (3) Any screen can quickly access the karaoke process. Multi-channel audio is received on the integrated driver and passenger screens, and aligned, synthesized, and distributed using technologies such as WifiDisplay, HDMI Display, and MediaRouter. Other hosts only need to play the images and sounds without downloading resources or synchronization status.
[0147] (4) Multiple voices can be synthesized in real time. The main drawback of existing technologies is that the audio recorded by each host needs to be sent to other hosts, and then each host synthesizes it independently. However, the technical solution of the embodiment of the present application uses the TinyAlas mechanism or microphone hardware driver mechanism to directly synthesize multiple channels of sound at the Hal layer, which is then sent to the app for synthesis with the accompaniment and finally output to each device. This avoids the need for waiting and synchronization when synthesizing multiple voices.
[0148] (5) Facilitate the generation of UGC (User Generated Content). In this application, the startup background of the integrated driver and passenger screens is linked to a karaoke service. The microphone records the voices of multiple people, which are then synthesized using TinyAlsa or a microphone driver. The synthesized data is fed to the karaoke program, which then synthesizes the obtained voices with the accompaniment. The delay time between the recorded voices, accompaniment, and MV is subtracted during the synthesis.
[0149] (6) The back-row host can access the karaoke more quickly after exiting it. Whether exiting or not, it can achieve quick access without downloading resources, status synchronization, etc.
[0150] (7) The logic of scoring and adding sound effects is unified.
[0151] (8) Support user clicks and user screen event sending;
[0152] Screen events including barrage, animation, clicks from other hosts, and touch events also need to be synchronized to the integrated driver and co-driver screens to implement UI changes, such as calling out the song request menu.
[0153] (9) Support user PK and duet.
[0154] Therefore, the joint singing method provided in the embodiment of the present application simplifies the complex scenarios of the existing solutions. Through this solution, while ensuring that all audio is uniformly synthesized by one host, uniformly scored by the same host, supports multi-person PK, and consistent UI rendering, the screens of other seats can be freely connected to the KTV scene without having to re-download audio and video resources.
[0155] like Figure 8 As shown, Figure 8 The following demonstrates the process of any user initiating a karaoke connection request using the current screen host (e.g., screen 1, screen 2, screen 3, or screen 4) as the interactive host, and the process of the main and co-pilot integrated screens processing the karaoke connection. Technically, the joint singing method provided in this embodiment includes the following steps:
[0156] (1) The user initiates karaoke through the interactive host (such as screen 1, screen 2, screen 3 or screen 4) and chooses whether to enable the online karaoke function according to his or her own wishes.
[0157] (2) The interactive host determines whether the user has input an instruction requesting online karaoke; if so, the karaoke mode is online karaoke; if not, the karaoke mode is stand-alone karaoke.
[0158] (3) If the karaoke mode is stand-alone karaoke, the interactive host uses the current screen to sing karaoke independently.
[0159] (4) If the karaoke mode is online karaoke, the interactive host sends a "karaoke application-multi-screen karaoke request" to the driver and co-driver integrated screen. After receiving the request, the driver and co-driver integrated screen determines "whether multi-screen karaoke can be started", that is, whether the current conditions allow karaoke.
[0160] (5) If the vehicle meets any abnormal conditions (such as dangerous driving, network abnormality, equipment abnormality), karaoke will not be allowed, and the user will be directly prompted that "the current situation does not allow it".
[0161] (6) If karaoke is allowed, the background service will be started and the karaoke environment will be initialized, that is, the main and co-pilot integrated screen host will be notified to start the karaoke service; the main and co-pilot integrated screen host will start the karaoke background program and turn on the "main and co-pilot integrated screen background karaoke service"; the various karaoke tasks (including "audio synthesis, scoring, audio addition, UI synthesis") will be executed through the service instructions configured in the "main and co-pilot integrated screen background karaoke service"; the execution of various karaoke tasks will generate synthesized audio and images in real time; it will facilitate the main and co-pilot integrated screen host to transmit the synthesized audio to the speaker and transmit the synthesized image to the main and co-pilot integrated screen and the screens of multiple interactive hosts.
[0162] (7) In terms of audio, the karaoke mode is first determined, that is, whether to transmit the audio to the integrated screen of the driver and the co-driver. When the karaoke mode is stand-alone karaoke, the audio is directly output to the speaker. When the karaoke mode is online karaoke, the recorded audio is directly sent to the integrated screen of the driver and the co-driver, synthesized and sound effects are added, and then output to the speaker.
[0163] (8) In terms of the picture, when singing karaoke online, the background service in the integrated screen of the driver and co-driver is responsible for scoring, audio synthesis, UI synthesis, MIDI, lyrics display, etc.; after the initial synthesis of the picture, the picture is rendered off-screen, the picture is synthesized into separate picture frames, and the picture frames are distributed to the integrated screen of the driver and co-driver and the screens of multiple interactive hosts, and obviously to the screens in front of each user.
[0164] The existing karaoke process does not use off-screen rendering or does not require off-screen rendering. Instead, each host independently outputs the screen, but in this process, it takes into account factors such as the playback progress and song status of the main and secondary integrated screens. In contrast, the embodiment of the application uses the main and secondary integrated screens to synthesize the karaoke app screen, using the LayerStack content managed by SurfaceFlinger in the Android system to synthesize into memory. The synthesized data is encoded using MediaCodec and rendered to other devices via the network or using EGL textures.
[0165] It can be seen that on the technical side, the joint singing method provided in the embodiment of the present application utilizes technologies such as opengles (OpenGraphics Library for Embedded Systems, an API for graphics rendering of embedded systems), EGL (interface standard, used to connect OpenGL ES with the window system of the underlying native platform), WifiP2p (a wireless local area network protocol that allows devices to be directly connected without a relay router), TinyALAS (Advanced Linux Sound Architecture, a lightweight audio library that provides a simplified interface to the audio framework), UDP (User Datagram Protocol), etc. to quickly implement a multi-party joint singing solution, avoiding network fluctuations, and supporting multi-screen PK, multi-screen interaction, multi-screen sound synthesis, etc., perfectly solving the network problems, audio and video synchronization problems, and multi-person karaoke problems encountered in in-vehicle multi-screen karaoke.
[0166] In summary, the joint singing method provided by the embodiment of the present application includes real-time recording and real-time playback in terms of audio real-time. During recording, a single host is used to record multiple audio channels, avoiding the problem of multi-channel audio synthesis on multiple machines and the problem of time-consuming transmission. In terms of picture consistency, the HDMI Display and WifiDisplay technologies of the Android host are used to achieve the real-time performance of the pictures of multiple hosts. In terms of karaoke resource downloading, karaoke resources are all downloaded by the same host, avoiding the problem of multiple hosts having to download accompaniment and midi before synchronizing playback when karaoke is performed online. The joint singing method provided by the embodiment of the present application has simple logic, convenient development, operation and maintenance, high degree of freedom, and low implementation cost. It solves the problems of audio real-time, picture consistency, and low efficiency of karaoke resource downloading that cannot be handled in existing scenarios. The program exception rate after access and use is low, and the online karaoke experience is better.
[0167] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0168] Based on the same inventive concept, embodiments of the present application also provide a joint singing device for implementing the aforementioned joint singing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more joint singing device embodiments provided below can be found in the aforementioned limitations of the joint singing method and will not be further elaborated here.
[0169] In one embodiment, Figure 9 As shown, a joint singing device is provided, which is applied to a first display terminal provided in a vehicle. The vehicle is also provided with at least two second display terminals. The first display terminal is respectively connected to each of the at least two second display terminals. The device includes: a first segment acquisition module 901, a first segment sending module 902, a synthesized singing voice generation module 903, a second segment generation module 904, and a second segment sending module 905, wherein:
[0170] The first segment acquisition module 901 is configured to obtain a video segment of the target song at the current playback progress in response to a joint singing request for the target song; the joint singing request is for at least two users in a vehicle to sing the target song together;
[0171] A first segment sending module 902 is configured to send a song video segment of the target song at the current playing progress to each of the at least two second display terminals for display to indicate the current playing progress of the target song;
[0172] The synthesized singing voice generation module 903 is used to collect the singing voices of each user according to the current singing progress, generate a synthesized singing voice segment of the current singing progress based on the singing voices of each user, and send the synthesized singing voice segment of the current singing progress to the audio playback device set in the vehicle for playback;
[0173] The second segment generation module 904 is configured to obtain singing evaluation information for the synthesized singing segment at the current singing progress; the singing evaluation information is used to characterize the singing effect of each user at the current singing progress of the target song; and to generate a combined singing video segment of the target song at the next playing progress by combining the singing evaluation information of the synthesized singing segment at the current singing progress and the song video segment at the next playing progress of the target song at the current playing progress.
[0174] The second clip sending module 905 is used to send the joint singing video clip of the target song in the next playback progress to each second display terminal of at least two second display terminals for display, so as to prompt each user of the singing effect of the current singing progress of the target song and the next singing progress of the target song.
[0175] In one embodiment, Figure 10As shown, a joint singing device is provided, which is applied to any one of at least two second display terminals provided in a vehicle. The vehicle is also provided with a first display terminal, and the second display terminal is communicatively connected to the first display terminal. The device includes: a first segment receiving module 1001, a first segment display module 1002, a second segment receiving module 1003, and a second segment display module 1004, wherein:
[0176] The first segment receiving module 1001 is configured to receive a video segment of a target song at a current playback progress sent by the first display terminal in response to a joint singing request; the joint singing request is for at least two users in a vehicle to sing the target song together;
[0177] The first segment display module 1002 is used to display a song video segment of the target song at the current playing progress to prompt each user to sing the current singing progress of the target song, so as to prompt the user to sing according to the current singing progress;
[0178] The second segment receiving module 1003 is configured to receive a joint singing video of a target song at the next playback progress of the current playback progress sent by the first display terminal; wherein the joint singing video segment of the target song at the next playback progress of the current playback progress includes singing evaluation information of the synthesized singing segment of the current singing progress and a song video segment of the target song at the next playback progress of the current playback progress; the synthesized singing segment of the current singing progress is generated by the singing voices of each user according to the current singing progress;
[0179] The second segment display module 1004 is used to display the joint singing video of the target song at the next playback progress of the current playback progress, so as to prompt the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
[0180] Each module in the aforementioned joint singing device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0181] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 11As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a joint singing method. The display unit of the computer device is used to produce a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0182] Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0183] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0184] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0185] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0186] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0187] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0188] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A joint singing method, characterized in that: The method is applied to a first display terminal provided in a vehicle, wherein the vehicle is further provided with at least two second display terminals, wherein the first display terminal is communicatively connected to each of the at least two second display terminals; the method comprises: In response to a joint singing request for a target song, obtaining a song video clip of the target song at a current playback progress; the joint singing request is for at least two users in the vehicle to sing the target song together; Sending a song video clip of the target song at the current playing progress to each of the at least two second display terminals for display to prompt the current singing progress of the target song; Collecting singing voices of each user for the current singing progress, generating a synthesized singing voice segment of the current singing progress based on the singing voices of each user, and sending the synthesized singing voice segment of the current singing progress to an audio playback device provided in the vehicle for playback; Generate singing evaluation information of the synthesized singing segment at the current singing progress; the singing evaluation information is used to characterize the singing effect of each user at the current singing progress of the target song; combine the singing evaluation information of the synthesized singing segment at the current singing progress and the song video segment of the target song at the next playback progress of the current playback progress, and generate a joint singing video segment of the target song at the next playback progress; The joint singing video clip of the target song at the next playback progress is sent to each second display terminal of the at least two second display terminals for display, so as to prompt each user of the singing effect of the current singing progress of the target song and the next singing progress of the target song.
2. The joint singing method according to claim 1, characterized in that: The step of obtaining a video clip of the target song at the current playback progress includes: Sending a resource download request for the target song to a server; receiving video data of the target song sent by the server according to the resource download request; A song video clip of the target song at the current playing progress is obtained from the video data of the target song.
3. The joint singing method according to claim 1, characterized in that: The step of generating singing evaluation information of the synthesized singing segment of the current singing progress includes: Identifying singing characteristics of each user singing the target song from the synthesized singing voice segment of the current singing progress; The singing evaluation information is generated based on the singing characteristics of each user singing the target song.
4. The joint singing method according to claim 1, characterized in that: The step of obtaining a song video clip of the target song at a current playing progress in response to a joint singing request for the target song includes: In response to a joint singing request for the target song sent by any of the second display terminals, a song video clip of the target song at the current playing progress is obtained.
5. The joint singing method according to claim 1, characterized in that: The step of obtaining a song video clip of the target song at a current playing progress in response to a joint singing request for the target song includes: sending a joint singing request for the target song to each of the second display terminals, and receiving a joint singing confirmation message sent by any of the second display terminals in response to the joint singing request; In response to the joint singing confirmation information, a song video clip of the target song at the current playing progress is obtained.
6. The joint singing method according to claim 1, characterized in that: The step of obtaining a song video clip of the target song at a current playing progress in response to a joint singing request for the target song includes: In response to a joint singing request for the target song, obtaining current operating status information corresponding to the vehicle; the current operating status information is used to indicate whether the driving state of the vehicle is safe and whether the equipment in the vehicle is operating normally; When the current working status information indicates that the driving status of the vehicle is safe and the equipment in the vehicle is operating normally, a song video clip of the target song at the current playing progress is obtained.
7. The joint singing method according to claim 1, characterized in that: The sending the joint singing video clip of the target song at the next playback progress to each of the at least two second display terminals for display includes: Rendering the joint singing video clip of the target song at the next playback progress to obtain a joint singing video clip including a song playback screen; The joint singing video clip including the song playing screen is sent to each of the second display terminals.
8. The joint singing method according to claim 1, characterized in that: After generating a joint singing video clip of the target song at the next playback progress by combining the singing evaluation information of the synthesized singing segment at the current playback progress and the song video clip of the target song at the next playback progress of the current playback progress, the method further includes: In response to the user's work creation instruction, the joint singing video clips of the target song at each playback progress are exported; and the exported joint singing video clips of the target song at each playback progress are edited according to the creation instructions of the work creation instruction to obtain a joint singing song work.
9. A joint singing method, characterized in that: The method is applied to any one of at least two second display terminals provided in a vehicle, wherein the vehicle is further provided with a first display terminal, and the second display terminal is communicatively connected to the first display terminal; the method comprises: receiving a song video clip of a target song at a current playing progress sent by the first display terminal in response to a joint singing request for at least two users in the vehicle to sing the target song together; Displaying a video clip of the target song at the current playing progress to prompt each user to sing the target song at the current progress, so as to prompt the user to sing according to the current progress; Receive a joint singing video clip of the target song at the next playback progress of the current playback progress sent by the first display terminal; wherein the joint singing video clip of the target song at the next playback progress of the current playback progress includes singing evaluation information of the synthesized singing clip of the current singing progress and a song video clip of the target song at the next playback progress of the current playback progress; the synthesized singing clip of the current singing progress is generated by the singing voices of each of the users for the current singing progress; Display a joint singing video clip of the target song at the next playback progress of the current playback progress to prompt the singing effect of the current singing progress of the target song and prompt the next singing progress of the target song.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.