Vehicle-mounted karaoke interaction method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610704143.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-09-08
AI Technical Summary
[0004]然而相关技术中提供的车载K歌方法,其呈现方式单一,缺乏娱乐性
[0010]本申请提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122715631A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of application software, and in particular to an in-vehicle karaoke interactive method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of intelligent vehicles, the intelligent cockpit has become a third living space integrating travel, work, and entertainment. Among them, the rise of the mobile karaoke room concept allows users to enjoy karaoke while driving.
[0003] In related technologies, in-car karaoke can be achieved through an in-vehicle terminal and a car microphone. The in-vehicle terminal displays a karaoke interface, allowing users to select songs and sing. In this scenario, users do not need to hold a physical microphone; they can record their performance directly using the built-in car microphone in the cabin.
[0004] However, the in-car karaoke methods provided by related technologies are monotonous and lack entertainment value. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for in-vehicle karaoke interaction, which can enrich the user experience of in-vehicle karaoke. The technical solution is as follows: According to one aspect of this application, an in-vehicle karaoke interactive method is provided, the method being executed by an in-vehicle terminal, the method comprising: Display the karaoke interface for the first song and play the accompaniment for the first song; Ambient audio, including human voices and accompaniment, is recorded using a vehicle-mounted microphone. If the human voice in the ambient audio matches the original audio of the first song, interactive effects are displayed on the karaoke interface.
[0006] According to another aspect of this application, an in-vehicle karaoke interactive device is provided, the device being used to implement an in-vehicle terminal, the device comprising: The display module is used to display the karaoke interface of the first song and play the accompaniment of the first song; A recording module for recording ambient audio via an in-vehicle microphone, the ambient audio including human voices and accompaniment; The display module is used to display interactive effects on the karaoke interface when the human voice in the ambient audio matches the original audio of the first song.
[0007] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the in-vehicle karaoke interactive method as described above.
[0008] According to another aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the in-vehicle karaoke interactive method as described above.
[0009] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the in-vehicle karaoke interactive method provided in various alternative implementations of the above aspects.
[0010] The beneficial effects of the technical solution provided in this application include at least the following: By introducing real-time pitch evaluation and highlight interaction mechanisms, this technology addresses the issues of limited presentation and lack of entertainment in in-car karaoke systems. Through real-time analysis of the matching degree between the user's voice and the original singer, it automatically triggers various interactive effects such as thumbs-up animations and cheering effects when a high matching degree is detected. This enhances the user's immersive experience while singing karaoke in the cabin, provides emotional value, compensates for the lack of entertainment in existing technologies, and improves user participation and enjoyment in in-car karaoke. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the structure of a computer system provided in an exemplary embodiment of this application; Figure 2 This is a flowchart of an exemplary embodiment of the in-vehicle karaoke interactive method provided in this application; Figure 3This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 4 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 5 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 6 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 7 This is a flowchart of an exemplary embodiment of the in-vehicle karaoke interactive method provided in this application; Figure 8 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 9 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 10 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 11 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 12 This is a schematic diagram of an in-vehicle karaoke interactive method provided in an exemplary embodiment of this application; Figure 13 This is a schematic diagram of the structure of an in-vehicle karaoke interactive device provided in an exemplary embodiment of this application; Figure 14 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application.
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] Figure 1 A schematic diagram of a computer system provided in an exemplary embodiment of this application is shown. The computer system may include an in-vehicle terminal 101 and a server 103.
[0016] For example, the in-vehicle karaoke interaction method shown in the embodiments of this application can be applied to an in-vehicle terminal, wherein the in-vehicle terminal 101 runs an application 102 that supports in-vehicle karaoke interaction.
[0017] For example, the in-vehicle karaoke interaction method provided in this application can be executed by a client on an in-vehicle terminal. This client is a client of an application that supports in-vehicle karaoke interaction. Any application with in-vehicle karaoke functionality can apply the in-vehicle karaoke interaction method provided in this application embodiment; therefore, this application is not limited to any type of application.
[0018] The vehicle-mounted terminal 101 includes a first memory and a first processor. The first memory stores a vehicle-mounted karaoke interactive program; the vehicle-mounted karaoke interactive program is invoked and executed by the first processor to implement the vehicle-mounted karaoke interactive method provided in this application. The first memory may include, but is not limited to, the following: Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM).
[0019] The first processor can consist of one or more integrated circuit chips. Optionally, the first processor can be a general-purpose processor, such as a central processing unit (CPU) or a network processor (NP). Optionally, the first processor can implement the in-vehicle karaoke interactive method provided in this application by running programs or code.
[0020] In one alternative embodiment, the vehicle terminal 101 and the server 103 can be interconnected via a wireless network.
[0021] Server 103 is used to provide background services for the client of vehicle terminal 101. Optionally, server 103 undertakes the main computing work and vehicle terminal 101 undertakes the secondary computing work; or, server 103 undertakes the secondary computing work and vehicle terminal 101 undertakes the main computing work; or, server 103 and vehicle terminal 101 use a distributed computing architecture for collaborative computing.
[0022] Server 103 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center.
[0023] Optionally, server 103 includes a second memory and a second processor. The second memory stores an in-vehicle karaoke interactive program; the in-vehicle karaoke interactive program is called by the second processor to implement the in-vehicle karaoke interactive method provided in this application. Optionally, the second memory may include, but is not limited to, the following: RAM, ROM, PROM, EPROM, EEPROM. Optionally, the second processor may be a general-purpose processor, such as a CPU or NP.
[0024] Figure 2 This is a flowchart illustrating an exemplary embodiment of an in-vehicle karaoke interactive method provided in this application. This method can be used for, for example... Figure 1 The vehicle-mounted terminal shown. The method includes the following steps.
[0025] Step 210: Display the karaoke interface for the first song and play the accompaniment for the first song.
[0026] For example, the method is executed by an in-vehicle terminal, which may be equipped with at least one display screen, on which the karaoke interface of the first song can be displayed. For instance, the karaoke interface of the first song can be displayed on the central control display screen, or on the passenger seat display screen.
[0027] The karaoke interface is the user interface displayed on the screen after the user activates the karaoke function on the in-vehicle terminal. The karaoke interface can display at least one of the following: song catalog, search box, recommendations and charts, lyrics of the currently played song, song progress bar, artist information, accompaniment and volume control, countdown timer, scoring system, recording and saving functions, etc.
[0028] The in-vehicle terminal may have multiple displays, and the karaoke interface displayed on different displays may vary. For example, on the karaoke initiator (such as a passenger in the front passenger seat using the front passenger display or the center console display, or other displays such as the rear entertainment screen), the interface will include complete operation controls, such as selecting songs, adjusting sound effects, and triggering interactions. On other displays (such as the rear entertainment screen or the center console display), the karaoke interface may focus more on viewing and experience, such as displaying only lyrics, song information, and interactive feedback effects, such as like animations and lighting effects, without including complex operation menus.
[0029] Optionally, in response to an action received on the karaoke interface of one display screen, other displays will also respond synchronously. For example, when a user clicks "switch songs" or "adjust accompaniment volume" on the central control screen, the karaoke interfaces on other displays will simultaneously update song information or volume status to ensure that the displayed content on all screens remains consistent. Similarly, when a user triggers a "like" interaction on the rear entertainment screen, other displays will simultaneously display a like animation or lighting effect, allowing all passengers in the vehicle to experience the interactive atmosphere.
[0030] For example, the central control display shows a karaoke interface, where users can request to play the first song. The karaoke interface for the first song is displayed, and the car's audio system plays the accompaniment audio of the first song. Optionally, the audio of the first song (including both the original vocals and the accompaniment) can also be played.
[0031] Step 220: Record ambient audio using the vehicle's microphone. The ambient audio includes human voices and accompaniment.
[0032] In a car karaoke scenario, the in-vehicle terminal captures the user's voice through a microphone installed inside the vehicle. Alternatively, the in-vehicle terminal can capture the user's voice through a handheld microphone connected to it.
[0033] For example, the in-vehicle terminal records ambient audio through a microphone. Since the in-vehicle audio system is playing the accompaniment of a song, the microphone will inevitably pick up the accompaniment playing in the cabin while collecting the user's singing voice. Therefore, the final recorded ambient audio is a mixed audio that includes human voice and accompaniment.
[0034] Step 230: If the human voice in the ambient audio matches the original audio of the first song, display interactive effects on the karaoke interface.
[0035] For example, human voice audio is extracted from ambient audio, and the user's vocal audio is matched with the original vocal audio of the first song (i.e., the original singer's audio). If the matching degree is high, interactive effects are automatically triggered.
[0036] For example, matching criteria can be set based on at least one dimension, such as audio pitch accuracy, rhythmic synchronization, pitch stability, and emotional expression matching. For instance: 1) The system uses a real-time pitch accuracy evaluation model to determine the degree of matching between the user's singing voice and the pitch curve of the original song. A lightweight model extracts feature vectors from the user's vocal audio and the original audio, and calculates their cosine similarity to obtain a matching score. When this score exceeds a preset threshold, the matching condition is met, triggering interactive effects.
[0037] 2) According to the pitch deviation of each note, when the deviation between the pitch of the human voice audio sung by the user and the standard pitch of the original singing audio is within a preset range (such as ±50 cents), it is determined that the matching degree condition is satisfied.
[0038] 3) Detect the alignment degree between the beat of the human voice audio sung by the user and the beat of the original singing audio, when the beat error between the human voice audio and the original singing audio is less than a preset time (such as 0.1 seconds), it is deemed that the matching degree condition is satisfied.
[0039] 4) Evaluate the fluctuation amplitude of the pitch during the user's continuous singing, if the pitch of the human voice audio remains stable and close to that of the original singing audio for several consecutive seconds, it is determined that the matching degree condition is satisfied.
[0040] 5) Analyze detailed features such as intensity, vibrato and portamento of the human voice audio, compare them with the emotional expression features of the original singing audio, and determine that the matching degree condition is satisfied when the matching degree is higher than a threshold.
[0041] In an alternative embodiment, scores of multiple dimensions including pitch, rhythm, stability and emotion can be weighted and summed, and it is determined that the matching degree condition is satisfied when the total score exceeds a matching degree threshold. The matching degree threshold can be dynamically adjusted according to the difficulty of the song or the user's historical singing level, for example, the threshold is lowered for high-difficulty songs, or more lenient matching conditions are provided for novice users.
[0042] In an alternative embodiment, the song can also be divided into paragraphs such as prelude, verse and chorus, and different matching degree conditions are set for different paragraphs. For example, a higher matching degree is required for the chorus to trigger a stronger interactive special effect.
[0043] Interactive special effects are visual, tactile or auditory feedback effects triggered during in-vehicle karaoke, which aim to simulate the easter egg experience of audience appreciation or gift-giving, thereby enhancing the user's sense of immersion and emotional value when karaoking in the cockpit. Optionally, the interactive special effects may include dynamic effects displayed on a display screen, such as rose special effect, heart animation, cheering GIF, ferris wheel gift special effect, cheering special effect, glow stick special effect, etc.
[0044] Optionally, the interactive special effects can also include linkage effects with on-board devices, for example: interior lighting effects, exterior lighting effects, seat vibration or rocking, air conditioner temperature or air volume adjustment, window lifting and lowering, car horn, etc. For example, the air conditioner air volume is automatically adjusted or the windows are automatically raised during the highlight moment of singing, so as to further enhance the user's sense of participation and fun.
[0045] For example, as Figure 3 shown, a glow stick special effect 301 is displayed on the karaoke interface; or, as Figure 4 shown, a cheering special effect 302 is displayed on the karaoke interface.
[0046] In one optional embodiment, the vehicle terminal can acquire the original accompaniment audio of the first song; mix the original accompaniment audio and the vocal audio to obtain the user's karaoke creation.
[0047] Optional, such as Figure 5 As shown, the in-vehicle terminal separates the user's vocals (i.e., human voice audio) from the recorded ambient audio, and enhances the user's vocals to obtain the user's dry vocals. The media music 402 of the first song undergoes acoustic separation to obtain the instrumental accompaniment (i.e., the original instrumental accompaniment audio). Subsequently, the user's dry vocals and the instrumental accompaniment are mixed in real time to obtain the user's finished product 403. Simultaneously, the in-vehicle terminal can also use the user's dry vocals and the original vocal audio to evaluate pitch accuracy and obtain a matching score. When the matching score meets certain conditions, a highlight interaction is triggered.
[0048] The mixing process is as follows: Figure 6 As shown, the instrumental music extracted from the media music of the first song is mixed with the user's dry vocals after recording enhancement to obtain the mixed user work 403.
[0049] Optionally, the user's karaoke performance may include interactive effects triggered when matching conditions are met. For example, when the user's karaoke performance is audio, it may include sound effects from interactive effects (such as clapping sound effects, cheering sound effects, etc.); when the user's karaoke performance is video, it may include dynamic visuals and / or sound effects from interactive effects.
[0050] In summary, the method provided in this embodiment introduces a real-time pitch evaluation and highlight interaction mechanism, solving the problems of monotonous presentation and lack of entertainment in in-vehicle karaoke in related technologies. By analyzing the matching degree between the user's singing and the original singer in real time, when a high matching degree is detected in the user's singing, various interactive effects such as thumbs-up animations and cheering effects are automatically triggered. This enhances the user's immersion in karaoke within the cabin, provides emotional value to the user, compensates for the lack of entertainment in related technologies, and improves the user's participation and enjoyment in in-vehicle karaoke.
[0051] An example is provided: an embodiment for calculating the matching score in real time.
[0052] Figure 7 This is a flowchart illustrating an exemplary embodiment of an in-vehicle karaoke interactive method provided in this application. This method can be used for, for example... Figure 1 The vehicle-mounted terminal shown is based on... Figure 2 In the illustrated embodiment, step 230 includes steps 231 to 234.
[0053] Step 210: Display the karaoke interface for the first song and play the accompaniment for the first song.
[0054] Step 220: Record ambient audio using the vehicle's microphone. The ambient audio includes human voices and accompaniment.
[0055] Step 231: Separate the vocals and accompaniment from the ambient audio to obtain the vocal audio and accompaniment audio.
[0056] For example, such as Figure 8 As shown, AGC (Automatic Gain Control) is used to adjust the signal strength of the ambient audio, AEC (Acoustic Echo Chancellor) is used to cancel the echo in the ambient audio, and then noise reduction processing is performed to remove noise. Figure 9 As shown, the ambient audio is then processed by sound separation 404 based on the original accompaniment audio of the first song to obtain the human voice audio and the accompaniment audio.
[0057] For example, the ambient audio is first preprocessed, including at least one of AGC, AEC, and noise reduction processing.
[0058] AGC automatically adjusts the signal strength of ambient audio to ensure that the volume of the audio signal remains at a stable and appropriate level, avoiding the impact of excessive or insufficient volume on subsequent processing.
[0059] Because the car audio system is playing background music, the microphone picks up both the user's singing and the audio from the system, creating an echo. AEC technology analyzes the reference signal from the audio playback to remove this echo interference from the ambient audio, making the recording closer to the user's pure vocals.
[0060] Noise reduction processing further removes background noise from the ambient audio, such as wind noise, tire noise, and air conditioning noise from a moving vehicle, resulting in a clearer audio signal.
[0061] After preprocessing, the ambient audio is then separated from the background music using a sound separation model (such as open-source models like mel-ro-former) based on the original instrumental audio of the first song. This sound separation model analyzes the spectral characteristics of the mixed audio and uses known original instrumental audio as supplementary information to more accurately separate the user's voice from the accompaniment, ultimately resulting in two independent audio streams: a clean vocal audio and an instrumental audio.
[0062] Step 232: Obtain the original audio of the first song.
[0063] The original vocal audio is the vocal portion of the original version of the song, that is, the pure vocal audio sung by the original singer. The purpose of obtaining the original vocal audio is to use it as a reference standard to evaluate the pitch matching of the user's vocal audio, thereby judging the accuracy of the user's singing and triggering corresponding interactive effects.
[0064] For example, the original audio of the first song can be obtained from a cloud music server or streaming platform via the network; or, if the first song has been downloaded to the local storage of the vehicle terminal, the vehicle terminal can directly read the original audio from the local storage; or, the original vocal track can be separated from the original music stream of the first song being played using a vocal separation model (such as open source models like mel-ro-former), thereby obtaining the original audio.
[0065] Step 233: Calculate the matching score based on the original singer's characteristics in the original audio and the vocal characteristics in the human voice audio.
[0066] For example, a lightweight feature extraction network is used to process the original vocal audio and the vocal audio separately. The lightweight feature extraction network can consist of at least one convolutional layer. It resamples the audio signal to 16kHz, extracts the 40-band Mel spectrum, and then extracts feature vectors using a lightweight CNN (Convolutional Neural Network). Feature vectors are accumulated for a period of time (e.g., 10 seconds) for both the original vocal audio and the vocal audio. The in-vehicle terminal then calculates the cosine similarity between these two accumulated feature vectors to obtain a matching score. This matching score is then processed through a fully connected layer and a sigmoid activation function, ultimately outputting a matching score between 0 and 1. A higher matching score indicates a higher degree of matching between the user's performance and the original vocals.
[0067] For example, while the user is singing the first song, the in-vehicle terminal continuously records ambient audio and extracts the human voice audio from it. A lightweight CNN feature extraction network processes each audio segment with minimal latency (e.g., tens of milliseconds per frame), resampling it to 16kHz, extracting the 40-band Mel spectrum, and outputting the corresponding feature vector. This process is streaming and real-time, without waiting for the entire song to finish.
[0068] The in-vehicle terminal maintains a sliding window with a length of 10 seconds. As the user begins singing, the terminal continuously stores the feature vector extracted at that moment into this window. For example, at the first second of the user's singing, the window only contains the features from the first second. At the fifth second, the window accumulates the features from the first to the fifth second. At the tenth second, the window has accumulated the complete 10-second feature vector. At the eleventh second, the window slides, discarding the features from the first second and adding the features from the eleventh second, at which point the window contains the features from the second to the eleventh second. At each moment, the in-vehicle terminal calculates the cosine similarity between the accumulated "10-second feature vector of the human voice audio" in the current window and the "10-second feature vector of the original vocal audio" extracted by the same network. This calculation is performed in real time, allowing the in-vehicle terminal to output a matching score every second (or even less).
[0069] For example, when a user starts singing, the in-vehicle terminal is accumulating features from the first 10 seconds. At this point, the window is not full, so scoring may not be performed or the scoring may be inaccurate. At the end of the 10th second of the performance, the in-vehicle terminal has accumulated a complete 10 seconds of user features for the first time. It then calculates the cosine similarity between these 10 seconds of user features and the features from the first 10 seconds of the original song, obtaining a score, for example, 0.85 (high match). The in-vehicle terminal determines that the condition is met and immediately triggers an interactive effect. From the 11th to the 20th second of the performance, the window slides, and the in-vehicle terminal continues to calculate the feature matching degree for the 2nd to the 11th second, the 3rd to the 12th second, and so on. If the user's pitch deviates at the 15th second, the score calculated by the in-vehicle terminal may drop to 0.60 (low match), and the interactive effect stops. From the 21st to the 30th second of the performance, the chorus begins, and the singing performance improves. The in-vehicle terminal calculates the feature matching degree for the 12th to the 22nd second, the score rises back to 0.90, and the in-vehicle terminal triggers the interactive effect again.
[0070] For example, such as Figure 10 As shown, feature extraction is performed on the user's dry voice 501 (i.e., human voice audio) and the music vocals 502 (i.e., original vocal audio). The extracted features are accumulated in the time domain to obtain the accumulated feature vectors of the user's dry voice and the music vocals, respectively. Real-time pitch evaluation is performed based on the two accumulated feature vectors to obtain the current time-time score 503 (i.e., matching score).
[0071] Feature extraction methods can be such as Figure 11 As shown, the audio is first resampled 601, and the sampling rate is adjusted to 16kHz. Then, the resampled audio is converted into Mel spectrum based on 40 frequency bands 602. Then, a lightweight CNN with shared weights is used to extract features from the Mel spectrum 603, and finally the feature vector is output 604.
[0072] In one optional embodiment, a short-time Fourier transform is performed on the original vocal audio to obtain at least one frame of the original vocal spectrum; a short-time Fourier transform is performed on the vocal audio to obtain at least one frame of the vocal spectrum. The at least one frame of the original vocal spectrum is input into a feature extraction network to obtain at least one frame of original vocal sub-features; the at least one frame of the vocal spectrum is input into a feature extraction network to obtain at least one frame of vocal sub-features. A target time window is determined according to a preset duration and the current playback progress; the at least one frame of original vocal sub-features corresponding to the target time window is concatenated to obtain the original vocal features; the at least one frame of vocal sub-features corresponding to the target time window is concatenated to obtain the vocal features. The cosine similarity between the original vocal features and the vocal features is calculated. The cosine similarity is sequentially input into a fully connected layer and an activation layer to obtain a matching score.
[0073] based on Figure 11 The feature extraction network shown can calculate the matching score in the following way: Figure 12 As shown, the user's vocals (701) and media music (702) are processed separately. The vocals (701), after recording enhancement, are combined with the music accompaniment obtained through vocal accompaniment separation in the real-time mixing stage to form the user's work. Simultaneously, the vocals and media music, after processing, are input into the feature extraction network (703) to obtain 10-second feature vectors. These two 10-second feature vectors are then input into the real-time pitch evaluation module (704). First, the cosine similarity is calculated, and then the vectors are processed through fully connected layers and activation layers (e.g., sigmoid activation layers) to finally output a matching score. The training process of the real-time pitch evaluation module (704) is as follows: During the training phase, based on training data containing binary labels (0 representing low matching score, 1 representing high matching score), the 10-second feature vectors output by the feature extraction network are input into the real-time pitch evaluation module for processing. The loss is calculated using the BCE loss function, and this loss is used to optimize the model to improve the accuracy of the model's matching score prediction.
[0074] Step 234: If the matching score is higher than the matching threshold, display interactive effects on the karaoke interface.
[0075] Optionally, in a car karaoke scenario, there may be multiple users singing karaoke simultaneously. In this case, the car microphones include at least two car microphones respectively located at at least two seats; the ambient audio includes at least two ambient sub-audio recordings obtained simultaneously by at least two car microphones.
[0076] When performing human voice audio separation, it is necessary to separate the human voice audio of each user separately: The accompaniment in at least two environmental sub-audios is removed using the original acoustic accompaniment audio of the first song to obtain at least two human voice sub-audios. These at least two human voice sub-audios are used for sound source localization or seat binding, as well as for extracting human voice audio from independent sound sources. Sound source localization is performed on the at least two human voice sub-audios to determine the sound source orientation of each seat in the vehicle. Based on the sound source orientation, a beamforming algorithm is applied to extract the preliminary human voice audio corresponding to each seat from the at least two human voice sub-audios. The at least two human voice sub-audios corresponding to at least two in-vehicle microphones are input into a speech separation model to obtain at least two human voice audios from independent sound sources; the speech separation model is used to separate mixed human voice signals. The at least two human voice sub-audios are matched with the preliminary human voice audios corresponding to at least two seats to determine the seats corresponding to the at least two human voice sub-audios.
[0077] Then, a matching score is assigned to each user based on their voice audio. That is, the matching score includes the matching score for the voice audio of each seat. Interactive effects are triggered based on the matching score.
[0078] When only one user's matching score meets the conditions, taking the first user in the first seat as an example, at least one of the following interactions can be triggered.
[0079] 1) If the matching score of the first seat is higher than the matching threshold, an interactive effect will be displayed on the in-vehicle display screen corresponding to the first seat, and a prompt message indicating that the first seat has obtained the interactive effect will be displayed on the display screens corresponding to the other seats.
[0080] 2) If the matching score of the first seat is higher than the matching threshold, control the interior lights to play a flowing light effect starting from the first seat; or, control the interior lights at the first seat to flash; or, control the interior lights of the other seats to play a flowing light effect pointing towards the first seat.
[0081] 3) If the matching score of the first seat is higher than the matching threshold, activate the seat massage function of the first seat.
[0082] When multiple users meet the matching score criteria, taking the first user in the first seat and the second user in the second seat as an example, at least one of the following interactions can be triggered.
[0083] 1) If the matching score of at least two seats is higher than the matching threshold at the same time, control the roof lights corresponding to at least two seats to flash; or, control the interior lights to play a flowing light effect starting from at least two seats.
[0084] 2) If the matching score of at least two seats is higher than the matching threshold at the same time, display the chorus gift-giving control on the in-vehicle display screens of the remaining seats except for the at least two seats; in response to the gift-giving operation triggered by the chorus gift-giving control, display the gift effect of receiving the gift on at least two in-vehicle display screens corresponding to the at least two seats.
[0085] For example, interactive effects may include any of the types mentioned above. The trigger type for an interactive effect can be determined based on the matching score, with each matching score range corresponding to a different type of interactive effect. Alternatively, the trigger type can be determined based on the duration for which the matching score is above a matching threshold; the longer the duration, the better the interactive effect. If the user's matching score remains above the matching threshold, various interactive effects can be continuously triggered. Alternatively, the trigger type can be determined based on the lyrics being sung when the matching score is above the matching threshold; the terminal device selects and triggers an interactive effect that matches the user's currently sung lyrics.
[0086] For example, the in-vehicle terminal triggers a corresponding interactive mechanism based on the meaning of the lyrics. Based on the current playback progress of the first song, a semantic recognition model is used to perform environmental semantic recognition on the context lyrics of the first song, obtaining environmental keywords. These environmental keywords describe the environment that fits the context lyrics. Based on the environmental keywords, the in-vehicle lights and / or the in-vehicle air conditioning are controlled.
[0087] For example, if the ambient keyword includes "wind," increase the airflow of the car's air conditioning. If the ambient keyword includes "cold," lower the temperature of the car's air conditioning. If the ambient keyword includes "hot," raise the temperature of the car's air conditioning. If the ambient keyword includes "bright," turn on the interior lights, or increase the brightness of the interior lights. If the ambient keyword includes "dark," turn off the interior lights, or decrease the brightness of the interior lights.
[0088] In one optional embodiment, the method provided in this application is executed when the in-vehicle karaoke mode is enabled. The in-vehicle terminal then enables the in-vehicle karaoke mode when preset vehicle safety conditions are met; wherein, the vehicle safety conditions include the vehicle speed being below a speed threshold or the vehicle being parked.
[0089] Optionally, the in-vehicle karaoke mode can also support multi-vehicle joint activation. For example, if at least two vehicles activate the in-vehicle karaoke mode and enter the same virtual room, the karaoke systems of at least two vehicles can be linked. Taking at least two vehicles (including vehicle 1, vehicle 2, vehicle 3, and vehicle 4) as an example, when a user in vehicle 1 sings and the matching score is higher than a threshold, joint interactive effects of at least two vehicles can be triggered. For example, the joint interactive effect could be the headlights of vehicle 1, vehicle 2, vehicle 3, and vehicle 4 lighting up sequentially; or, the joint interactive effect could be vehicle 2, vehicle 3, and vehicle 4 honking their horns simultaneously.
[0090] In one alternative embodiment, to avoid excessive noise from in-car karaoke, both in-vehicle and out-of-vehicle noise can be measured. If the in-vehicle noise is higher than the out-of-vehicle noise, and the in-vehicle noise exceeds a noise threshold, the car windows are closed.
[0091] In summary, the method provided in this embodiment introduces a real-time pitch evaluation and highlight interaction mechanism, solving the problems of monotonous presentation and lack of entertainment in in-vehicle karaoke in related technologies. By analyzing the matching degree between the user's singing and the original singer in real time, when a high matching degree is detected in the user's singing, various interactive effects such as thumbs-up animations and cheering effects are automatically triggered. This enhances the user's immersion in karaoke within the cabin, provides emotional value to the user, compensates for the lack of entertainment in related technologies, and improves the user's participation and enjoyment in in-vehicle karaoke.
[0092] The method provided in this embodiment designs a lightweight real-time evaluation model for user pitch accuracy, which can be deployed on edge platforms with limited system computing power, such as in-vehicle systems. This ultra-lightweight model is a real-time scoring model for user pitch highlights. Real-time requirements are considered during training. This model does not require strict extraction and matching of user pitch accuracy, nor does it require multi-pitch time alignment. Instead, it focuses on evaluating the user's pitch accuracy within a short period, achieving lighter and faster end-to-end scoring inference. This pitch evaluation model prioritizes pitch evaluation over strict pitch scoring. While providing users with real-time playback of their performances (online audio mixes), this method also provides interactive displays during periods of high-quality singing, enhancing the user's karaoke immersion.
[0093] The method provided in this embodiment solves the problems of existing in-car karaoke presentation methods being monotonous and lacking in entertainment value by judging the matching degree between the user's voice and the original audio in real time during the in-car karaoke process and displaying interactive effects when conditions are met. It provides users with immediate positive feedback, enhances the immersiveness and emotional value of singing, and improves the fun and interactivity of in-car karaoke.
[0094] The method provided in this embodiment offers a specific and quantifiable interactive triggering mechanism by separating ambient audio, acquiring the original vocal audio, and calculating a matching score. Compared to directly judging mixed audio, separating the pure vocals and then comparing them with the original vocal characteristics can more accurately and objectively evaluate the user's singing level, thereby ensuring that the triggering timing of interactive effects is more precise and reasonable.
[0095] The method provided in this embodiment achieves a lightweight, real-time pitch evaluation method through short-time Fourier transform, feature extraction network, sliding time window, and cosine similarity calculation. This method eliminates the need for complex pitch alignment processing and can dynamically output a matching score in real time based on the user's singing performance over a past period (e.g., 10 seconds) with low computational overhead. This ensures the immediacy and smoothness of interactive feedback, making it particularly suitable for deployment on platforms with limited computing power, such as in-vehicle infotainment systems.
[0096] The method provided in this embodiment utilizes multi-channel recording via in-vehicle microphones at multiple seats, combined with sound source localization, beamforming, and speech separation technologies, to accurately separate the independent vocal audio of each user from mixed ambient audio. This provides a technological foundation for karaoke scenarios with multiple seats and multiple users, enabling the system to identify and distinguish singers in different seats, thus creating conditions for providing personalized interactive feedback for different seats.
[0097] The method provided in this embodiment greatly enhances the fun and personalized experience of multi-user karaoke scenarios by providing an independent matching score for each seat and triggering differentiated interactive feedback accordingly. For example, when a user sings well, not only will their own screen display special effects, but the screens of other seats will also receive a notification. At the same time, the interior lights and seats will also produce synchronized effects (such as flowing light, flashing, and massage), creating an immersive atmosphere similar to a stage performance and enhancing the interaction and emotional resonance between users.
[0098] The method provided in this embodiment offers a joint interactive feedback mechanism for multi-seat chorus scenarios. When multiple seats simultaneously achieve a high matching degree, the system triggers coordinated lighting effects (such as multi-starting-point flowing light) and displays a "Chorus Gift Sending" control on the screens of other seats, allowing other passengers to send gifts to the chorus members. This simulates the interactive mode in live streaming or KTV, encourages multi-person collaborative singing, and further enhances the social attributes and entertainment value of in-vehicle karaoke.
[0099] The method provided in this embodiment generates a user-exclusive karaoke performance by mixing the original accompaniment with the user's vocals in real time. This provides users with instant playback and saving of their performances, fulfilling their potential need to record and share their singing, and enhancing the product's practical value and user engagement.
[0100] The method provided in this embodiment analyzes the context of the lyrics using a semantic recognition model and intelligently controls the in-car lights and air conditioning accordingly, achieving a deep integration of the karaoke experience and the cabin environment. For example, when the lyrics mention "wind," the air conditioning automatically increases the airflow. This cross-modal linkage greatly enhances the user's immersion and sense of scene engagement, upgrading the karaoke experience from a simple auditory enjoyment to a comprehensive sensory experience.
[0101] The method provided in this embodiment maps semantic keywords in lyrics (such as wind, cold, hot, bright, and dark) to specific control commands for vehicle air conditioning and lighting, offering a simple, intuitive, and practical linkage control logic. This enables the system to automatically and accurately adjust the cabin environment according to the mood of the lyrics, without requiring manual operation by the user, further enhancing intelligence and convenience.
[0102] The method provided in this embodiment activates the karaoke mode by setting vehicle safety conditions (such as vehicle speed below a threshold or parking), ensuring driving safety. This avoids the distraction risks that may arise from using the karaoke function in dangerous scenarios such as high-speed driving, demonstrating a high level of attention to user safety and enabling the in-car karaoke function to provide entertainment for users while ensuring safety.
[0103] The method provided in this embodiment automatically closes the car windows when the noise inside the vehicle is significantly higher than the noise outside and affects the karaoke experience, by comparing the noise levels inside and outside the vehicle. This effectively isolates external noise while avoiding noise disturbance to others, improving the acoustic environment inside the cabin and thus optimizing the user's karaoke and in-ear monitoring experience. It ensures that the user can hear their own singing and accompaniment more clearly, improving the quality of their performance.
[0104] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0105] It should be noted that the order of the method steps provided in the embodiments of this application can be appropriately adjusted, and the steps can also be added or removed as appropriate. The steps in the above embodiments can also be arbitrarily combined to obtain new embodiments. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.
[0106] Figure 13 This is a schematic diagram of the structure of an in-vehicle karaoke interactive device provided in an exemplary embodiment of this application. The device is used to implement an in-vehicle terminal and includes: Display module 1001 is used to display the karaoke interface of the first song and play the accompaniment of the first song; Recording module 1002 is used to record ambient audio via a vehicle-mounted microphone, the ambient audio including human voices and accompaniment; The display module 1001 is used to display interactive effects on the karaoke interface when the human voice in the ambient audio meets the matching degree condition with the original audio of the first song.
[0107] In an optional embodiment, the device further includes: The scoring module 1003 is used to separate the human voice and accompaniment in the ambient audio to obtain the human voice audio and the accompaniment audio. The scoring module 1003 is used to obtain the original audio of the first song; The scoring module 1003 is used to calculate a matching score based on the original singer characteristics of the original audio and the human voice characteristics of the human voice audio. The display module 1001 is used to display interactive effects on the karaoke interface when the matching score is higher than the matching threshold.
[0108] In one optional embodiment, the scoring module 1003 is used to perform a short-time Fourier transform on the original vocal audio to obtain at least one frame of the original vocal spectrum; and to perform a short-time Fourier transform on the human voice audio to obtain at least one frame of the human voice spectrum. The scoring module 1003 is used to input the at least one frame of original vocal spectrum into the feature extraction network to obtain at least one frame of original vocal sub-features; and to input the at least one frame of human voice spectrum into the feature extraction network to obtain at least one frame of human voice sub-features. The scoring module 1003 is used to determine a target time window according to a preset duration and the current playback progress; to splice at least one frame of original vocal features corresponding to the target time window to obtain the original vocal features; and to splice at least one frame of human voice features corresponding to the target time window to obtain the human voice features. The scoring module 1003 is used to calculate the cosine similarity between the original vocal features and the human vocal features; The scoring module 1003 is used to input the cosine similarity into the fully connected layer and the activation layer in sequence to obtain the matching score.
[0109] In one alternative embodiment, the vehicle-mounted microphones include at least two vehicle-mounted microphones respectively disposed at at least two seats; The ambient audio includes at least two ambient sub-audios recorded simultaneously by the at least two vehicle-mounted microphones; The scoring module 1003 is used to eliminate the accompaniment in the at least two environmental sub-audios using the original accompaniment audio of the first song to obtain at least two human voice sub-audios; the at least two human voice sub-audios are used for sound source localization or seat binding, and for human voice audio extraction from independent sound sources. The scoring module 1003 is used to locate the sound source of the at least two human voice sub-audio, and determine the sound source location of each seat in the vehicle. The scoring module 1003 is used to extract preliminary human voice audio corresponding to each seat from the at least two human voice sub-audio files by applying a beamforming algorithm based on the sound source location. The scoring module 1003 is used to input the at least two human voice sub-audios corresponding to at least two vehicle microphones into the speech separation model to obtain at least two human voice audios from independent sound sources; the speech separation model is used to separate mixed human voice signals. The scoring module 1003 is used to match the at least two human voice audios with the preliminary human voice audios corresponding to the at least two seats, and to determine the seats corresponding to the at least two human voice audios.
[0110] In one optional embodiment, the matching score includes a matching score corresponding to the human voice audio for each seat; the device further includes: The display module 1001 is used to display the interactive effect on the in-vehicle display screen corresponding to the first seat when the matching score of the first seat is higher than the matching threshold, and to display a prompt message that the first seat has obtained the interactive effect on the display screens corresponding to the other seats besides the first seat. The control module 1004 is used to control the interior lights to play a flowing light effect starting from the first seat when the matching score of the first seat is higher than the matching score threshold; or, control the interior lights at the first seat to flash; or, control the interior lights of the other seats to play a flowing light effect pointing towards the first seat. The control module 1004 is used to activate the seat massage function of the first seat when the matching score of the first seat is higher than the matching threshold.
[0111] In one optional embodiment, the matching score includes a matching score corresponding to the human voice audio for each seat; the device further includes: Control module 1004 is used to control the roof lights corresponding to the at least two seats to flash when the matching scores of at least two seats are simultaneously higher than the matching threshold; or to control the interior lights to play a flowing light effect starting from the at least two seats. The display module 1001 is used to display a chorus gift-giving control on the vehicle display screens of the seats other than the at least two seats when the matching score of at least two seats is higher than the matching threshold at the same time; in response to triggering the gift-giving operation of the chorus gift-giving control, a gift effect of receiving a gift is displayed on at least two vehicle display screens corresponding to the at least two seats.
[0112] In an optional embodiment, the device further includes: The generation module 1005 is used to obtain the original accompaniment audio of the first song; The generation module 1005 is used to mix the original accompaniment audio and the human voice audio to obtain the user's karaoke work.
[0113] In an optional embodiment, the device further includes: The control module 1004 is used to perform environmental semantic recognition on the context lyrics of the first song according to the current playback progress of the first song, and obtain environmental keywords; the environmental keywords are used to describe the environment that fits the context lyrics. The control module 1004 is used to control the interior lights and / or the vehicle air conditioner according to the environmental keywords.
[0114] In one alternative embodiment, the control module 1004 is configured to perform at least one of the following: When the environmental keyword includes wind, increase the air volume of the vehicle air conditioner; When the environmental keyword includes "cold", lower the temperature of the vehicle air conditioner; When the environmental keyword includes heat, the temperature of the vehicle air conditioner is increased; When the environmental keyword includes "bright", turn on the interior lights, or increase the brightness of the interior lights; When the environmental keyword includes darkness, turn off the interior lights, or reduce the brightness of the interior lights.
[0115] In an optional embodiment, the device operates when the in-car karaoke mode is activated, and the device further includes: Control module 1004 is used to activate the in-vehicle karaoke mode when preset vehicle safety conditions are met. The vehicle safety conditions include the vehicle speed being below a speed threshold or the vehicle being parked.
[0116] In an optional embodiment, the device further includes: Control module 1004 is used to measure in-vehicle noise and out-of-vehicle noise; The control module 1004 is used to close the car window when the noise inside the vehicle is higher than the noise outside the vehicle and the noise inside the vehicle is higher than a noise threshold.
[0117] It should be noted that the in-vehicle karaoke interactive device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the in-vehicle karaoke interactive device and the in-vehicle karaoke interactive method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0118] Embodiments of this application also provide a computer device, comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the in-vehicle karaoke interactive method provided in the above-described method embodiments. This computer device can be implemented as an in-vehicle terminal.
[0119] For example, Figure 14 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application.
[0120] Typically, computer device 1700 includes a processor 1701 and a memory 1702.
[0121] Processor 1701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0122] The memory 1702 may include one or more computer-readable storage media, which may be non-transitory. The memory 1702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1702 is used to store at least one instruction, which is executed by the processor 1701 to implement the in-vehicle karaoke interactive method provided in the method embodiments of this application.
[0123] In some embodiments, the computer device 1700 may also optionally include a peripheral device interface 1703 and at least one peripheral device. The processor 1701, memory 1702, and peripheral device interface 1703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1703 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1704, a display screen 1705, a camera assembly 1706, an audio circuit 1707, and a power supply 1708.
[0124] Peripheral device interface 1703 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1701 and memory 1702. In some embodiments, processor 1701, memory 1702 and peripheral device interface 1703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1701, memory 1702 and peripheral device interface 1703 can be implemented on separate chips or circuit boards, which is not limited in this application embodiment.
[0125] The radio frequency (RF) circuit 1704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1704 can communicate with other computer devices via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1704 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0126] Display screen 1705 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1705 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1701 for processing. In this case, display screen 1705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1705, which is located on the front panel of computer device 1700; in other embodiments, there may be at least two display screens 1705, respectively located on different surfaces of computer device 1700 or in a folded design; in still other embodiments, display screen 1705 may be a flexible display screen, located on a curved or folded surface of computer device 1700. Furthermore, display screen 1705 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1705 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0127] The camera assembly 1706 is used to acquire images or videos. Optionally, the camera assembly 1706 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the computer device 1700, and the rear-facing camera is located on the back of the computer device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1706 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0128] The audio circuit 1707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1701 for processing, or to the radio frequency circuit 1704 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations within the computer device 1700. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1701 or the radio frequency circuit 1704 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1707 may also include a headphone jack.
[0129] Power supply 1708 is used to supply power to the various components in computer device 1700. Power supply 1708 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1708 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0130] In some embodiments, the computer device 1700 further includes one or more sensors 1709. The one or more sensors 1709 include, but are not limited to, an accelerometer 1710, a gyroscope 1711, a pressure sensor 1712, an optical sensor 1713, and a proximity sensor 1714.
[0131] Accelerometer 1710 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 1700. For example, accelerometer 1710 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1701 can control touchscreen display 1705 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1710. Accelerometer 1710 can also be used for games or for acquiring user motion data.
[0132] The gyroscope sensor 1711 can detect the orientation and rotation angle of the computer device 1700. The gyroscope sensor 1711 can work in conjunction with the accelerometer sensor 1710 to acquire 3D motion data from the user on the computer device 1700. Based on the data acquired by the gyroscope sensor 1711, the processor 1701 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0133] Pressure sensor 1712 can be disposed on the side bezel of computer device 1700 and / or on the lower layer of touch display screen 1705. When pressure sensor 1712 is disposed on the side bezel of computer device 1700, it can detect the user's grip signal on computer device 1700, and processor 1701 can perform left / right hand recognition or quick operation based on the grip signal collected by pressure sensor 1712. When pressure sensor 1712 is disposed on the lower layer of touch display screen 1705, processor 1701 can control operable controls on the UI interface based on the user's pressure operation on touch display screen 1705. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0134] Optical sensor 1713 is used to collect ambient light intensity. In one embodiment, processor 1701 can control the display brightness of touch display screen 1705 based on the ambient light intensity collected by optical sensor 1713. Specifically, when the ambient light intensity is high, the display brightness of touch display screen 1705 is increased; when the ambient light intensity is low, the display brightness of touch display screen 1705 is decreased. In another embodiment, processor 1701 can also dynamically adjust the shooting parameters of camera assembly 1706 based on the ambient light intensity collected by optical sensor 1713.
[0135] The proximity sensor 1714, also known as a distance sensor, is typically located on the front panel of the computer device 1700. The proximity sensor 1714 is used to detect the distance between the user and the front of the computer device 1700. In one embodiment, when the proximity sensor 1714 detects that the distance between the user and the front of the computer device 1700 is gradually decreasing, the processor 1701 controls the touch display screen 1705 to switch from a screen-on state to a screen-off state; when the proximity sensor 1714 detects that the distance between the user and the front of the computer device 1700 is gradually increasing, the processor 1701 controls the touch display screen 1705 to switch from a screen-off state to a screen-on state.
[0136] Those skilled in the art will understand that Figure 14 The structure shown does not constitute a limitation on the computer device 1700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0137] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. When the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor of a computer device, the in-vehicle karaoke interactive method provided in the above-described method embodiments is implemented.
[0138] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the in-vehicle karaoke interactive method provided in the above-described method embodiments.
[0139] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0140] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent switching, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for in-vehicle karaoke interaction, characterized in that, The method is executed by an in-vehicle terminal, and the method includes: Display the karaoke interface for the first song and play the accompaniment for the first song; Ambient audio, including human voices and accompaniment, is recorded using a vehicle-mounted microphone. If the human voice in the ambient audio matches the original audio of the first song, interactive effects are displayed on the karaoke interface.
2. The method according to claim 1, characterized in that, When the human voice in the ambient audio matches the original audio of the first song, interactive effects are displayed on the karaoke interface, including: Separate the human voice and accompaniment from the ambient audio to obtain the human voice audio and the accompaniment audio; Obtain the original audio of the first song; A matching score is calculated based on the original vocal characteristics of the original audio and the vocal characteristics of the human voice audio. If the matching score is higher than the matching threshold, interactive effects will be displayed on the karaoke interface.
3. The method according to claim 2, characterized in that, The step of calculating a matching score based on the original vocal features of the original audio and the vocal features of the human voice audio includes: Perform a short-time Fourier transform on the original vocal audio to obtain at least one frame of the original vocal spectrum; perform a short-time Fourier transform on the human voice audio to obtain at least one frame of the human voice spectrum; The original vocal spectrum of at least one frame is input into the feature extraction network to obtain at least one frame of original vocal sub-features; the human voice spectrum of at least one frame is input into the feature extraction network to obtain at least one frame of human voice sub-features. Determine the target time window based on the preset duration and the current playback progress; stitch together at least one frame of original vocal features corresponding to the target time window to obtain the original vocal features; stitch together at least one frame of human voice features corresponding to the target time window to obtain the human voice features; Calculate the cosine similarity between the original vocal features and the human vocal features; The cosine similarity is sequentially input into the fully connected layer and the activation layer to obtain the matching score.
4. The method according to claim 2, characterized in that, The vehicle-mounted microphones include at least two vehicle-mounted microphones respectively installed at at least two seats; The ambient audio includes at least two ambient sub-audios recorded simultaneously by the at least two vehicle-mounted microphones; The process of separating the human voice and accompaniment from the ambient audio to obtain the human voice audio and accompaniment audio includes: The original accompaniment audio of the first song is used to remove the accompaniment in the at least two environmental sub-audio to obtain at least two human voice sub-audio; the at least two human voice sub-audio are used for sound source localization or seat binding, and for human voice audio extraction from independent sound sources; The sound source is located for the at least two human voice sub-audios, and the sound source location of each seat in the vehicle is determined. Based on the sound source location, a beamforming algorithm is applied to extract the preliminary human voice audio corresponding to each seat from the at least two human voice sub-audio files; The at least two human voice sub-audios corresponding to at least two vehicle-mounted microphones are input into the speech separation model to obtain at least two human voice audios from independent sound sources; the speech separation model is used to separate mixed human voice signals. Match the at least two human voice audios with the preliminary human voice audios corresponding to the at least two seats to determine the seats corresponding to the at least two human voice audios.
5. The method according to claim 4, characterized in that, The matching score includes a matching score for the human voice audio of each seat; the method also includes at least one of the following: If the matching score of the first seat is higher than the matching threshold, the interactive effect is displayed on the in-vehicle display screen corresponding to the first seat, and a prompt message indicating that the first seat has obtained the interactive effect is displayed on the display screens corresponding to the other seats besides the first seat. If the matching score of the first seat is higher than the matching threshold, control the interior lights to play a flowing light effect starting from the first seat; or, control the interior lights at the first seat to flash; or, control the interior lights of the other seats to play a flowing light effect pointing towards the first seat. If the matching score of the first seat is higher than the matching threshold, the seat massage function of the first seat will be activated.
6. The method according to claim 4, characterized in that, The matching score includes a matching score for the human voice audio of each seat; the method also includes at least one of the following: If the matching score for at least two seats is higher than the matching threshold at the same time, control the roof lights corresponding to the at least two seats to flash; or, control the interior lights to play a flowing light effect starting from the at least two seats. If the matching score for at least two seats is higher than the matching threshold at the same time, a chorus gift-giving control is displayed on the in-vehicle display screens of the remaining seats, excluding the at least two seats; in response to triggering the gift-giving operation of the chorus gift-giving control, a gift effect of receiving a gift is displayed on at least two in-vehicle display screens corresponding to the at least two seats.
7. The method according to any one of claims 2 to 6, characterized in that, The method further includes: Obtain the original instrumental audio of the first song; The original background music and the vocal audio are mixed to create the user's karaoke performance.
8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Based on the current playback progress of the first song, a semantic recognition model is used to perform environmental semantic recognition on the context lyrics of the first song to obtain environmental keywords; the environmental keywords are used to describe the environment that fits the context lyrics. Control the interior lights and / or the vehicle air conditioning based on the environmental keywords.
9. The method according to claim 8, characterized in that, The step of controlling the vehicle interior lights and / or the vehicle air conditioning based on the environmental keywords includes at least one of the following: When the environmental keyword includes wind, increase the air volume of the vehicle air conditioner; When the environmental keyword includes "cold", lower the temperature of the vehicle air conditioner; When the environmental keyword includes heat, the temperature of the vehicle air conditioner is increased; When the environmental keyword includes "bright", turn on the interior lights, or increase the brightness of the interior lights; When the environmental keyword includes darkness, turn off the interior lights, or reduce the brightness of the interior lights.
10. The method according to any one of claims 1 to 6, characterized in that, The method is executed when the in-car karaoke mode is enabled, and the method further includes: The in-vehicle karaoke mode is activated when the preset vehicle safety conditions are met. The vehicle safety conditions include the vehicle speed being below a speed threshold or the vehicle being parked.
11. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Measure interior and exterior noise levels; When the noise inside the vehicle is higher than the noise outside the vehicle, and the noise inside the vehicle is higher than the noise threshold, close the vehicle windows.
12. A car-based karaoke interactive device, characterized in that, The device is used to implement a vehicle-mounted terminal, and the device includes: The display module is used to display the karaoke interface of the first song and play the accompaniment of the first song; A recording module for recording ambient audio via an in-vehicle microphone, the ambient audio including human voices and accompaniment; The display module is used to display interactive effects on the karaoke interface when the human voice in the ambient audio matches the original audio of the first song.
13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the in-vehicle karaoke interactive method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one program, which is loaded and executed by a processor to implement the in-vehicle karaoke interactive method as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the in-vehicle karaoke interactive method as described in any one of claims 1 to 11.