A method, device, storage medium, and program product for multi-vehicle collaborative karaoke.

CN122575316APending Publication Date: 2026-08-14BEJING ANGEL VOICE DIGITAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这些处理环节的时间开销各不相同,导致各车辆在播放人声和伴奏时难以保持精确同步,容易出现人声与伴奏错位、不同车辆播放不一致等问题,影响多车协同K歌的实时互动体验

Benefits of technology

[0027]1、本申请通过将人声数据与伴奏数据进行链路解耦,在局域网建立低时延音频链路,并针对主唱与接收角色确立不同的时延补偿基准,消减了不同车辆间处理路径不一致造成的音频错位,提升了人声与伴奏的同步一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575316A_ABST
    Figure CN122575316A_ABST
Patent Text Reader

Abstract

This invention provides a method, device, storage medium, and program product for multi-vehicle collaborative karaoke, primarily addressing the problem of poor synchronization between vocals and accompaniment in multi-vehicle karaoke. The method includes: vehicle terminals receiving role information, low-latency audio link parameters, and reference time information from the cloud; receiving accompaniment audio via a wide area network and configuring either lead singer or receiver mode based on the role; broadcasting vocals via a local area link and processing them locally in lead singer mode, and receiving vocals via a local area link in receiver mode; determining different vocal path latency parameters based on the mode; compensating for the accompaniment output time based on the reference time and latency parameters; and locally mixing and outputting the vocals and the compensated accompaniment. This invention achieves high-precision synchronization between vocals and accompaniment in multi-vehicle scenarios through a dual-link separation architecture and differentiated latency compensation, thus improving the collaborative singing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of vehicle networking and in-vehicle multimedia technology, and in particular to a multi-vehicle collaborative karaoke method, device, storage medium and program product. Background Technology

[0002] With the development of in-vehicle entertainment systems, karaoke functionality has expanded from traditional KTV settings to in-vehicle scenarios. In outdoor settings such as RV campsites and road trips, convoys of multiple vehicles often require collaborative karaoke interaction across vehicles. Common in-vehicle karaoke solutions primarily target single-vehicle scenarios, where users capture their voice through the in-vehicle microphone, mix it with locally played accompaniment, and then output the result, meeting the karaoke needs within a single vehicle.

[0003] In a multi-vehicle collaborative karaoke scenario, all vehicles need to sing the same song together. This requires that the singer's voice be heard by other vehicles in real time to achieve collaborative singing; each vehicle needs to play the same accompaniment audio, lyrics subtitles, and music video content to maintain consistency in the singing.

[0004] Network-based multi-person karaoke solutions exist in related technologies. However, in practical applications, the transmission latency of human voice signals directly affects the singing experience. When the transmission latency exceeds 50ms, singers will clearly perceive the delay, affecting the coordination of collaborative singing. Furthermore, in outdoor vehicle convoy scenarios, vehicles may be scattered, network signal quality may be unstable, and human voice transmission is prone to latency fluctuations or interruptions.

[0005] Furthermore, the backing audio and music video content are large in size, requiring high network bandwidth when multiple devices request them simultaneously. When network conditions are limited, slow loading or playback stuttering can easily occur.

[0006] Furthermore, during multi-vehicle collaborative singing, there is a time lag in the process of different vehicles receiving and processing audio signals. For example, when a singer in one vehicle is rapping, their voice needs to go through acquisition and transmission processes before reaching other vehicles; while the accompaniment played by that vehicle itself undergoes a different processing flow. The varying time costs of these processing steps make it difficult for each vehicle to maintain precise synchronization when playing vocals and accompaniment, easily leading to problems such as misalignment of vocals and accompaniment, and inconsistent playback between different vehicles, thus affecting the real-time interactive experience of multi-vehicle collaborative karaoke. Summary of the Invention

[0007] This application provides a method, device, storage medium, and program product for multi-vehicle collaborative karaoke, which can improve the real-time interactive experience in multi-vehicle collaborative karaoke scenarios.

[0008] Firstly, this application provides a multi-vehicle collaborative karaoke method, which includes: receiving current role information, first low-latency wireless audio link parameters, and reference time information from a cloud server, wherein the first low-latency wireless audio link parameters are used to establish a first low-latency wireless audio link; receiving the accompaniment audio corresponding to the current song through a second wireless communication link, and configuring the vehicle terminal to lead singer mode or receive mode according to the current role information, wherein the first low-latency wireless audio link is a local area communication link within the fleet, and the second wireless communication link is a wide area network communication link; when in lead singer mode, transmitting the collected local microphone voice to other vehicle terminals in the fleet through the first low-latency wireless audio link, and transmitting the local microphone voice through the local audio input path. The audio is fed into the local audio processing module. When in receiving mode, the system receives the vocals sent by the current lead vocal vehicle terminal according to the parameters of the first low-latency wireless audio link, and sends the received vocals into the local audio processing module. The system determines the corresponding vocal path delay parameters according to the current mode. In lead vocal mode, the vocal path delay parameters are the local vocal input path delay, and in receiving mode, they are the vocal reception delay of the first low-latency wireless audio link. The system synchronizes with the cloud server through the second wireless communication link and compensates for the output time of the accompaniment audio based on the reference time information and the vocal path delay parameters. The system mixes the vocals entering the local audio processing module with the compensated accompaniment audio locally in the vehicle terminal to obtain synchronized playback audio and output it.

[0009] In the above embodiments, the system decouples latency-sensitive vocal data from bandwidth-intensive accompaniment data, avoiding vocal transmission delays caused by wide area network fluctuations in traditional pure cloud architectures. By establishing a low-latency audio link within the fleet's local area network and defining local input latency and wireless reception latency as compensation benchmarks for the lead singer and receiver roles respectively, the system can reverse-align the output time of the local accompaniment using a unified reference time as a standard. This role-differentiated dual-link latency compensation reduces audio misalignment caused by inconsistent processing paths between different vehicles, improving the synchronization consistency of vocals and accompaniment during multi-vehicle collaborative singing.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, after the vocals and compensated accompaniment audio entering the local audio processing module are mixed locally on the vehicle terminal to obtain synchronized playback audio and output, the method further includes: receiving a role switching instruction issued by a cloud server, the role switching instruction including at least the next lead vocal vehicle terminal identifier, the target operating frequency corresponding to the next lead vocal vehicle terminal identifier, the transmit / receive mode, and the switching effective time; when the switching effective time is reached, reconfiguring the vehicle terminal into lead vocal mode or receive mode according to the next lead vocal vehicle terminal identifier and the transmit / receive mode, and updating the parameters of the first low-latency wireless audio link.

[0011] In the above embodiments, the system pre-issues role-switching instructions containing a clear effective time and target configuration via a cloud server, changing the traditional method of relying on real-time handshake protocols for state transitions. Upon receiving the instruction, each vehicle terminal first completes parameter parsing and resource preparation, and synchronously executes the switching of transmit / receive modes and frequency at the designated time node. This mechanism ensures a unified timing throughout the entire convoy during lead vocal handover, avoiding audio overlap or gaps caused by instruction delays or differences in processing speed among vehicles, thus guaranteeing the continuity of the collaborative singing process.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, after reconfiguring the vehicle terminal into a singing mode or a receiving mode according to the next lead vocal vehicle terminal identifier and transmission / reception mode, and updating the first low-latency wireless audio link parameters, the method further includes: re-determining the vocal path delay parameters corresponding to the updated mode; when the updated mode is a singing mode, using the local vocal input path delay as a compensation benchmark; when the updated mode is a receiving mode, using the vocal reception delay of the first low-latency wireless audio link as a compensation benchmark; and re-compensating the output time of the accompaniment audio corresponding to the next song based on the re-determined vocal path delay parameters.

[0013] In the above embodiments, the system triggers a recalibration process for latency parameters after a mode update to address changes in the physical transmission path caused by a change in the lead singer's role. Since the data flow between the new lead singer and the new receiver is reversed, the system accordingly switches the compensation benchmark from the original link reception latency to the local input latency, or from the original local input latency to the new link reception latency. Based on this re-established benchmark, compensation calculations are performed on the accompaniment output time of the next song, enabling the local mixing logic to adapt to the dynamic reconstruction of the spatial topology and data flow within the vehicle, maintaining audiovisual synchronization during cross-track and cross-role handovers.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, when the vehicle terminal is in receiving mode, the corresponding human voice path delay parameter is determined according to the current mode. Specifically, this includes: parsing the transmission timestamp from the human voice data received through the first low-latency wireless audio link, the transmission timestamp being recorded by the current lead vehicle terminal based on reference time information when transmitting human voice data; obtaining the reception timestamp of the received human voice data based on the local clock of the vehicle terminal synchronized with the reference time information; calculating the time difference between the reception timestamp and the transmission timestamp, and determining the time difference as the human voice reception delay.

[0015] In the above embodiment, the system calculates the difference between the transmission timestamp and the reception timestamp generated based on a unified reference time. The receiving end parses the transmission time record carried in the voice data packet and obtains the actual arrival time by combining it with the local synchronization clock, thus quantifying the overall air interface transmission time and the underlying hardware buffering time. This measurement method based on the actual data packet flow cycle objectively reflects the actual communication cost under the current spatial distance and electromagnetic environment, providing reliable data support for audio output compensation.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, during the process of mixing the human voice and the compensated accompaniment audio in the local vehicle terminal, the method further includes: periodically repeating the steps of parsing the sending timestamp, obtaining the receiving timestamp, and calculating the time difference to obtain the updated human voice receiving delay; comparing the updated human voice receiving delay with the current human voice path delay parameter; when the difference between the updated human voice receiving delay and the current human voice path delay parameter exceeds a preset threshold, updating the human voice path delay parameter to the updated human voice receiving delay, and redetermining the compensation amount for the output time of the accompaniment audio based on the updated human voice path delay parameter.

[0017] In the above embodiments, the system introduces a periodic monitoring and threshold decision mechanism for human voice reception delay to address wireless link quality fluctuations caused by relative vehicle movement. By comparing newly acquired measurements with currently effective delay parameters, the system only triggers parameter updates and recalculation of compensation amounts when the difference exceeds a preset threshold. This mechanism strikes a balance between absorbing minor network jitter and responding to trending delay changes, avoiding audio pops or stutters caused by frequent adjustments to the accompaniment reading pointer, and maintaining the stability of the local mixing output.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, while receiving the accompaniment audio corresponding to the current song through the second wireless communication link, the method further includes: receiving lyrics subtitles and / or video images corresponding to the current song through the second wireless communication link; compensating the display time of the lyrics subtitles and / or video images based on reference time information and vocal path delay parameters; and synchronously displaying the compensated lyrics subtitles and / or video images while outputting synchronously played audio.

[0019] In the above embodiments, the system extends the compensation logic based on vocal path delay parameters from the audio dimension to the visual presentation dimension. By using the same reference time information and delay compensation amount to perform a complete translation of the display timeline of the lyrics subtitles and video images distributed over the wide area network, the system ensures that the progression rhythm of the visual cues and the compensated accompaniment audio maintain a strict lock-on state. This unified alignment processing of multimodal data allows singers in each vehicle to obtain an interactive experience that matches the auditory feedback when viewing the screen prompts, improving the guiding effect of collaborative singing.

[0020] In conjunction with some embodiments of the first aspect, in some embodiments, during the process of mixing the vocals and compensated accompaniment audio entering the local audio processing module locally on the vehicle terminal, the method further includes: monitoring the connection status of the first low-latency wireless audio link; when the first low-latency wireless audio link is detected to be interrupted or the signal quality is lower than a preset threshold, sending an abnormal notification to the cloud server through the second wireless communication link; receiving a processing instruction issued by the cloud server in response to the abnormal notification, the processing instruction being used to instruct the vehicle terminal to perform at least one of the following operations: switching to a backup operating frequency, triggering a reselection of the lead vocal vehicle terminal, or downgrading to a accompaniment-only playback mode.

[0021] In the above embodiments, the system constructs a fault-tolerant mechanism for complex outdoor electromagnetic environments by implementing link status monitoring on local terminals and combining it with global scheduling on cloud servers. When a local low-latency link encounters severe interference or interruption, the terminal actively reports the abnormal status, and the cloud uniformly issues processing strategies such as switching to backup frequencies, reselecting the main channel, or downgrading playback. This end-to-cloud collaborative anomaly handling process limits the negative impact of local communication failures to a controllable range, preventing a single vehicle disconnection from causing the entire convoy to fall into a state of disordered noise output, thus improving the overall robustness of the multi-vehicle collaborative system.

[0022] Secondly, embodiments of this application provide a multi-vehicle collaborative karaoke device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the multi-vehicle collaborative karaoke device to perform the method described in the first aspect and any possible implementation thereof.

[0023] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a multi-vehicle collaborative karaoke device, cause the multi-vehicle collaborative karaoke device to execute the method described in the first aspect and any possible implementation thereof.

[0024] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a multi-vehicle collaborative karaoke device, cause the multi-vehicle collaborative karaoke device to perform the method described in the first aspect and any possible implementation thereof.

[0025] Understandably, the multi-vehicle collaborative karaoke device provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.

[0026] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0027] 1. This application decouples the vocal data from the accompaniment data, establishes a low-latency audio link in the local area network, and establishes different latency compensation benchmarks for the lead singer and the receiving role, thereby reducing audio misalignment caused by inconsistent processing paths between different vehicles and improving the synchronization consistency of vocals and accompaniment.

[0028] 2. This application uses the cloud to pre-deliver role switching instructions containing a clear effective time and target configuration, so that each vehicle can synchronously execute mode flipping and frequency switching at the specified time node, avoiding audio overlap or gaps caused by network latency or processing speed differences, and ensuring the continuity of collaborative singing.

[0029] 3. This application uses the difference between the sending and receiving timestamps generated based on a unified reference time to objectively quantify the actual communication cost, and dynamically updates the compensation amount when the delay change exceeds the limit, effectively dealing with network fluctuations caused by the relative movement of outdoor vehicles and maintaining the stability of local mixing output. Attached Figure Description

[0030] Figure 1 This is a flowchart illustrating a multi-vehicle collaborative karaoke method in an embodiment of this application;

[0031] Figure 2 This is another flowchart illustrating the multi-vehicle collaborative karaoke method in this application embodiment;

[0032] Figure 3 This is a schematic diagram of the physical device structure of a multi-vehicle collaborative karaoke device in the embodiments of this application. Detailed Implementation

[0033] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.

[0034] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0035] The following describes a scenario where the multi-vehicle collaborative karaoke method described in this application is used.

[0036] Multi-vehicle collaborative karaoke scenarios can be found in RV campsites, self-driving camping, car enthusiast gatherings, or parked vehicle displays. The multiple vehicle terminals can be from different brands, have different vehicle operating systems, or different audio hardware platforms, but each vehicle can receive control commands, receive accompaniment resources, and perform local mixing through a unified application layer protocol.

[0037] In some specific instances, the system divides the data to be transmitted into two categories: real-time voice data and content distribution data. The first category of real-time voice data is more sensitive to end-to-end latency and jitter, while the second category of content distribution data focuses more on coverage, content integrity, and caching efficiency. Therefore, different wireless communication links are used to carry the two types of data.

[0038] In some specific instances, the cloud server does not carry out the real-time cloud mixing loop for the lead vocals, but mainly undertakes room management, role arrangement, distribution of low-latency link parameters, unified reference time release, song request queue management and exception coordination, thereby ensuring that the most real-time sensitive vocals are transmitted within a short-distance low-latency link.

[0039] In some preferred embodiments, the multi-vehicle collaborative karaoke system includes at least a cloud server, multiple vehicle-mounted terminals, a first wireless communication unit, a second wireless communication unit, and a local audio processing module for each vehicle-mounted terminal. The first wireless communication unit is preferably a UHF / UHF wireless microphone dongle or an integrated low-latency audio transceiver module, used for short-range broadcasting of the lead singer's voice within the vehicle fleet. The second wireless communication unit is preferably a 4G / 5G cellular communication module, used for interacting with the cloud server to exchange song selection information, accompaniment, lyrics, video, time information, and error handling commands.

[0040] In some preferred embodiments, each vehicle performs a device self-test and readiness confirmation before starting the current song. The device self-test includes at least: a first low-latency wireless audio link transceiver unit availability check, a second wireless communication link connectivity check, a pre-buffering check of the current song accompaniment, a local clock skew check, and a local vocal input path delay calibration value readability check.

[0041] In some preferred embodiments, when the second wireless communication link is in a weak network state but the first low-latency wireless audio link in the fleet remains available, the team leader vehicle or the pre-designated edge coordination vehicle can temporarily assume the local coordination function. The local coordination function includes at least: maintaining the current lead singer vehicle identifier locally, forwarding the most recently valid unified reference time information, publishing the unified start time of the next song, and transmitting the song request queue changes and role switching records generated during the local period back to the cloud server after the connection is restored in the cloud.

[0042] The following describes the process of the method provided in this implementation, based on the above scenario. Please refer to... Figure 1 This is a flowchart illustrating a multi-vehicle collaborative karaoke method in an embodiment of this application.

[0043] S101: Receive the current role information, first low-latency wireless audio link parameters and reference time information sent by the cloud server. The first low-latency wireless audio link parameters are used to establish the first low-latency wireless audio link.

[0044] The current role information refers to the singing identity identifier assigned to the vehicle terminal by the cloud server based on the song request queue and fleet topology. This includes at least the lead singer identifier or receiver identifier, the corresponding transmit / receive mode configuration, and the effective time period for the role. The first low-latency wireless audio link parameters refer to the configuration set used to establish a local voice transmission channel within the fleet. Specifically, this includes the operating frequency (e.g., a specified frequency point within the 430MHz-440MHz range), transmit power level, audio encoding format (e.g., ADPCM or Opus low-latency encoding), frame length configuration (e.g., 10ms or 20ms), CRC check mode, and device address binding table. Reference time information refers to the unified time reference generated by the cloud server based on the Network Time Protocol (NTP) or Global Navigation Satellite System (GNSS) timing source, distributed to each vehicle terminal with microsecond or nanosecond precision. It serves as the synchronization anchor point for all time-related operations within the fleet and typically includes the current reference timestamp, clock deviation tolerance range, and a preview of the next synchronization time.

[0045] After a vehicle joins the room and completes identity authentication, the cloud server comprehensively decides on the role allocation scheme based on factors such as the current song selection status, vehicle location, historical RF quality assessment results, and user active selection, and uniformly issues configuration information before the song starts or before role switching. Specifically, the vehicle terminal receives a JSON or Protobuf format configuration packet containing the above three types of information through a second wireless communication link (4G / 5G cellular network). After local parsing, it first verifies the validity of the timestamp. If the deviation from the local clock exceeds a preset threshold (e.g., 50ms), it actively initiates a time synchronization request. Subsequently, based on the role identifier, it switches the local state machine to the lead singer standby or receiver standby state, and writes the parameters of the first low-latency wireless audio link into the register of the UHF transceiver module, completing frequency locking, address filtering table configuration, and encoder initialization. To ensure synchronized start-up of the entire fleet, the cloud will issue a unified start-up time instruction only after all vehicles have completed configuration and reported their readiness status. This three-stage mechanism of pre-distribution configuration, centralized timing, and readiness confirmation ensures that each vehicle has completed resource readiness and parameter alignment before the song starts, avoiding the problem of loss of synchronization during broadcast due to network jitter in traditional real-time handshake protocols.

[0046] In some preferred embodiments, the configuration package issued by the cloud server includes two parts: a common parameter template and vehicle-specific parameters. The common parameter template defines the operating frequency of the first low-latency wireless audio link uniformly applicable to the current song, the unified reference time information, the target start time, and the current lead singer's vehicle identifier. The vehicle-specific parameters define the receiving gain, address filtering table, local compensation coefficient correction value, and device compatibility configuration corresponding to each vehicle terminal. If a vehicle terminal detects that the version number of the received configuration package is lower than the local cache version, the field integrity verification fails, or the role's effective time has expired, it actively initiates a parameter re-retrieval request through the second wireless communication link, and maintains the previous valid configuration unchanged until the updated configuration is obtained.

[0047] S102. Receive the accompaniment audio corresponding to the current song through the second wireless communication link, and configure the vehicle terminal into lead singer mode or receiving mode according to the current role information. The first low-latency wireless audio link is a local area communication link within the fleet, and the second wireless communication link is a wide area network communication link.

[0048] The accompaniment audio refers to the purely instrumental part of the current song, typically stored in AAC, MP3, or lossless FLAC formats on a cloud content delivery network (CDN). After being transmitted to the vehicle terminal via a second wireless communication link, it is stored in a local buffer for local mixing with the vocals. The lead vocal mode configures the vehicle terminal as the vocal signal source, responsible for collecting local microphone input and broadcasting it to other vehicles in the convoy via a first low-latency wireless audio link, while simultaneously monitoring the mixing of vocals and accompaniment locally. The receiving mode configures the vehicle terminal as the vocal receiver, only listening to the vocal data transmitted by the current lead vocal vehicle on the first low-latency wireless audio link, mixing it with the locally played accompaniment, and outputting no external local microphone signal. The first low-latency wireless audio link is a local broadcast link within the convoy based on the UHF band, with a typical air interface latency of less than 10ms, requiring no internet relay. The second wireless communication link is a wide-area cellular network link based on operator base stations, used for high-bandwidth content distribution and control signaling interaction.

[0049] After receiving the role information and link parameters from the cloud, the vehicle terminal initiates the accompaniment resource preloading process. It pulls the accompaniment audio file of the current song from the CDN edge node through the second wireless communication link, and fills the local playback buffer by using a streaming processing method of downloading and decoding each piece in segments to ensure that the cache depth is sufficient for at least 10 seconds of playback before the unified start time arrives.

[0050] Specifically, the terminal determines whether it should enter lead vocal mode or receive mode based on the identifier field in the role information: if designated as lead vocal, it activates the local audio input channel, sets the microphone device connected via USB or 3.5mm interface as the active input source, and configures the UHF transmission module to enter the transmit-ready state; if designated as receive, it configures the UHF receiving module to monitor the frequency and device address corresponding to the lead vocal vehicle, and simultaneously turns off or mutes the local microphone input to avoid echo interference. The core of this dual-link separation architecture is that: the accompaniment data volume is large (a single song can reach several MB to tens of MB) but the real-time requirements are relatively relaxed, and time alignment can be achieved through wide area network pre-caching and local scheduling; while the vocal data volume is small (several KB per second) but is extremely sensitive to end-to-end latency, and must be directly transmitted through a short-range low-latency link to ensure the real-time performance of the chorus, thereby avoiding the long-tail latency and bandwidth congestion risks caused by uploading all audio data to the cloud and then re-downloading it.

[0051] In some preferred embodiments, when the accompaniment resource for the current song is already pre-stored locally on the vehicle terminal, the cloud server does not need to repeatedly send the complete accompaniment file. Instead, it only sends the version number, integrity check value, unified start point, and corresponding lyrics subtitles and video image resource identifiers for the accompaniment resource. After verifying the validity of the local cached resource, the vehicle terminal directly reads the accompaniment from the local cache and performs output timing compensation according to the unified start point. Furthermore, when the bandwidth of the second wireless communication link is limited, the vehicle terminal prioritizes the acquisition and caching of the accompaniment audio and lyrics subtitles. The video image can be downgraded to a low bitrate video stream, static cover image, or background animation based on the bandwidth level.

[0052] S103. When in lead vocal mode, the collected local microphone voice is sent to other vehicle terminals in the fleet through the first low-latency wireless audio link, and the local microphone voice is sent to the local audio processing module through the local audio input path.

[0053] In this context, the local microphone voice refers to the vocal signal of the lead singer, which is captured in real time by the vehicle's terminal audio input device (such as a USB microphone, in-vehicle microphone, or 3.5mm jack microphone). This signal is then processed through analog-to-digital conversion (ADC), automatic gain control (AGC), and noise suppression (NS) to form a digital audio stream. The first low-latency wireless audio link serves as the transmission channel for the vocal data in this step. The lead singer's vehicle encapsulates the processed vocal frames into data packets containing timestamps, sequence numbers, and CRC checksums, and broadcasts them synchronously to all receiving vehicles in the convoy via the UHF transmission module. The local audio input path refers to the complete data flow link from the microphone pickup hole through the hardware acquisition circuit, drive buffer, operating system audio subsystem, application layer receiving queue, and finally into the audio processing module. Its latency is determined by the fixed hardware delay, drive buffer depth, and system scheduling overhead.

[0054] When the vehicle terminal is in lead vocal mode, the system needs to handle two parallel paths simultaneously: external broadcasting of vocals and local monitoring. Specifically, the audio acquisition thread reads PCM data from the microphone device at a fixed sampling rate (e.g., 48kHz) and frame length (e.g., 20ms). After preprocessing, the audio frame is encapsulated into a wireless data packet in the transmission branch: a transmission timestamp generated based on a unified reference time (with an accuracy of no less than 1ms), an incrementing frame sequence number, and a CRC-16 checksum are written into the packet header. Then, it is broadcast to other receiving vehicles in the convoy through the configured UHF transmission module. In the local monitoring branch, a copy of the same audio frame is directly sent to the vocal input port of the local audio processing module without going through the wireless transmission link, thus ensuring that the lead vocal vehicle itself can hear the mixing effect of vocals and accompaniment in real time. To prevent feedback, the local monitoring path dynamically adjusts the monitor gain or enables the adaptive echo cancellation (AEC) algorithm according to the speaker layout and the in-vehicle acoustic environment, so that the lead vocalist will not generate positive feedback feedback when the vocals output from the speakers are picked up by the microphone again while monitoring in the vehicle.

[0055] In some preferred embodiments, the lead vocal vehicle can simultaneously connect to a main microphone and a secondary microphone. The system first performs gain equalization and clock alignment locally on the vocal signals input from multiple microphones, then synthesizes them into a unified vocal transmission stream, which is broadcast to other vehicle terminals via a first low-latency wireless audio link. Simultaneously, the vocal component corresponding to the main microphone is retained as the local main monitoring component and sent to the local audio processing module. For in-vehicle monitor scenarios, the system can also automatically adjust the monitor gain, reverberation intensity, and echo cancellation coefficient according to the relative position between the speaker and the microphone to reduce the risk of feedback when multiple microphones are input and improve the clarity of the lead vocalist's hearing.

[0056] S104. When in receiving mode, receive the human voice sent by the current lead singer vehicle terminal according to the first low-latency wireless audio link parameters, and send the received human voice to the local audio processing module.

[0057] The first low-latency wireless audio link parameter, in receive mode, guides the UHF receiving module to lock onto the transmitter frequency, device address, and decoding configuration of the lead singer vehicle, ensuring that the vehicle can accurately capture the over-the-air vocal data packets and filter out non-target signals. The current lead singer vehicle terminal refers to the specific vehicle designated as the vocal source by the cloud server during the current song's performance. Its device identifier and operating frequency have been explicitly distributed to all receiving vehicles in the fleet through the role information. The received vocal data includes the audio frame payload, transmission timestamp, frame sequence number, and verification information. After unpacking, the receiving vehicle performs CRC verification, frame loss detection, and timing reordering before sending the decoded PCM audio stream to the vocal input port of the local audio processing module.

[0058] When the vehicle terminal is in receive mode, the local microphone acquisition function is no longer activated. Instead, the UHF receiving module is configured to listen and continuously scan for wireless data packets on the specified frequency. Specifically, the receiving module receives only data packets from the current lead vocal vehicle based on the issued device address filtering table, and performs hardware-level filtering on signals from other vehicles or interference sources. Each time a complete data packet is received, the driver layer immediately extracts the CRC checksum from the packet header for verification. If the verification passes, the audio frame payload and timestamp information are uploaded to the application layer's receive queue. If the verification fails, the packet is discarded and a frame loss event is recorded. The application layer's receive thread retrieves audio frames from the queue and checks for packet loss or out-of-order delivery based on the frame sequence number. For lost frames, a strategy of repeating the previous frame or using silence padding is employed for error correction. For out-of-order frames, they are rearranged according to the sequence number before being sent to the decoder. The decoded PCM vocal data is written to the vocal input buffer of the local audio processing module, awaiting mixing with the locally played accompaniment audio. This pipelined processing method of receiving, verifying, decoding, and sending ensures data integrity while keeping the additional processing latency on the receiving side within 5ms. Combined with the air interface transmission latency (usually less than 10ms), the total delay for the receiving vehicle to hear the lead singer's voice can be stabilized at around 15ms, meeting the experience requirements for real-time chorus.

[0059] In some preferred embodiments, different vehicle terminals in receiving mode do not necessarily output the exact same mixing result. The system can differentiate the gain ratio of vocals and accompaniment based on vehicle location, speaker type, or application scenario. For example, vehicles closer to the lead singer's vehicle and serving as live audience playback devices can appropriately reduce the volume of vocals and enhance the atmosphere of the accompaniment, while vehicles farther from the lead singer's vehicle or serving as main chorus monitoring devices can increase vocal clarity and mid-frequency energy, thereby creating local mixing versions suitable for different spatial locations.

[0060] S105. Determine the corresponding vocal path delay parameter according to the current mode. In lead vocal mode, the vocal path delay parameter is the local vocal input path delay, and in receiving mode, it is the vocal reception delay of the first low-latency wireless audio link.

[0061] The vocal path delay parameter refers to the end-to-end time overhead from the generation of the vocal signal to its entry into the local audio processing module. This parameter uses different measurement benchmarks depending on the role of the vehicle. The local vocal input path delay refers to the inherent delay in the process of the microphone picking up the sound signal, which is then processed through hardware acquisition, drive buffering, and system scheduling before finally reaching the input of the audio processing module in lead vocal mode. This value is mainly determined by the audio interface type (USB latency approximately 5-10ms, 3.5mm analog interface latency approximately 2-5ms), drive buffer depth (usually configured for 2-4 audio frames), and operating system scheduling granularity, typically ranging from 10-30ms. The vocal reception delay refers to the total time taken in lead vocal vehicle to send vocal data packets to the local vehicle for reception and delivery to the audio processing module in receive mode. This includes air interface transmission delay, receiver hardware processing delay, verification and decoding overhead, and software queue buffering time. This value is dynamically affected by factors such as vehicle distance, electromagnetic environment, and received signal strength.

[0062] Before starting playback of the current song, the system needs to determine the latency reference value for accompaniment output timing compensation based on the vehicle's current role. Specifically, if the vehicle is in lead vocal mode, the system reads the pre-determined local vocal input path latency value from the local configuration file or device calibration file. This value is usually obtained and persistently stored through an automatic calibration process (e.g., playing test audio with known latency and measuring the round-trip latency through the microphone, taking the one-way value) or manual calibration after the vehicle's first startup or after an audio device change.

[0063] If the vehicle is in receiving mode, a real-time latency measurement mechanism is activated. This involves parsing the transmission timestamp carried in the received voice data packet, combining it with the vehicle's local clock synchronized with the reference time information to obtain the reception timestamp, and calculating the difference between the two as the current voice reception latency. Since both the vehicle terminal and the lead singer's vehicle terminal have achieved high-precision time synchronization with the cloud server via a second wireless communication link, their local clocks are on the same time base, thus ensuring the accuracy of the transmission latency calculated based on the timestamp.

[0064] The reason for using latency parameters based on role differentiation is that the vocals in the lead singer's vehicle can enter the local mixing process without wireless transmission, and the key bottleneck lies in the local hardware acquisition link; while the vocals in the receiving vehicle must be transmitted over the air interface, and the key bottleneck lies in the propagation and processing latency of the wireless link. By selecting corresponding measurement methods for the actual physical characteristics of the two types of paths, the system can more accurately quantify the actual arrival time of the vocal data stream, providing a reliable alignment benchmark for subsequent accompaniment output compensation.

[0065] In some preferred embodiments, the vocal path delay parameter can be further decomposed into at least two components. The local vocal input path delay includes at least audio acquisition hardware delay, driver buffer delay, and application layer scheduling delay. The vocal reception delay in receiving mode includes at least air interface transmission delay, receiving-side unpacking and decoding delay, and local playback scheduling delay. The system can maintain local delay calibration files for different vehicle models, different USB audio transceiver models, or different audio driver versions. When entering vocal mode or receiving mode, it prioritizes reading the empirical initial values ​​from the corresponding files and makes corrections based on real-time measurement results to shorten the initial compensation convergence time.

[0066] In some preferred embodiments, when the vehicle terminal is in receiving mode, the corresponding voice path delay parameter is determined according to the current mode. Specifically, this includes: parsing the transmission timestamp from the voice data received through the first low-latency wireless audio link, the transmission timestamp being recorded by the current lead vehicle terminal based on reference time information when transmitting the voice data; obtaining the reception timestamp of the received voice data based on the local clock of the vehicle terminal synchronized with the reference time information; calculating the time difference between the reception timestamp and the transmission timestamp, and determining the time difference as the voice reception delay.

[0067] In receive mode, latency measurement is triggered immediately after the first low-latency wireless audio link is established and the first valid audio packet is received. Specifically, the system parses T_send from the CRC-checked data packet and simultaneously reads the local synchronization clock to obtain T_recv, calculating Δt = T_recv − T_send as the link latency sample value. Since both ends' clocks are synchronized to the same reference time, the difference directly reflects the actual propagation and processing delay, without the need for additional clock deviation correction. To suppress occasional air interface jitter, the median or weighted average of the differences over several consecutive frames (e.g., 5 frames) can be used as the final initial reference, while obviously abnormal outliers are removed.

[0068] S106. Time synchronization is performed with the cloud server via the second wireless communication link, and the output timing of the accompaniment audio is compensated based on the reference time information and the human voice path delay parameters.

[0069] Time synchronization refers to the vehicle terminal periodically calibrating its clock with the cloud server via a second wireless communication link. This is achieved using a variant of the Simplified Network Time Protocol (SNTP) or the Precise Time Protocol (PTP) to ensure continuous alignment between the local clock and the unified reference time in the cloud. The synchronization period is typically set to 10-60 seconds, with a clock deviation tolerance of ±5ms. The output time of the accompaniment audio refers to the precise point in time when the local audio processing module reads the audio frame from the accompaniment decoding buffer and sends it to the mixer for synthesis. This moment directly determines the alignment of the accompaniment and vocals on the timeline. The compensation logic can be implemented by adjusting the read pointer offset of the decoding buffer, modifying the audio timestamp, or inserting / deleting silence samples.

[0070] After determining the vocal path delay parameters, the system needs to calculate the output compensation amount of the accompaniment audio to achieve audiovisual synchronization. Specifically, the vehicle terminal sends a time synchronization request to the cloud server through the second wireless communication link, receives the high-precision reference timestamp T_cloud returned by the cloud and the local time T_local when the timestamp is received, calculates the one-way network delay (T_local-T_cloud) / 2 and calibrates the local clock accordingly to keep the deviation between the local clock and the cloud reference time within the tolerance range. Subsequently, based on the unified start time T_start, the current time T_now, and the vocal path delay parameter Δt_voice, the target time when the accompaniment should start output is calculated as T_output=T_start-Δt_voice, that is, the accompaniment needs to start playing Δt_voice time in advance to ensure that when the vocal signal reaches the mixer after acquisition or transmission, the accompaniment has progressed to the corresponding musical measure.

[0071] In vocal mode, Δt_voice represents the local input path latency (e.g., 15ms), and the compensated accompaniment will be output 15ms earlier. In receive mode, Δt_voice represents the wireless receive latency (e.g., 12ms), and the compensated accompaniment will be output 12ms earlier. To ensure a smooth compensation process, the system employs a gradual strategy each time the compensation amount is updated: if the difference between the old and new compensation amounts is less than 5ms, a seamless switch is achieved by fine-tuning the decoder buffer read pointer; if the difference exceeds 5ms, a jump adjustment is performed at the next audio frame boundary, triggering a short fade-in / fade-out process to avoid popping sounds.

[0072] In some preferred embodiments, when the accompaniment resource comes from the local cache, the vehicle terminal does not need to recalculate the absolute start time. Instead, it directly uses the local sample position corresponding to the unified start anchor point issued by the cloud server as a benchmark, and achieves compensation by adjusting the reading start point and continuous reading step of the accompaniment buffer. Furthermore, when the vehicle is in in-vehicle monitoring mode, in-vehicle amplification mode, or in-vehicle and in-vehicle simultaneous output mode, the system can configure compensation coefficient correction values ​​for different output layouts. Among them, since the in-vehicle amplification path may pass through the power amplifier and external speakers, an additional output link correction amount can be superimposed on the vocal path delay parameter to improve the auditory synchronization between different playback channels.

[0073] S107. The human voice and the compensated accompaniment audio entering the local audio processing module are mixed locally on the vehicle terminal to obtain synchronized playback audio and output it.

[0074] After both vocal and accompaniment data enter the local audio processing module, the system initiates a real-time mixing process to generate the final playback audio. Specifically, the mixer synchronously reads audio frames from the vocal and accompaniment input buffers according to a unified sampling clock (e.g., one frame every 20ms), and performs a weighted summation operation on the two signals based on sample points.

[0075] The mixer calculates the output value for each sample point based on the user-preset vocal volume V_vocal and accompaniment volume V_music: Output[i] = V_vocal × Vocal[i] + V_music × Music[i], where Vocal[i] and Music[i] are the amplitude values ​​of the vocals and accompaniment at the i-th sample point, respectively. To prevent signal overflow or clipping distortion after mixing, the mixer performs dynamic range detection before superposition. If the predicted output peak value may exceed the maximum representation range of digital audio (such as ±32767 for 16-bit PCM), it automatically reduces the gain of the two signals or enables the limiter for peak clipping protection.

[0076] After mixing, the output audio stream passes through optional sound effect processing units (such as adding KTV reverb effects, bass boost, or virtual surround sound depending on the song type) and final output gain adjustment. It is then sent to the audio driver layer to be written to the hardware playback buffer, converted into an analog signal by the vehicle audio system's DAC (digital-to-analog converter), and used to drive the speakers. Depending on the vehicle configuration, the output path can be divided into in-vehicle monitoring mode (output to in-vehicle speakers for passengers), external amplification mode (output to external amplifiers and outdoor speakers for creating a camping atmosphere), or a mixed mode (output to both in-vehicle and external devices simultaneously).

[0077] In some preferred embodiments, the local audio processing module can simultaneously generate an in-vehicle monitoring mix and an external amplification mix. The in-vehicle monitoring mix prioritizes the clarity of the lead vocals, employing higher vocal gain and lower reverberation intensity. The external amplification mix prioritizes the chorus atmosphere and coverage, appropriately increasing the accompaniment soundstage width and environmental reverberation ratio. Both mixes can share the same vocal path delay parameters and a unified start-up baseline, but their final gain allocation, sound effect parameters, and output device selection are controlled independently.

[0078] In some preferred embodiments, during the process of mixing the human voice and the compensated accompaniment audio entering the local audio processing module locally on the vehicle terminal, the method further includes:

[0079] The steps of parsing the sending timestamp, obtaining the receiving timestamp, and calculating the time difference are repeated periodically to obtain the updated human voice reception delay;

[0080] Compare the updated voice reception delay with the current voice path delay parameter;

[0081] When the difference between the updated voice reception delay and the current voice path delay parameter exceeds a preset threshold, the voice path delay parameter is updated to the updated voice reception delay, and the compensation amount for the output time of the accompaniment audio is re-determined based on the updated voice path delay parameter.

[0082] Among them, the preset threshold refers to the difference boundary for judging whether the delay has drifted significantly. It can be configured with a default fixed value (such as 3ms to 5ms) or dynamically adjusted according to the current vehicle speed or network fluctuation level; the compensation amount refers to the sample point offset applied to the accompaniment buffer read pointer after conversion based on the delay parameter.

[0083] This mechanism operates while the vehicle terminal is in receive mode and mixing is continuously outputting. The detection period is typically set to 500ms to 1000ms to balance real-time response capability and processing resource consumption. Specifically, the system periodically resamples the current link delay using the difference between send and receive timestamps and compares it with the already effective vocal path delay parameters. If the difference does not exceed the threshold, the current accompaniment read pointer offset remains unchanged to prevent audio popping caused by frequent pointer adjustments due to instantaneous network jitter. If the difference exceeds the threshold, the new measurement value is updated to the current vocal path delay parameters, and a gradual strategy is used to fine-tune the read pointer to ensure a smooth transition of compensation and avoid audio distortion caused by abrupt changes.

[0084] In some preferred embodiments, the preset threshold, detection period, and compensation update strategy can be dynamically switched according to the song type. For songs with a fast tempo, dense drum beats, or a high proportion of rap, the system uses a shorter detection period and a smaller threshold to improve transient beat alignment; for ballads or songs with many long notes, the system uses a longer detection period and a larger threshold, and increases the weight of the compensation smooth transition to reduce subtle auditory jitter caused by frequent compensation. The song type can be provided directly from the song request resource metadata, or it can be carried synchronously by the cloud server when distributing song information.

[0085] The above embodiments describe dual-link latency compensation based on role differentiation. In real-world scenarios, users within a convoy often need to take turns singing. To ensure the continuity of audio-visual synchronization across vehicles during lead singer handover, this application also provides, for example... Figure 2 The method embodiment shown illustrates the process of dynamic role switching and compensation benchmark reconstruction.

[0086] The following describes the process of the method provided in this implementation. Please refer to [link / reference]. Figure 2 This is a flowchart illustrating a multi-vehicle collaborative karaoke method in an embodiment of this application.

[0087] S201. Receive the role switching instruction sent by the cloud server. The role switching instruction includes at least the next lead singer vehicle terminal identifier, the target working frequency corresponding to the next lead singer vehicle terminal identifier, the transmit / receive mode, and the switching effective time.

[0088] Among them, the role switching command refers to the role reconfiguration control command uniformly issued by the cloud server to all vehicle terminals in the fleet through the second wireless communication link based on the progress of the song request queue or user operation; the target working frequency refers to the UHF frequency point that the new lead singer vehicle will use on the first low latency wireless audio link, and the receiving mode vehicles in the fleet will lock the receiving frequency accordingly.

[0089] Role switching is typically driven by events such as the current song finishing, the song request queue advancing, or a user actively triggering the end of the current lead singer's role. The cloud server issues a command in advance before the switch takes effect, allowing each terminal time for parameter parsing and resource preparation. Specifically, upon receiving the command, each vehicle terminal immediately parses the fields, extracts the next lead singer's identifier, target operating frequency, transmit / receive mode, and switch time, registers the switch trigger event in a local timer, and then can autonomously execute the switch without waiting for any additional network notifications.

[0090] S202. When the handover takes effect, the vehicle terminal is reconfigured into lead mode or receive mode according to the next lead vehicle terminal identifier and transmit / receive mode, and the parameters of the first low-latency wireless audio link are updated.

[0091] Among them, the switching effective time refers to the configuration flip trigger point registered by the local timer of each vehicle terminal based on the clock synchronized with the reference time information; the transmit / receive mode update refers to reconfiguring the working direction of the vehicle's UHF transmit / receive module according to the new role, with the lead role corresponding to the transmit mode and the receiving role corresponding to the receive / listen mode.

[0092] When the switch takes effect, the local timer triggers an atomic configuration flip process without waiting for any external instructions. Specifically, the state machine switches to the mode corresponding to the new role: if this vehicle is designated as the new lead singer, the local microphone acquisition channel is activated, and the UHF module is switched to the transmit-ready state, while a transmission timestamp based on the latest reference time is written into the header of the transmitted data packet; if this vehicle switches to the receiver role, the local microphone input is turned off, the UHF receiver module is re-locked to the target operating frequency and device address of the new lead singer's vehicle, and listening to the vocal data packets emitted by the new lead singer begins.

[0093] S203. Redetermine the voice path delay parameters corresponding to the updated mode;

[0094] Among them, the vocal path delay parameter corresponding to the updated mode refers to the accompaniment output compensation reference delay value re-determined based on the new role of the vehicle terminal after the role switch is completed. When the updated mode is for the lead singer, the local vocal input path delay is taken, and when the updated mode is for receiving, the measured vocal reception delay of the first low-latency wireless audio link is taken. The physical source, measurement method and numerical range of the two are fundamentally different and cannot be used interchangeably.

[0095] Specifically, if the update is to lead vocal mode, the pre-stored audio input path delay value is read from the local device calibration file and used directly as the new compensation benchmark. If the update is to receive mode, a real-time measurement process is initiated, parsing the transmission timestamp in the first batch of audio data packets sent by the new lead vocal vehicle, calculating the reception delay in conjunction with the local synchronization clock, and completing the calibration using the median of the initial few frames. The fundamental reason for redetermining the delay parameters is that the physical propagation path of the vocal data stream has fundamentally changed after the role is reversed—the measurement path corresponding to the original benchmark value no longer exists. Using the old benchmark will introduce systematic compensation errors, causing a continuous offset between the vocals and accompaniment on the timeline. Therefore, recalibration based on the new path is necessary.

[0096] S204. When the updated mode is the lead vocal mode, the local vocal input path latency will be used as the compensation benchmark.

[0097] Among them, the compensation benchmark refers to the time delay reference value calculated by the amount of time the accompaniment is output in advance. Its accuracy affects the time alignment accuracy between the vocals and the accompaniment at the mixing output.

[0098] Specifically, the local latency calibration value Δt_local corresponding to the current audio input device type is read from the device calibration file. The accompaniment audio should start outputting Δt_local before the unified start time T_start, so that the arrival time of the accompaniment signal at the mixer is precisely aligned with the arrival time of the vocals captured by the local microphone. The local acquisition latency is used as the compensation benchmark rather than the wireless link latency because the vocals in the vocal vehicle flow entirely within the local acquisition link and do not undergo any over-the-air propagation. The source of its latency is unrelated to the wireless transmission path at the receiving end. If the wireless link latency is mistakenly used as the compensation benchmark, too much or too little advance will be applied to the accompaniment, resulting in a continuously perceptible misalignment of the vocals and accompaniment when the vocal vehicle is monitoring itself.

[0099] S205. When the updated mode is the receiving mode, the voice reception delay of the first low-latency wireless audio link is used as the compensation benchmark.

[0100] Specifically, the sending timestamp T_send is parsed from the audio data packets sent by the new lead singer's vehicle. Combined with the receiving time T_recv recorded by the local synchronization clock, Δt_recv = T_recv − T_send is calculated. This value is set as the compensation amount for the current accompaniment output time, causing the accompaniment to be output Δt_recv ahead of time. This ensures that the vocals arriving locally via the air interface are precisely aligned on the timeline with the accompaniment signal in the mixer. The core principle of using the air interface reception delay as a benchmark is that the vocals heard by the receiving vehicle will inevitably be later than the actual vocals. If an equal amount of advance compensation is not applied to the accompaniment, the accompaniment beat will lead the vocals in the local mix.

[0101] S206. Based on the redefined vocal path delay parameters, the output time of the accompaniment audio for the next song is recompensated.

[0102] Specifically, the system calculates the target time when the accompaniment should actually start output, T_output = T_start_next − Δt_new, based on the updated vocal path delay parameter Δt_new (local input path delay in lead vocal mode and measured wireless link reception delay in receive mode), combined with the unified start time T_start_next of the next song sent from the cloud. That is, the accompaniment needs to start playing Δt_new time in advance.

[0103] If the difference between the old compensation amount and the new compensation amount is small (e.g., less than 5ms), a smooth transition is achieved by fine-tuning the read pointer offset of the decoding buffer; if the difference is large, the new compensation amount is applied directly when the next song starts playing, and a short fade-in process is performed at the moment of start-up to avoid sudden popping sounds.

[0104] In some preferred embodiments, the triggering conditions for role switching, in addition to the natural end of the current song, also include: the user actively ending the current lead singer, the team leader's vehicle forcibly inserting a song, the current lead singer's vehicle reporting that it is quitting the performance, and the cloud server triggering a reselection of the lead singer due to a link abnormality of the current lead singer's vehicle.

[0105] For the different trigger sources mentioned above, the cloud server prioritizes generating a role switching instruction that includes the next lead vehicle terminal identifier, target operating frequency, and the effective switching time, and reserves at least one parameter preparation window before the effective switching time. Within the parameter preparation window, each vehicle terminal completes target frequency pre-locking, clearing the tail frame of old data packet reception, and preparing local resources corresponding to the new role in advance to reduce the probability of loss of synchronization caused by old frame residue or new frame not being ready at the moment of switching.

[0106] Furthermore, in some preferred embodiments, the system also supports multi-vocalist extension modes, including dual-frequency dual-vocalist mode, time-division round singing mode, or main and secondary voice parallel transmission mode.

[0107] In some preferred embodiments, while receiving the accompaniment audio corresponding to the current song via a second wireless communication link, the method further includes:

[0108] Receive lyrics and / or video footage corresponding to the current song via a second wireless communication link;

[0109] Based on reference time information and human voice path delay parameters, the display time of lyrics subtitles and / or video images is compensated;

[0110] While outputting synchronized audio playback, it simultaneously displays compensated lyrics subtitles and / or video footage.

[0111] Among them, lyrics subtitles refer to the word-by-word or sentence-by-sentence singing prompts marked with timestamps, usually stored in LRC or enhanced JSON format, and each subtitle fragment is associated with a display timestamp relative to the start time of the song; video footage refers to the MV (Music Video) video stream corresponding to the current song, which is transmitted to the local machine in segments through a second wireless communication link in H.264 / H.265 and other codec formats, and each video frame also carries a presentation timestamp based on the song's timeline; display time compensation refers to applying the same time offset as the accompaniment audio to the above timestamps or frame presentation times, so that the progression rhythm of visual elements is strictly locked in time with the time-delay compensated accompaniment.

[0112] This step is executed in parallel with the reception of the accompaniment audio, completing the parsing of the subtitle file and video pre-caching before the song begins. Specifically, after receiving the lyrics subtitle file, the system iterates through all time stamps, adds an compensation amount Δt (i.e., the vocal path delay parameter corresponding to the current character) to each original display time t_lyric, generates the adjusted display time t_lyric_adjusted=t_lyric+Δt, and writes it to the local subtitle rendering queue; for the video frame, after reading the PTS (Presentation Time Stamp) of each frame, the decoder also applies the same compensation amount to ensure that the frame presentation time is consistent with the accompaniment output time by a certain advance.

[0113] In some preferred embodiments, when the available bandwidth of the second wireless communication link is lower than the video continuous playback threshold, the system prioritizes the normal reception and display of the accompaniment audio and lyrics subtitles, and downgrades the video image to low bitrate video, keyframe sampling playback, static cover image or pure background animation in sequence, so as to prioritize maintaining auditory synchronization and basic visual cues during the singing process.

[0114] Furthermore, the lead singer vehicle terminal and the receiving mode vehicle terminal can display prompt interfaces with different granularities. The lead singer vehicle terminal prioritizes displaying word-by-word rhythm prompts, real-time scoring, and a countdown to the start of singing, while the receiving mode vehicle terminal prioritizes displaying large-print lyrics, chorus prompts, or simplified rhythm guidance.

[0115] In some preferred embodiments, during the process of mixing the vocals and compensated accompaniment audio entering the local audio processing module locally on the vehicle terminal, the method further includes: monitoring the connection status of the first low-latency wireless audio link; when the first low-latency wireless audio link is detected to be interrupted or the signal quality is lower than a preset threshold, sending an abnormal notification to the cloud server through the second wireless communication link; receiving a processing instruction issued by the cloud server in response to the abnormal notification, the processing instruction being used to instruct the vehicle terminal to perform at least one of the following operations: switching to a backup operating frequency, triggering a reselection of the lead vocal vehicle terminal, or downgrading to a accompaniment-only playback mode.

[0116] Among them, the connection status refers to the real-time communication quality indicators of the first low-latency wireless audio link, including quantifiable parameters such as Received Signal Strength Indicator (RSSI), packet loss rate, consecutive frame loss count, and CRC check failure rate; the preset threshold refers to the critical conditions that trigger abnormal reporting, such as RSSI being lower than -85dBm for more than 2 seconds, packet loss rate exceeding 10%, or 5 consecutive frame check failures.

[0117] This monitoring mechanism operates continuously during mixing output, with a detection cycle typically set between 100ms and 500ms. Specifically, the UHF receiver module driver layer periodically reports the current RSSI value, recent packet loss statistics, and checksum error count to the application layer. The application layer compares these indicators with preset thresholds one by one: if a link interruption is detected (e.g., no valid data packets are received for one second) or signal quality deterioration is detected (e.g., RSSI drops below the threshold and packet loss rate increases sharply), an abnormal notification is immediately sent to the cloud via the second wireless communication link, carrying the vehicle identifier, the current lead vocalist identifier, the operating frequency, and the fault type. After receiving the notification, the cloud server makes a decision based on the simultaneous reporting from other vehicles in the fleet: if only a few vehicles are affected and the backup frequency is available, a frequency switching command is issued; if the current lead vocalist's vehicle has a link abnormality or most receiving vehicles report faults simultaneously, a lead vocalist reselection process is triggered, selecting a vehicle with a better RF environment from the fleet as the new lead vocalist; if the overall link quality is insufficient to support vocal transmission, a degradation command is issued, and each vehicle switches to accompaniment-only mode, suspending vocal broadcasting and reception.

[0118] In some preferred embodiments, a local first-level fallback is triggered when the vehicle terminal first detects an anomaly in the first low-latency wireless audio link. The local first-level fallback includes at least: freezing the current accompaniment compensation amount, performing a short-term hold or mute fill on the most recent effective vocal frame, displaying a "vocalist link anomaly" prompt on the user interface, and waiting for the cloud server to return processing instructions without immediately interrupting the accompaniment playback.

[0119] When the link quality is restored and the recovery threshold is continuously met for a preset duration, the vehicle terminal executes a switchback process. The switchback process includes at least the following: reacquiring the current lead vocal vehicle terminal identifier and effective operating frequency, restoring normal voice reception and latency measurement, performing smooth regression on the compensation amount, and canceling the abnormal prompt.

[0120] In some preferred embodiments, before publishing the unified start time of the current song, the cloud server first confirms that each vehicle meets the readiness conditions. The readiness conditions include at least: the accompaniment cache of the current song reaches a first threshold, the lyrics resource is loaded, the first low-latency wireless audio link parameters have been successfully configured, the clock deviation between each vehicle and the reference time information is less than a preset tolerance, and the input path calibration value of the current lead singer vehicle is available.

[0121] In some preferred embodiments, each vehicle completes device authentication and session binding when joining a room. The cloud server assigns a unique vehicle identifier to each vehicle within the room and associates the vehicle identifier with a local low-latency wireless audio device identifier to reduce the risk of nearby non-vehicle equipment mistakenly accessing the same frequency.

[0122] In some preferred embodiments, when the second wireless communication link experiences a short-term anomaly but the local cache already contains sufficient accompaniment and lyrics, each vehicle can continue playing the current song according to the most recent valid unified reference time information, thereby achieving scenario adaptation in weak networks.

[0123] In some preferred embodiments, before a vehicle joins a karaoke room, the team leader vehicle or cloud server first completes room creation, vehicle identity authentication, device binding, and frequency legality verification. The frequency legality verification is used to filter the allocable target operating frequencies based on the current regional regulatory rules, the available frequency pool, and the occupancy of nearby interference.

[0124] In some preferred embodiments, the cloud server initiates a readiness voting process before announcing the unified start time. Each vehicle terminal sequentially reports the status of the accompaniment cache, lyrics resources, video resources, the configuration status of the first low-latency wireless audio link, the clock synchronization status, and the vocalist's input path calibration status. The unified start time is only issued when the preset readiness ratio is reached or all users are ready.

[0125] Furthermore, in some preferred embodiments, nearby vehicles can automatically discover the fleet to be joined based on mechanisms such as room code, geofencing, or Bluetooth proximity discovery, and join the room after confirmation by the team leader vehicle; when the fleet size is large, the cloud server can also divide the fleet into multiple synchronization groups to reduce the synchronization and scheduling complexity in large-scale fleets.

[0126] The multi-vehicle collaborative karaoke device in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference needed]. Figure 3 This is a schematic diagram of the physical device structure of a multi-vehicle collaborative karaoke device in the embodiments of this application.

[0127] It should be noted that, Figure 3 The structure of the multi-vehicle collaborative karaoke device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0128] like Figure 3As shown, the multi-vehicle collaborative karaoke device includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage section 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0129] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0130] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0131] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0133] Specifically, the multi-vehicle collaborative karaoke device in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the multi-vehicle collaborative karaoke method provided in the above embodiment.

[0134] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the multi-vehicle collaborative karaoke device described in the above embodiments; or it may exist independently and not assembled into the multi-vehicle collaborative karaoke device. The storage medium carries one or more computer programs, which, when executed by a processor of the multi-vehicle collaborative karaoke device, cause the multi-vehicle collaborative karaoke device to implement the multi-vehicle collaborative karaoke method provided in the above embodiments.

[0135] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0136] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0137] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A multi-vehicle collaborative karaoke method, applied to a multi-vehicle collaborative karaoke system, the system comprising a cloud server and multiple vehicle terminals, characterized in that, The method includes: Receive current role information, first low-latency wireless audio link parameters and reference time information sent by the cloud server, wherein the first low-latency wireless audio link parameters are used to establish the first low-latency wireless audio link; The vehicle terminal receives the accompaniment audio corresponding to the current song through the second wireless communication link, and configures the vehicle terminal into lead singer mode or receiving mode according to the current role information. The first low-latency wireless audio link is a local area communication link within the fleet, and the second wireless communication link is a wide area network communication link. When in the lead vocal mode, the collected local microphone voice is sent to other vehicle terminals in the fleet through the first low-latency wireless audio link, and the local microphone voice is sent to the local audio processing module through the local audio input path. When in the receiving mode, the voice sent by the current lead singer vehicle terminal is received according to the first low-latency wireless audio link parameters, and the received voice is sent to the local audio processing module. The corresponding vocal path delay parameter is determined according to the current mode. The vocal path delay parameter is the local vocal input path delay in the lead vocal mode and the vocal reception delay of the first low-latency wireless audio link in the receiving mode. The second wireless communication link is used to synchronize time with the cloud server, and the output time of the accompaniment audio is compensated based on the reference time information and the human voice path delay parameter. The human voice and the compensated accompaniment audio entering the local audio processing module are mixed locally on the vehicle terminal to obtain synchronized playback audio and output it.

2. The method according to claim 1, characterized in that, After mixing the vocals and compensated accompaniment audio that have entered the local audio processing module locally with the vehicle terminal to obtain synchronized playback audio and then outputting it, the method further includes: Receive a role switching instruction sent by the cloud server. The role switching instruction includes at least the next lead singer vehicle terminal identifier, the target operating frequency corresponding to the next lead singer vehicle terminal identifier, the transmit / receive mode, and the switching effective time. When the switching takes effect, the vehicle terminal is reconfigured to the lead vocal mode or the receive mode according to the next lead vocal vehicle terminal identifier and the transmit / receive mode, and the first low-latency wireless audio link parameters are updated.

3. The method according to claim 2, characterized in that, After reconfiguring the vehicle terminal to the lead vocal mode or the receive mode according to the next lead vocal vehicle terminal identifier and the transmit / receive mode, and updating the first low-latency wireless audio link parameters, the method further includes: The voice path delay parameters corresponding to the updated mode are redefined; When the updated mode is the lead vocal mode, the local vocal input path delay is used as the compensation benchmark. When the updated mode is the receiving mode, the human voice reception delay of the first low-latency wireless audio link is used as the compensation benchmark. Based on the redefined vocal path delay parameters, the output timing of the accompaniment audio for the next song is recompensated.

4. The method according to claim 1, characterized in that, When the vehicle terminal is in the receiving mode, determining the corresponding voice path delay parameter based on the current mode specifically includes: The transmission timestamp is parsed from the human voice data received through the first low-latency wireless audio link. The transmission timestamp is recorded by the current lead singer vehicle terminal based on the reference time information when transmitting the human voice data. Based on the local clock of the vehicle terminal synchronized with the reference time information, the receiving timestamp of the received human voice data is obtained. Calculate the time difference between the received timestamp and the sent timestamp, and determine the time difference as the human voice reception delay.

5. The method according to claim 4, characterized in that, The method further includes the following steps during the process of mixing the vocals and the compensated accompaniment audio entering the local audio processing module locally on the vehicle terminal: The steps of parsing the sending timestamp, obtaining the receiving timestamp, and calculating the time difference are repeated periodically to obtain the updated human voice reception delay; Compare the updated human voice reception delay with the current human voice path delay parameter; When the difference between the updated human voice reception delay and the current human voice path delay parameter exceeds a preset threshold, the human voice path delay parameter is updated to the updated human voice reception delay, and the compensation amount for the output time of the accompaniment audio is re-determined based on the updated human voice path delay parameter.

6. The method according to claim 1, characterized in that, While receiving the accompaniment audio corresponding to the current song via the second wireless communication link, the method further includes: Receive lyrics subtitles and / or video footage corresponding to the current song through the second wireless communication link; Based on the reference time information and the human voice path delay parameters, the display time of the lyrics subtitles and / or the video frame is compensated; While outputting the synchronized audio, the compensated lyrics subtitles and / or the video footage are displayed simultaneously.

7. The method according to claim 1, characterized in that, The method further includes the following steps during the process of mixing the vocals and the compensated accompaniment audio entering the local audio processing module locally on the vehicle terminal: Monitor the connection status of the first low-latency wireless audio link; When the first low-latency wireless audio link is detected to be interrupted or the signal quality is lower than a preset threshold, an abnormal notification is sent to the cloud server through the second wireless communication link. The cloud server receives a processing instruction in response to the abnormal notification. The processing instruction is used to instruct the vehicle terminal to perform at least one of the following operations: switch to a backup operating frequency, trigger a reselection of the lead vocal vehicle terminal, or downgrade to a accompaniment-only playback mode.

8. A multi-vehicle collaborative karaoke device, characterized in that, The multi-vehicle collaborative karaoke device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the multi-vehicle collaborative karaoke device to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the multi-vehicle collaborative karaoke device, the multi-vehicle collaborative karaoke device performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on a multi-vehicle collaborative karaoke device, the multi-vehicle collaborative karaoke device performs the method as described in any one of claims 1-7.