Multi-passenger seat area-oriented in-vehicle KTV partition sound pickup and mixing method

By constructing a fixed sound path and a virtual microphone set in the vehicle, the problem of microphone dependence in the in-vehicle KTV system was solved, and clear zoned sound pickup and synchronous playback in multiple seating areas were achieved, improving the interactivity and immersion of the in-vehicle multi-occupant KTV.

CN120766640BActive Publication Date: 2025-11-28SHANGHAI DAYIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261672.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-28
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing in-vehicle karaoke systems rely on microphones, which suffer from inconvenience, space occupation, echo interference, and insufficient interaction among multiple passengers. In particular, it is difficult to achieve clear zoned sound pickup and synchronized playback when multiple seats are singing together at the same time.

Method used

By constructing a fixed acoustic path inside the vehicle, a sound field topology map of the seating area is generated, forming a virtual microphone set. By using lead singer templates and leak templates for staggered matching and annotation, combined with resonance gating technology, independent acquisition and processing of human voices in each seating area is achieved, ensuring synchronous and consistent audio output.

Benefits of technology

It achieves high-quality zonal acquisition and intelligent mixing of human voices in the complex acoustic environment inside the vehicle, eliminating echo interference and inter-seat interference, improving the interactivity and immersion of multiple passengers singing at the same time, and ensuring the clarity and rhythmic consistency of the audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766640B_ABST
    Figure CN120766640B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of voice data processing, and further relates to a multi-occupant zone-oriented vehicle-mounted microphone-free KTV partition sound pickup and mixing method, which comprises the following steps: step 1: controlling a microphone array to enter a partition sound pickup mode, and sending a partition sound pickup stream and an accompaniment signal to an audio channel manager according to an uplink channel; step 2: after the audio channel manager receives the partition sound pickup stream and the accompaniment signal, performing seat zone multi-shadow separation and resonance gate process, aligning each seat zone reconstruction result with a beat skeleton according to an accompaniment phase anchor, obtaining a seat zone aligned dry sound stream, and obtaining a processed audio signal after spatial reverberation and vocal optimization of the seat zone aligned dry sound stream of each seat zone and the accompaniment signal. The present application solves the problems of unclear partition, distorted sound quality and synchronization difficulty of the existing microphone-free KTV under the condition of multiple sound sources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of voice data processing, and specifically relates to a multi-occupant zone-oriented vehicle-mounted microphone-free KTV partition sound pickup and mixing method. BACKGROUND

[0002] With the development of intelligent networked vehicles, in-vehicle entertainment systems have gradually become one of the important value-added functions of vehicles. Traditional vehicle-mounted entertainment is mainly based on radio, CD playback and Bluetooth music playback. In recent years, with the popularization of mobile Internet and intelligent voice technology, users' demand for diversified entertainment in vehicles has increased. Under this background, vehicle-mounted KTV functions have gradually entered the industry's field of vision and become an important exploration direction for vehicle manufacturers and audio technology manufacturers. Existing vehicle-mounted KTVs mostly rely on microphone sound pickup and sound system playback. Users sing through wireless microphones or wired microphones, and the system then mixes the sound and the accompaniment to form an experience similar to that of a KTV environment.

[0003] However, such traditional solutions based on physical microphones have many limitations in actual use. First, the microphone device needs to be held by the user or installed, which has problems such as inconvenience of use, occupation of in-vehicle space, and additional hardware costs. Second, microphones are extremely susceptible to echo and howling in a limited in-vehicle space, especially when the sound system is close to the microphone. The difficulty of echo cancellation increases significantly, resulting in a decline in singing experience. Third, in a multi-occupant scenario, providing only one or two microphones will make the interaction between passengers insufficient, failing to meet the needs of simultaneous singing, role singing, or multi-person interaction in multiple zones.

[0004] To solve the microphone dependency problem, some manufacturers have proposed a "microphone-free KTV" solution, which directly uses the vehicle-mounted microphone array (usually arranged on the ceiling, doors, or center console) to pick up the human voice and mix it with the accompaniment for playback. The advantage of this approach is that it avoids the inconvenience of holding a microphone, and it also reduces system costs by using existing vehicle-mounted hardware. However, the existing "microphone-free KTV" solution still has several difficulties. SUMMARY

[0005] The main purpose of the present application is to provide a multi-occupant seat area-oriented vehicle-mounted microphone-free KTV partition sound pickup and mixing method, which realizes independent collection and processing of each seat area voice by constructing an in-vehicle fixed sound path and generating a seat area sound field topology map, combined with a virtual microphone set; further, the main singer template and the leakage template are used to stagger match and mark the partition sound pickup stream, effectively separate the main singer shadow, echo shadow and leakage shadow, and suppress non-target fragments according to the path order; when multiple seat areas sing at the same time, the occluded syllable is intermittently reconstructed through the resonance gate, and the beat is aligned by anchoring the accompaniment phase, ensuring the synchronization and consistency of the dry sound stream output by each seat area; finally, combined with spatial reverberation and voice optimization processing, the natural, clear and rich-level audio signal is output. The present application not only overcomes the problems of unclear partition, echo interference and rhythm misalignment of existing microphone-free KTV technology under multi-sound source conditions, but also improves the interactivity and immersion of multiple passengers singing at the same time in the vehicle, significantly improving the overall experience and application value of the vehicle-mounted entertainment system.

[0006] In order to solve the above problems, the technical scheme of the present application is as follows:

[0007] Reference Figure 1 A multi-occupant seat area-oriented vehicle-mounted microphone-free KTV partition sound pickup and mixing method, the method comprising the following steps:

[0008] Step 1: control the microphone array to enter the partition sound pickup mode, perform echo cancellation, noise reduction and voice optimization on each sound pickup collected by the microphone array to obtain a partition sound pickup stream, and send the partition sound pickup stream and the accompaniment signal to the audio path manager according to the uplink path;

[0009] Step 2: after the audio path manager receives the partition sound pickup stream and the accompaniment signal, execute the seat area multiple shadow separation and resonance gate process, which includes: generating a seat area sound field topology map according to the in-vehicle fixed sound path, and forming a virtual microphone set based on the sound field topology map; generating a main singer template and a leakage template with the virtual microphone set, stagger matching the partition sound pickup stream, marking the main singer shadow, echo shadow and leakage shadow, and suppressing non-target fragments according to the path order; when multiple seat areas sing at the same time, intermittently reconstruct the occluded syllable according to the resonance gate; align the reconstruction results of each seat area with the beat skeleton by anchoring the accompaniment phase, obtain the seat area aligned dry sound stream, and after spatial reverberation and voice optimization of the seat area aligned dry sound stream and the accompaniment signal of each seat area, obtain the processed audio signal.

[0010] Further, after receiving the control instruction for triggering the partition pickup mode in step 1, the geometric parameters and the seat area pointing parameter set of the microphone array are loaded, the sampling rate is set to 48000 Hz, the quantization bit depth is set to 24 bits, the frame length is set to 20 ms, the frame shift is set to 10 ms, and a unified time scale is established based on the microphone array master clock; the sampling frequency deviation of the accompaniment signal and the microphone array collected signal is controlled within ± 50 parts per million by using a sampling rate fine-tuning resampling method, so that they can be compared under the unified time scale; in order to reduce the overall time delay, the uplink path target end-to-end delay budget of the partition pickup mode is set to not more than 80 ms.

[0011] Further, the seat area reference coordinates and the in-vehicle fixed sound paths are constructed by the following process: taking the geometric center of the vehicle as the origin, the forward direction as the positive direction of the x-axis, the left side as the positive direction of the y-axis, and the upper side as the positive direction of the z-axis, the seat area reference coordinates are constructed; taking the center of each seat headrest as the reference, moving 70 mm along the x-axis, moving 120 mm along the z-axis, and the y-axis coordinate being located on the seat center line, the mouth reference point coordinates are obtained; the driver's seat, the front passenger seat, the rear left seat, the rear middle seat, and the rear right seat are sequentially labeled as seat area one to seat area five; the three-dimensional coordinates of each array pickup unit in the microphone array are recorded; the direct propagation path from each seat area mouth reference point to each array pickup unit and the reflection path generated on the reflection surface are defined as the in-vehicle fixed sound path.

[0012] Further, in step 2, the seat area reference coordinates, the mouth reference point coordinates of each seat area, the three-dimensional coordinates of each array pickup unit, and the spatial pointing direction are read to determine the front windshield, the rear windshield, the left side window, the right side window, the instrument panel, the door trim panel, the roof liner, and the floor as reflection surfaces, and determine the seat backrest, the headrest, the instrument panel body, the central control channel, and the trunk partition as non-penetrable bodies; for each seat area and each array pickup unit, enumerate the direct propagation path, the first reflection path, and the second reflection path, sample at 10 mm intervals along the path, and if any sampling point falls within the non-penetrable body outer envelope boundary, the path is removed; for the remaining paths, the arrival index is assigned from short to long according to the geometric length, and the path type and reflection surface sequence are recorded, and the time delay is converted and aligned to an integer sampling point at a speed of 343 meters per second and a sampling rate of 48000 Hz.

[0013] Further, in step 2, the process of generating the seat sound field topology map based on the fixed sound path in the vehicle includes: performing seat crossing detection on the reserved path, using a 50mm expansion in each direction for the non-native seat area surrounding boundary, and marking the leakage risk level as high if at least one crossing occurs; calculating the minimum distance of the path to the non-native seat area mouth reference point, and marking it as high if it is less than or equal to 0.25 meters, and marking it as medium if it is between 0.25 meters and 0.40 meters; calculating the angle between the path incident direction and the spatial pointing direction of the array pickup unit, and marking the leakage risk level as high if the non-native seat area path is less than or equal to 25 degrees, and marking the leakage risk level as medium if it is between 25 degrees and 50 degrees; additionally marking the secondary reflection path as medium in the leakage risk level; taking the mouth reference point and the array pickup unit as nodes, and taking the reserved path as edges, and adding the attributes of arrival index, path type, reflection surface sequence, delay sampling point and leakage risk level; organizing in layers within each seat area subgraph, first sorting by arrival index, then sorting by path type as direct propagation priority over one reflection, one reflection priority over two reflections, and sorting by leakage risk level from low to high within the same type, to obtain the seat sound field topology map.

[0014] Further, in step 2, the process of forming a virtual microphone set based on the sound field topology map includes: selecting the array pickup unit with the smallest arrival index and the direct propagation path as the main channel under frame length 20 milliseconds, frame shift 10 milliseconds, if the short-time energy of the main channel is less than the quiet baseline by more than 6 decibel threshold, then sequentially select a direct propagation path with smaller index, if none of the direct propagation paths meet the condition, then select a one-reflection path; aligning the selected channels with the corresponding delay sampling points, and using a 5 millisecond linear cross transition for channel switching; when the evaluated speech intelligibility meets the preset rules, select at most 2 one-reflection paths with arrival index not greater than 3 and leakage risk level low or medium as support channels, and mix them in with a gain of 6 decibels lower than the main channel; applying a 12 decibel fixed attenuation to the non-native seat area incident direction with a high leakage risk level; suspending the support channels when the short-time energy is less than the quiet baseline by more than 3 decibel threshold for 300 milliseconds; when the similarity to the adjacent seat area virtual microphone reaches 0.8 and lasts for more than 100 milliseconds, add its direction to the side suppression direction set; output at least one main virtual microphone for each seat area, if there are two symmetric direct propagation paths with arrival index difference not greater than 2 and leakage risk level low, then additionally output one auxiliary virtual microphone.

[0015] Further, in step 2, select continuous voiced candidate segments on each seat area main virtual microphone with frame length 20 milliseconds, frame shift 10 milliseconds and sampling rate 48000 hertz, extract beat key points, energy envelope, frequency band energy ratio and rise-fall transient, and form the main singer template; collect leakage samples on adjacent seat areas based on the seat sound field topology map and arrival index, and align them by delay sampling points, record the lag, frequency band energy ratio and direction consistency, and form the leakage template.

[0016] Further, in step 2, the lead singer similarity and the leakage similarity are calculated respectively on the partitioned pickup stream in a sliding window with the lead singer template and the leakage template, and the echo similarity is calculated combining the accompaniment beat kick and the reserved path delay sample point number; each window is preliminarily judged as a lead singer shadow, an echo shadow or a leakage shadow according to a preset threshold, conflicts are resolved by score priority and inheritance of the previous window, and continuous judgment of 3 windows is used as the confirmation criterion; when the lead singer shadow is determined, the path components with arrival index from 2 are implemented with intensity positively correlated with the arrival index, and the path with high leakage risk level is additionally attenuated; when the echo shadow is determined, the paths within the range of ± 15 milliseconds of the beat kick are uniformly attenuated, and the paths with arrival index not less than 3 are additionally attenuated; when the leakage shadow is determined, the determined source path and its adjacent paths are gradedly suppressed by a fixed attenuation amount; the rise and fall transitions and the maximum total attenuation upper limit are set for all suppression operations; the timestamped shadow category labeling sequence and the suppression instruction sequence are generated, and the suppressed partitioned pickup stream is output.

[0017] Further, in step 2, the partitioned pickup streams of all seats and the virtual microphone are frame-level detected, when two or more seats are confirmed as lead singer shadow overlap in the same time window, the envelope reference and the beat key points of the lead singer template and the leakage template are combined to determine the occluded syllable segment of the target seat; the formants are extracted in the neighborhood before and after the occluded syllable segment, the formant gating map of open, semi-open and closed states is formed according to the frequency band and time, and the side suppression mark is set according to the leakage template of the adjacent seat at the corresponding lag position, which is used to trigger the subsequent leakage compensation; the occluded syllable segment is baseline recovered with the lead singer template as the envelope reference; in the time slice marked as open state in the formant gating map, the unoccluded detail fragments of the same seat are re-injected according to the level and boundary transition rules; when the side suppression mark is triggered, the leakage template matched therewith is introduced for compensation superposition to obtain the reconstruction result.

[0018] The multi-occupant seat area-oriented vehicle-mounted KTV partition sound pickup and mixing method of the present application has the following beneficial effects: high-quality voice partition collection and intelligent mixing can be realized in a complex acoustic environment in the vehicle, thereby effectively improving the KTV experience in the vehicle. By establishing a fixed sound path in the vehicle and generating a seat area sound field topology map, the present application forms a virtual microphone set that can accurately correspond to the mouth reference point of each seat area, ensuring that the voice of each occupant can be independently captured and recognized. This way not only avoids the space occupation and operational inconvenience brought by traditional handheld microphones, but also eliminates the sound pickup distortion caused by multi-path echoes and mutual interference between seat areas. Further, by generating and interleaving the main singer template and the leakage template, the system can label and distinguish the main singer shadow, echo shadow and leakage shadow that exist simultaneously in multiple seat areas, and through the way of suppressing non-target segments according to the path order, the clean separation of voice signals is realized. Especially when multiple seat areas sing at the same time, the present application introduces a resonance gating mechanism that can intermittently reconstruct the obstructed syllables, avoiding the sense of broken chorus caused by missing voice segments. At the same time, taking the accompaniment phase anchoring as a unified reference, the reconstruction results of each seat area are strictly aligned with the beat skeleton, so that the final output of the seat area-aligned dry sound stream remains synchronized with the accompaniment, ensuring the rhythm consistency of the singing process. Finally, the present application generates processed audio signals with spatial hierarchy and natural voice quality for each seat area through spatial reverberation and voice optimization processing, thereby realizing an immersive KTV experience in the vehicle with multiple users singing simultaneously and without mutual interference. The implementation of the present application not only solves the problems of unclear partition, distorted sound quality and synchronization difficulty of existing microphone-free KTV in the case of multiple sound sources, but also provides a solution with high robustness and high scalability, significantly improving the overall value and user experience of multi-occupant interactive entertainment in the vehicle. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The method flowchart of the multi-occupant seat area-oriented vehicle-mounted KTV partition sound pickup and mixing method provided for the embodiments of the present application is shown in the figure.

[0020] Figure 2 The schematic diagram of the in-vehicle fixed sound path propagation path in the embodiments of the present application is shown in the figure.

[0021] Figure 3 The in-vehicle seat area arrangement and microphone array schematic diagram in the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0022] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0023] Reference Figure 1 The vehicle-mounted KTV partition pickup and mixing method for multi-occupant seat areas includes the following steps:

[0024] Step 1: Control the microphone array to enter the partition pickup mode, and perform echo cancellation, noise reduction and voice optimization on each pickup collected by the microphone array to obtain a partition pickup stream. At the same time, the partition pickup stream and the accompaniment signal are sent to the audio channel manager according to the uplink channel.

[0025] After receiving the control instruction, a seat area list is established according to the correspondence between the vehicle seats and the microphones. Each channel of the microphone array is enabled in turn, and time consistency and channel connectivity verification are completed. Under the condition that the vehicle maintains environmental stability, a silent segment is collected, and the baseline characteristics of environmental noise are extracted as a reference for subsequent noise reduction and voice optimization. Each pickup of the microphone array is mapped to a number of seat channel sets according to the seat area list. A processing flow that cycles by time slice is established to ensure that each seat channel set enters processing at the same time base. For each time slice, the corresponding each pickup is read to form the original seat pickup segment of the time slice. The accompaniment signal currently played is obtained for each seat channel set as an echo reference, and the original seat pickup segment is subjected to echo path estimation and reference alignment. The components associated with the accompaniment signal are first offset, and then the remaining distorted echo components are subjected to nonlinear suppression and boundary smoothing to obtain a seat pickup segment that has completed echo cancellation processing. During the entire processing process, the convergence rhythm of echo cancellation is dynamically adjusted according to the voice activity detection result to avoid introducing excessive processing in the silent area.

[0026] For the seat pickup segment that has completed echo cancellation processing, the ambient noise baseline characteristics obtained in the mute segment are used to suppress the steady-state noise and perform transient suppression on the impulsive interference; the noise reduction strength is controlled to maintain the naturalness of the human voice in the voice activity detection segment, and the noise reduction strength is increased to reduce the background noise in the voice activity detection segment; distortion compensation and edge continuity repair are performed on the processed segment to maintain the continuity between the segment and the adjacent time slice. The seat pickup segment that has completed noise reduction processing is optimized for human voice, and the loudness equalization, dynamic range control, sibilance suppression and formant modification are sequentially implemented, and the spatial direction of the human voice is fixed according to the seat list to obtain stable lead vocal perception; the original speech characteristics and beat boundaries are preserved during the human voice optimization process to avoid affecting the time relationship with the accompaniment signal in the subsequent steps. The seat pickup segments processed by echo cancellation, noise reduction and human voice optimization are spliced into corresponding partition pickup streams in time sequence, and seat identifiers and timestamp identifiers are added to each partition pickup stream to ensure the synchronization between multiple seat areas; the accompaniment signal corresponding to the time period of the partition pickup stream is taken out under the same time base to generate an accompaniment signal segment with a timestamp identifier; the partition pickup stream and the accompaniment signal are packaged into a transmissible data unit according to the uplink path, and are sent to the audio path manager in the order of the seat list for use in subsequent steps.

[0027] Step 2: After the audio path manager receives the partition pickup stream and the accompaniment signal, it performs seat multi-shadow separation and resonance gating processes, which include:

[0028] Specifically, a seat sound field topology map is generated according to the in-vehicle fixed sound path, and a virtual microphone set is formed based on the sound field topology map, including: according to the definition of the seat reference coordinates and the mouth reference point coordinates, the three-dimensional coordinates of each seat's mouth reference point and the array pickup unit are read, and an index table is established according to seat one to seat five; the front windshield, rear windshield, left side window, right side window, instrument panel, door trim, roof lining and floor are taken as the main reflection surfaces, and the geometric boundaries and normal directions of each reflection surface are recorded; the above information is written into the in-vehicle fixed sound path candidate library as the only input for subsequent path enumeration and screening. For each seat and each array pickup unit, a path set is generated in the following order: direct propagation path, one reflection path, two reflection paths. The direct propagation path is obtained by connecting the straight line between the mouth reference point and the array pickup unit; the one reflection path is obtained by connecting the mirror image point and the array pickup unit after geometrically mirroring the mouth reference point on a single reflection surface; the two reflection paths are obtained by connecting the final mirror image point and the array pickup unit after geometrically mirroring on two different reflection surfaces. The reachability of each candidate path is determined, and if the straight line segment or segmented polyline segment is blocked by the geometric boundary of the in-vehicle obstacle, the path is removed; the path type, the sequence of reflection surfaces involved, and the relative level of geometric length are recorded for the paths that are not removed.

[0029] According to the geometric length from short to long, arrange the arrival index of each reserved path, and consecutively number from index 1; mark the path type as direct path, one-time reflection path or two-time reflection path; write the sequence of the reflection surface involved in the path label in a fixed order; according to the spatial pointing relationship between the seat area range passed by the path and the array pickup unit, mark the leakage risk level of each path as low, medium or high; according to the reflection surface material and the incident angle range, mark the reflection stability level of each path as high, medium or low. After completion, the path attribute table in seat area units is obtained. A graph structure is established with the mouth reference point and the array pickup unit as nodes and the fixed in-vehicle sound path determined to be reserved as edges; the arrival index, path type, reflection surface sequence, geometric length level, leakage risk level and reflection stability level are attached to each edge; the whole graph is divided into five subgraphs according to the seat area, and each subgraph is further layered according to the arrival index, so that the layer with smaller index has higher priority, forming a layered structure from near to far and from direct to multiple reflections. The seat area sound field topology graph is thus generated.

[0030] For each seat area, direction sampling is performed on all array pickup units in the corresponding subgraph. The direction pointing to the mouth reference point from the array pickup unit is defined as the main pickup direction, and the direction communicating with the adjacent seat area and having a medium or high leakage risk level is defined as the side suppression direction. In each layer, a pickup priority is assigned to each edge according to the path type and the reflection stability level. The direct path priority is higher than the one-time reflection path, and the one-time reflection path priority is higher than the two-time reflection path. A weight schedule is generated for each edge, which includes the pickup priority, the boundary smoothing time length, the switching order with the adjacent layer, the suppression intensity level for the leakage risk level, and the silence threshold level when the voice activity detection is silent. According to the weight schedule of step five, the following rules are executed frame by frame under the time organization of the frame length of 20 milliseconds and the frame shift of 10 milliseconds set in step 1: the array pickup unit signal corresponding to the edge pair with the minimum arrival index and the direct path type is preferentially selected as the main channel; if the main channel short-time energy is insufficient or the voice activity detection is silent, the array pickup unit signals corresponding to the one-time reflection path and the two-time reflection path are sequentially supplemented according to the pickup priority, and boundary smoothing is applied within the frame boundary during the supplement; the array pickup unit signal corresponding to the edge marked as the side suppression direction and having a medium or high leakage risk level is subjected to attenuation corresponding to the suppression intensity level, and the same arrival index order as the main channel is maintained within the frame; a fixed order is executed for the cross-layer switching, and multiple switching within the same frame is prohibited to ensure stability. The signal obtained by the above frame-by-frame synthesis is defined as the virtual microphone of the seat area, and at least one virtual microphone is formed as the main virtual microphone; when the seat area sound field topology graph shows that there are two groups of main pickup directions that are symmetrically valid and have a low leakage risk level, one virtual microphone is allowed to be additionally formed as an auxiliary virtual microphone in the same seat area to enhance the audibility of the edge seat.

[0031] In a continuous time period, the short-time spectral shape and beat variation of the main virtual microphone and the auxiliary virtual microphone are compared for similarity. If the association degree with the mouth reference point of the corresponding seat area decreases and the association degree with the adjacent seat area increases, the leakage risk level of the adjacent seat area related edge is upgraded by one level, and the pickup priority of the corresponding edge in the current seat area is downgraded by one level. If a silent state lasts for several frames, the weight schedule of the current layer is restored to the default setting. Through the above process, a virtual microphone set including the virtual microphone of each seat area, the corresponding array pickup unit list, the weight schedule, the arrival index table, and the leakage suppression table is output for subsequent steps.

[0032] The intelligibility is evaluated frame by frame on the main virtual microphone signal of each seat area. If any two of the following conditions are met, it is determined that the intelligibility is insufficient:

[0033] Voice on-off boundary deviation: within a continuous voiced period, the voice on-set and off-set boundaries of the main virtual microphone in the seat zone are detected with reference to the beat boundaries of the accompaniment beat skeleton; when the time deviation of the on-set or off-set boundary exceeds 30 milliseconds and occurs 3 times or more within a 500-millisecond window, it is recorded as a boundary deviation unqualified. High-frequency detail loss: in a transient segment containing a fast rising edge, the high-frequency energy proportion is calculated; when the high-frequency energy proportion is less than 15% of the total energy and the situation accumulates for 120 milliseconds or more within a 300-millisecond window, it is recorded as a high-frequency detail loss unqualified. The transient segment is determined by a rising edge threshold, which is set to an adjacent frame energy increase of 6 decibels and a duration of 5 to 15 milliseconds. Discontinuous distortion insertion: a voice insertion event occurs in a segment determined to be voiced, with a single insertion duration of 20 milliseconds or more; when 3 or more occur within 1 second, it is recorded as a discontinuous distortion unqualified. Voice insertion is determined by a threshold of 3 decibels above the silence baseline. Residual noise masking: in a pause segment that should be blank, the average energy is continuously higher than 3 decibels above the silence baseline for more than 200 milliseconds; or the same texture is detected to occur more than 2 times in the pause segment, which is recorded as a residual noise masking unqualified. Direction consistency decrease: with reference to the main sampling direction, the incident direction of the main virtual microphone is estimated; when the deviation angle of the incident direction relative to the main sampling direction is greater than 25 degrees and the duration is more than 100 milliseconds, it is recorded as a direction consistency unqualified. The incident direction is estimated according to the arrival index of the selected path in the seat zone sound field topology and the spatial pointing direction of the array pickup unit. Echo residual synchronization: at the time point near the accompaniment beat kick, the main virtual microphone appears a peak value synchronized with the beat kick and the amplitude is higher than 6 decibels above the silence baseline; when 3 or more occur continuously, it is recorded as an echo residual unqualified. Pronunciation intensity imbalance: within a continuous three-syllable range, the difference between the peak value of any syllable and its adjacent valley value is less than 6 decibels, and the imbalance phenomenon occurs in two or more syllables at the same time, which is recorded as a pronunciation intensity imbalance unqualified. Syllable boundaries are determined by energy rising and falling inflection points.

[0034] The execution criteria are as follows: in each frame processing period, the above 7 indicators are calculated in turn; when any 2 are determined to be unqualified, the "insufficient intelligibility" flag is triggered; once marked, immediately enable the support channel mixing according to the previously agreed process, and keep the 5-millisecond linear cross transition of channel switching and the corresponding delay alignment; if "insufficient intelligibility" is still continuously triggered within 200 milliseconds, add 1 path with arrival index not greater than 3 and leakage risk level low to the support direction set to participate in mixing; if "insufficient intelligibility" is not triggered again after the addition for 300 milliseconds, restore the use of only the main channel; all determination results and processing actions are recorded with frame timestamps and written into the process log of the seat zone virtual microphone for subsequent steps of interleaving matching, labeling and alignment.

[0035] Specifically, the lead vocal template and the leakage template are generated based on the virtual microphone set, the partitioned pickup stream is cross-matched, the lead vocal shadow, echo shadow and leakage shadow are marked, and the non-target segments are suppressed according to the path order, including: inputting the virtual microphone set of each seat area, the partitioned pickup stream, the accompaniment signal, the seat area sound field topology map, the arrival index and the delay sampling point number of each reserved path. The processing is performed with a frame length of 20 milliseconds and a frame shift of 10 milliseconds, and all signals are executed with a sampling rate of 48000 Hz, a quantization bit depth of 24 bits and a unified time scale.

[0036] In the main virtual microphone of each seat area, a continuous segment that meets the following three conditions is found as a candidate segment: the voice activity detection is voiced and the continuous duration is not less than 400 milliseconds; the average energy of the segment is higher than the quiet baseline by 9 decibels or more; and the beat skeleton is aligned with the accompaniment signal, covering 4 or more consecutive beats. Normalization and alignment: normalize the candidate segment according to the reference level of the main virtual microphone of the seat area; if the candidate segment is spliced from multiple array pickup units, align each segment according to the delay sampling point number of the corresponding path before splicing.

[0037] For each candidate segment, the following four types of elements are extracted to form the lead vocal template entry of the segment: beat key point sequence: records the relative time of each beat kick and weak beat, allowing a deviation of not more than ±10 milliseconds; energy envelope curve: records the normalized energy envelope at an interval of 10 milliseconds; frequency band energy ratio: divides the audible frequency into low, medium and high frequency bands, and records the ratio of the energy of the three bands (expressed in percentage); rise and fall transient form: records the rise time and fall time, respectively expressed in milliseconds. Representative template determination: select up to 3 lead vocal templates as representative templates of the seat area according to the singing situation covered by the segment, corresponding to strong, weak and continuous singing segments; if there are less than 3 available candidates, select all available templates in descending order of frequency.

[0038] In the seat area sound field topology map of the seat area, find the edges from adjacent seat areas with medium or high leakage risk level, and record the arrival index and delay sampling point number of these edges to form a leakage source list. Leakage sample collection: select a segment that meets the following two conditions as a leakage sample on the main virtual microphone of the adjacent seat area: the average energy is 6 to 12 decibels lower than the average energy of the lead vocal template; and align with the beat skeleton of the adjacent seat area, covering 2 or more consecutive beats. Path unification processing: align the leakage sample according to the delay sampling point number of the corresponding edge in the leakage source list; if the edge is a one-time reflection path, reduce the high frequency band energy of the leakage sample by 3 decibels; if the edge is a two-time reflection path, reduce the high frequency band energy by an additional 3 decibels based on the 3 decibel reduction.

[0039] Leak template elements: record the following elements for the aligned leak sample to form a leak template entry: arrival index and corresponding delay sample point number; energy envelope flatness, expressed as the decibel value of envelope peak-valley difference; time lag with the seat area beat skeleton, expressed in milliseconds; frequency band energy proportion and direction consistency, direction consistency is grouped by the angle between the incident direction and the spatial pointing direction of the array pickup unit, less than or equal to 25 degrees is strong consistency, between 25 degrees and 50 degrees is moderate consistency, and greater than 50 degrees is weak consistency. According to the occurrence frequency and direction consistency of the leak source edge, at most 2 leak templates are retained for each adjacent seat area.

[0040] On the partitioned pickup stream of each seat area, a sliding window with the same length as the template is traversed with a step of 10 milliseconds. Three scores are calculated between each window and each lead vocal template, and the lead vocal similarity (range 0 to 1) is obtained by weighting: beat key point alignment score, accounting for 50%. If the average deviation of the key points in the window and the template key points is not more than 20 milliseconds, it is recorded as 1, the deviation is between 20 milliseconds and 40 milliseconds, and it is reduced from 1 to 0 in a linear manner, and more than 40 milliseconds is recorded as 0; energy envelope similarity score, accounting for 30%. The normalized energy difference between the window and the template is linearly mapped from 0 to the maximum difference to 1 to 0; frequency band proportion similarity score, accounting for 20%. The absolute difference of the three proportion ratios is less than or equal to 10% for 1, any proportion absolute difference reaches 30% for 0, and the rest is interpolated in a linear manner. Take the highest lead vocal similarity of all lead vocal templates in this window as the lead vocal similarity of the window.

[0041] Echo similarity calculation: In each window, the number of peaks synchronized with the accompaniment beat kick and the amplitude proportion are counted, and combined with the delay constraint corresponding to the arrival index to obtain the echo similarity (range 0 to 1): synchronization rate score, proportion 70%. A peak appearing within ±10 milliseconds of each kick point is considered a hit, with a hit rate of 1, otherwise it is counted as 0; delay compliance score, proportion 30%. If the time difference between the window peak and the kick point is equal to the time corresponding to the delay sampling point number of any reserved path, and the deviation is not more than ±10 milliseconds, it is recorded as 1, the deviation is between ±10 milliseconds and ±30 milliseconds, and it is reduced from 1 to 0 in a linear manner, and more than ±30 milliseconds is recorded as 0. Leakage similarity calculation: Calculate three scores between each window and each leakage template and get the leakage similarity (range 0 to 1): delay compliance score, proportion 40%. If the lag of the peak time of the window relative to the beat skeleton of the seat area is consistent with the lag of the leakage template and the deviation is not more than ±10 milliseconds, it is recorded as 1, the deviation is between ±10 milliseconds and ±30 milliseconds, and it is reduced from 1 to 0 in a linear manner; direction consistency score, proportion 30%. If the leakage template direction consistency is strong consistency, it is recorded as 1, medium consistency is recorded as 0.5, and weak consistency is recorded as 0; frequency band proportion matching score, proportion 30%. If the absolute difference of the three proportions is less than or equal to 15%, it is recorded as 1, and if the absolute difference of any proportion reaches 35%, it is recorded as 0, and the rest is interpolated in a linear manner. Take the highest leakage similarity of the window to all leakage templates as the leakage similarity of the window.

[0042] Shadow category labeling: For each window, preliminary judgment is made according to the following thresholds: lead vocal shadow: lead vocal similarity is greater than or equal to 0.70, and leakage similarity is less than or equal to 0.40, and echo similarity is less than or equal to 0.50; echo shadow: echo similarity is greater than or equal to 0.70, and lead vocal similarity is less than or equal to 0.50; leakage shadow: leakage similarity is greater than or equal to 0.60, and lead vocal similarity is less than or equal to 0.60. If more than two categories meet the preliminary judgment conditions at the same time, the category with the highest score is selected; if the difference between the highest scores is less than 0.05, the category confirmed in the last window is retained; if there is no confirmed category in the last window, it is temporarily recorded as a pending category and combined with the next window for judgment. The same category is determined in the next 3 windows, and the category is confirmed; if the same category appears twice in 3 windows and the other is a pending category, the category is also confirmed.

[0043] Lead shadow window suppression: When the window is identified as a lead shadow, apply hierarchical suppression to all components corresponding to paths with arrival precedence index greater than or equal to 2 in the seat: 6 dB attenuation for arrival precedence index 2, 12 dB attenuation for arrival precedence index 3 or 4, 18 dB attenuation for arrival precedence index greater than or equal to 5; if the path leakage risk level is high, an additional 6 dB attenuation is applied on top of the above. The rise and fall of the suppression is executed with a 5 ms linear transition. Echo shadow window suppression: When the window is identified as an echo shadow, uniformly attenuate all path components within the range of ± 15 ms of the beat kick by 12 dB; outside this range, it remains unchanged; if the echo shadow persists for 3 consecutive windows, an additional 6 dB attenuation is applied to paths with arrival precedence index greater than or equal to 3. Leakage shadow window suppression: When the window is identified as a leakage shadow, identify the adjacent seat source path corresponding to the leakage template, and attenuate the path component by 18 dB; attenuate the path components of the homologous adjacent path with an arrival precedence index difference not greater than 1 by 12 dB; attenuate the path components of the remaining non- seat source and leakage risk level of medium or high by 6 dB. Continuity and upper limit constraint: the total attenuation after any single suppression does not exceed 24 dB; if the category changes more than twice within 200 ms, the transition time of all suppression rise and fall is increased to 10 ms to avoid audible mutations. Record the lead similarity, echo similarity and leakage similarity, shadow category, path list involved in suppression, corresponding arrival precedence index and attenuation amount for each window, and add a timestamp. Generate a shadow category label sequence and a suppression instruction sequence as input for subsequent steps of discontinuous reconstruction and phase anchoring; keep the partitioned pickup stream after suppression as an updated version for subsequent processing of the seat.

[0044] Specifically, when multiple seats sing at the same time, the occluded syllable is discontinuously reconstructed according to the resonance gating; the reconstruction results of each seat are aligned with the beat skeleton by the accompaniment phase anchoring to obtain seat-aligned dry sound streams; after spatial reverberation and human voice optimization of the seat-aligned dry sound streams of each seat and the accompaniment signal, the processed audio signal is obtained, including: under the uniform time scale, sampling rate of 48000 Hz, quantization bit depth of 24 bits, frame length of 20 ms, frame shift of 10 ms time organization, the partitioned pickup stream and virtual microphone of all seats are detected frame by frame to determine whether two or more seats are labeled as lead shadows in the same time window, and if the continuous overlapping time length reaches or exceeds 50 ms, it is determined that multiple seats sing at the same time. In each determined time window, the main virtual microphone of the target seat is divided by syllable boundary; if the syllable has an energy trough in the middle section and is lower than the envelope reference value of the lead template by more than 6 dB, and the leakage similarity of the adjacent seat leakage template reaches or exceeds 0.60, mark the trough section of the occluded syllable section; the boundaries of the occluded syllable section are determined by the energy turning point and the envelope reference value intersection point, and the minimum reconstruction unit is not less than 30 ms.

[0045] The target seat zone main virtual microphone is within the range of 100 milliseconds before and after the occluded syllable segment, and the duration and relative energy of the stable peak are counted in the low frequency band, the medium frequency band and the high frequency band respectively. The peak value with duration reaching or exceeding 40 milliseconds and relative energy higher than the silence baseline by more than 9 decibels is identified as a formant. The formant gating map is generated with the formant as the center: the segment within the frequency band where the formant is located and the range of 10 milliseconds before and after it is marked as "open", the segment without formant and with energy lower than the silence baseline by more than 3 decibels is marked as "closed", and the remaining segment is marked as "half open". The minimum holding time of "open" state is 30 milliseconds, the minimum holding time of "closed" state is 20 milliseconds, and "half open" state is used for state transition and keeps for no less than 10 milliseconds. When the arrival index corresponding to the adjacent seat zone leakage template is less than or equal to 3 and the direction consistency is strong consistency, the time point matching the lag of the leakage template is added with a "side suppression mark" on the formant gating map of the target seat zone as a trigger condition for subsequent leakage compensation.

[0046] The discontinuous reconstruction (performed for each occluded syllable segment) process includes: baseline recovery: using the main singer template entry most similar to the syllable as the envelope reference to generate the target envelope within the occluded syllable segment; the residual signal within the occluded syllable segment is amplitude recovered according to the target envelope, so that the peak-to-valley difference is not less than 70% of the template peak-to-valley difference. When recovering, a 5-millisecond linear cross transition is set before and after the segment to avoid boundary mutation. Detail restoration: search for an unoccluded segment with the same vowel category or a beat key point sequence similarity reaching or exceeding 0.70 within the same seat zone in the last 1 second, and cut off a detail segment of no more than 120 milliseconds as a formant re-injection source; within the time slice where the formant gating map is "open", superimpose the detail segment to the target envelope with a mixing level not higher than 6 decibels of the current frame main virtual microphone level, and use 10-millisecond linear cross transition for mixing and removal. Leakage compensation: when the "side suppression mark" is triggered, select the entry with the most matching lag from the leakage template, align the number of delay sampling points according to the entry, and perform compensation superposition on the corresponding time slice within the occluded syllable segment, with the compensation level set to 80% of the leakage component estimated level of the time slice and linearly started and stopped within 5 milliseconds. If the feature peak value with consistent direction with the entry is still detected after compensation, add a 20% compensation. Consistency check: after reconstruction, check whether the beat key point sequence deviation of the syllable is less than or equal to ±20 milliseconds, and whether the energy envelope peak-to-valley difference reaches or exceeds 70% of the corresponding main singer template peak-to-valley difference. If either is not met, perform a one-time smooth stretching or compression on the entire syllable, with the absolute time length of stretching or compression not exceeding 20 milliseconds, and maintain a 5-millisecond cross transition at both ends. Output the reconstruction result of the syllable after completion.

[0047] The process of accompaniment phase anchoring and seat area alignment dry sound stream generation includes: extracting the beat skeleton from the accompaniment signal, and determining the time points of each strong beat and weak beat. The time deviation of the target seat area at the "starting point", "core point" and "ending point" of the reconstructed syllable is calculated with the adjacent anchor points. When the absolute deviation of any key point is less than or equal to 15 milliseconds, the micro-shift correction is directly performed on the segment where the key point is located; when the absolute deviation is between 15 milliseconds and 40 milliseconds, the correction is completed by shortening or lengthening the adjacent pause segment without changing the internal structure of the syllable, and the single adjustment does not exceed 20 milliseconds; when the absolute deviation is greater than 40 milliseconds, it is distributed to the subsequent 3 anchor points, and the adjustment of a single anchor point does not exceed 20 milliseconds, and the cumulative adjustment does not exceed 50 milliseconds within any 1 second. Each corrected syllable is spliced in time sequence to form the seat area alignment dry sound stream; a 5 millisecond linear cross transition is uniformly applied at the syllable splicing place, and a time stamp is added to ensure synchronization during subsequent mixing.

[0048] The process of spatial reverberation and voice optimization and mixing with accompaniment output includes: spatial reverberation: according to the seat area sound field topology, select the 3 paths with the smallest index of direct propagation and arrival time as the early reflection reference, the relative level of early reflection is set to be 6 decibels, 9 decibels and 12 decibels less than the main direct level in turn, and the start and stop of early reflection are executed with 5 millisecond linear transition; then superimpose the decay type reverberation tail, the tail length is pre-set according to the volume category of the vehicle and decreases by 100 milliseconds until the overall decay is within 3 decibels of the quiet baseline. Voice optimization: the seat area alignment dry sound stream is first executed for loudness equalization, and then executed for two sections of dynamic range control, the first section is used to control the fast peak, the starting time is 10 milliseconds and the recovery time is 80 milliseconds; the second section is used to control the slow fluctuation, the starting time is 60 milliseconds and the recovery time is 300 milliseconds; then execute the sibilance suppression and formant modification to make the frequency band energy ratio close to the target ratio of the main singer template, with a deviation of not more than 10%. Mixing with accompaniment: compare the average energy of the seat area alignment dry sound stream and the accompaniment signal in a 200 millisecond sliding window, when the accompaniment is 6 decibels higher than the voice and above, reduce the accompaniment by 3 decibels; when the voice is 6 decibels higher than the accompaniment and above, reduce the voice by 1 decibel to avoid excessive prominence; all gain changes are realized with 10 millisecond linear transition. After mixing, the transient peak is fixed to 6 decibels by the peak limiter. Output: output the processed audio signal after adding spatial reverberation and voice optimization for each seat area, and keep the time alignment mark with the beat skeleton for vehicle playback and subsequent scoring.

[0049] A complete embodiment of the present application is given below:

[0050] I. Vehicle and sampling parameter setting: the sampling rate is set to Hz audio sampling rate), quantization bit depth 24 bits. Frame length and frame shift are seconds ( seconds seconds seconds. Corresponding sampling points are ; wherein is the number of sampling points per frame, is the number of sampling points of frame shift. The origin of the in-vehicle coordinate system is at the geometric center of the passenger compartment, the forward direction is the positive direction of the axis, the left side is the positive direction of the axis, and the upper side is the positive direction of the axis. The seat area mouth reference point (the mouth reference point is the spatial position of the singer's mouth) is valued at: seat area one mouth reference point meters; seat area two mouth reference point meters. The microphone array pickup unit (the array pickup unit is a spatially fixed microphone position) is selected as 6: meters, meters, meters, meters, meters, meters. The song beat speed is selected as ( beats per minute), and the beat period is ; wherein is the time between adjacent strong beats.

[0051] II. Fixed sound path and arrival index example (direct path from seat area one to several microphones): the Euclidean distance between the path geometric length ( and ) is ; wherein is the path length. The speed of sound in air is selected as meters per second ( sound speed), and the arrival time is , corresponding to the aligned delay sampling point number ; wherein is the propagation time, is the delay point number.

[0052] Calculation example: direct path from seat area one to microphone four ; .

[0053] Calculation example: direct path from seat area one to microphone one and microphone two ; .

[0054] Therefore, the arrival index of seat area one is microphone four, microphone one, microphone two in order, one, two, three.

[0055] Frame-level synthesis of seat zone 1 virtual microphone (main channel and support channels): frame-wise short-time root-mean-square envelope ; wherein is the energy envelope of the th frame, is the waveform sample sequence. If the main channel threshold is a relative quiet baseline in decibels (in logarithmic scale), when the frame energy of microphone 4 satisfies the threshold, microphone 4 is selected as the main channel and aligned with the delay point. When is below the threshold, microphone 1 and microphone 2 are examined in turn. When the intelligibility is insufficient (satisfying any two of the seven indicators as previously specified), 1-2 support channels are selected from the first reflection or the direct path of the neighbor, and mixed in at a level decibels lower than the main channel. The level conversion is linear gain ; wherein is the linear gain, decibels when .

[0056] Main singer template and leakage template instances: beat skeleton time points are (when is the th strong beat moment, is the song alignment starting point). A main singer candidate segment covering 4 strong beats is taken seconds long, and a leakage candidate segment is taken seconds long. The beat key point sequence is aligned with the segment within milliseconds, which is considered perfect alignment; the energy envelope is normalized to a maximum value of 1 every milliseconds; the average of the three-band energy ratios (low, medium, high) in the segment is set to . Seat zone 2 sings in the same time period. The direct path length from seat zone 2 to microphone 4 is corresponding to a delay point. The relative delay difference from seat zone 1 to microphone 4 is ; the leakage template records a lag of points, a band energy ratio of about , and a direction consistency of moderate consistency (angle of about 35 degrees).

[0057] One window calculation of interleaved matching and shadow marking: the sliding window length is equal to the template length, and the step is seconds. Within the window seconds, three scores are calculated with the main singer template. The singer similarity is composed of three parts ; wherein is the singer similarity, is the beat alignment score, for envelope similarity score, for band proportion score. Let the window beat average deviation be mapped to , envelope absolute difference be mapped to , three segment proportion difference be mapped to , respectively . Echo similarity is calculated with beat synchronization and delay coincidence (let the score be). Leak similarity combines lag , moderate direction consistency and band proportion match (let the score be). According to threshold rule (lead vocal shadow threshold ), the window is labeled as lead vocal shadow.

[0058] Simultaneous singing and occluded syllable intermittent reconstruction of a segment: in seconds, the energy of the seat area of a virtual microphone is lower than the template by decibels, and the peak value of the leak similarity reaches . It is determined as an occluded syllable segment. The template peak-to-valley difference is decibels, and the target is at least recovered, that is decibels. Set the gain of the segment to ; and apply a linear crossfade of milliseconds at both ends of the segment. Search for similar syllable fragments in the last seconds in the same seat area seconds, cut off milliseconds and align them to superimpose them at a mixing level of decibels, that is , the injection interval is limited to the time slice when the resonance gate is "on". Align the leak template lag points, superimpose the offset component in the occluded segment, and set the offset level to of the leak estimation level. If the remaining peak value consistent with the direction of the leak template is still visible, add compensation. After reconstruction, the beat key point deviation is milliseconds, less than milliseconds threshold; the envelope peak-to-valley difference reaches decibels, meeting requirements, and the reconstruction is completed.

[0059] Accompaniment phase anchoring and alignment: the accompaniment beat sequence is seconds. Calculate the deviation at the starting point, core point, and end point of the reconstructed syllable: for example, the starting point deviation is milliseconds. Shorten the adjacent pause segment by milliseconds and move the entire syllable slightly milliseconds, so that the absolute deviation is not greater than milliseconds. All micro-movements are limited to single milliseconds, and the use of milliseconds linear cross-fade. All aligned syllables are time-stamped to obtain the seat zone aligned dry sound stream.

[0060] Spatial reverberation, voice optimization and accompaniment mixing output: Early reflections: select the 3 shortest first reflection paths associated with the seat zone, and set the level difference of the early reflections relative to the main direct sound level as , , dB, start-stop transition milliseconds. Reverberation tail envelope ; wherein is the tail envelope, is the tail time constant. Set seconds, so that the tail is about seconds below the quiet baseline dB. Preceding loudness equalization, then two sections of dynamic range control. The first section starts milliseconds, and the recovery time is milliseconds, and the second section starts milliseconds, and the recovery time is milliseconds; then sibilance suppression and formant modification, so that the frequency band ratio converges to the deviation of the main vocal template not more than . Compare the average energy of the voice and accompaniment in a milliseconds sliding window. When the accompaniment is dB and above higher than the voice, the accompaniment is attenuated by dB; when the voice is dB and above higher than the accompaniment, the voice is attenuated by dB. All gain changes are realized with milliseconds linear transition. The peak limiter leaves dB margin, linear coefficient ; respectively obtain the processed audio signals of seat zone one and seat zone two, each signal carries a time alignment mark with the beat skeleton, which is used for partitioned playback and scoring.

[0061] As Figure 2As shown, the sound path analysis system of the present application details the sound propagation paths and their geometric characteristics from the mouth reference point of the seating zone to each array pickup unit. In the figure, the mouth reference point of the first seating zone is taken as the sound source starting point, and a certain array pickup unit is taken as the end point, showing three different types of sound propagation paths: direct propagation path, once-reflected path, and twice-reflected path. The direct propagation path is represented by a solid line, which is the shortest geometric path from the mouth reference point of the first seating zone to the array pickup unit. This path does not pass through any reflecting surface, and the sound signal propagates in a straight line, with the shortest propagation time and the smallest energy loss. The direct propagation path has the highest priority and the lowest risk level of leakage in the sound field topology map. The once-reflected path is represented by a dashed line and includes two typical cases. The first is the reflection path through the roof lining, where the sound travels from the mouth reference point of the first seating zone, reflects off the surface of the roof lining, and then propagates to the array pickup unit. The second is the reflection path through the floor, where the sound reflects off the surface of the floor before reaching the target location. The once-reflected path has a longer propagation distance than the direct propagation path, and the corresponding arrival time is also delayed accordingly, giving it a higher arrival index value in the system. The twice-reflected path is represented by a dotted and dashed line and involves two consecutive reflection processes. The twice-reflected path shown in the figure first reflects off the left window surface for the first time, then reflects off the roof lining surface for the second time, and finally reaches the array pickup unit. The twice-reflected path has the longest propagation distance and the largest propagation delay, and its signal strength is relatively weak due to the energy attenuation caused by multiple reflections. The interior reflecting surfaces include the front windshield, rear windshield, left window, right window, instrument panel, door trim panel, roof lining, and floor surfaces. These reflecting surfaces have different acoustic characteristics and reflection coefficients, affecting the spectral characteristics and energy distribution of the reflected sound. In contrast, the non-traversable bodies include the seat backrest, headrest, instrument panel body, center console channel, and trunk partition, which block the sound propagation path. The system analyzes all possible paths between each seating zone and each array pickup unit, sampling and detecting at 10mm intervals along the path. If any sampling point falls within the outer boundary of the non-traversable body, the path is automatically excluded. The remaining paths are sorted by geometric length from short to long and assigned a corresponding arrival index, while recording path type, reflection surface sequence, and other attribute information, providing basic data for subsequent sound field topology map construction and virtual microphone formation.

[0062] As Figure 3As shown, the vehicle-mounted microphone-free KTV partition sound pickup system of the present application adopts a specific seat zone division scheme and microphone array arrangement structure in the vehicle interior. The vehicle interior space is divided into five independent seat zones, which are sequentially identified as seat zone one to seat zone five from front to back and from left to right. Specifically, seat zone one is located at the driver's seat, seat zone two is located at the front passenger seat, and seat zone three, seat zone four, and seat zone five are located at the rear left seat, the rear center seat, and the rear right seat, respectively. Each seat zone is identified as a rectangular area, and its boundary range is determined based on the normal sitting posture and sound production position of the occupant, ensuring that the main activity range of the occupant in each seat zone is covered. The microphone array system includes eight array pickup units, which are labeled as Mic 1 to Mic 8, and adopts a distributed arrangement strategy to achieve effective coverage of the entire vehicle space. In the front area of the vehicle, Mic 1 to Mic 5 are distributed in a horizontal line, with Mic 1 located on the left side of the front of the vehicle, Mic 2, Mic 3, and Mic 4 arranged in sequence to the right in the central area of the front of the vehicle, and Mic 5 located on the right side of the front of the vehicle. This front array arrangement can effectively pick up the sound signals of the two front seat zones, while providing appropriate directional support for the rear seat zones. In the side area of the vehicle, Mic 6 is arranged in the middle left position, and Mic 7 is arranged in the middle right position, forming a symmetrical lateral pickup configuration, mainly responsible for picking up sound signals from the side and providing spatial positioning reference. In the rear area of the vehicle, Mic 8 is located in the central position of the rear, and is specially responsible for sound pickup of the rear seat zones. The spatial geometric configuration of the entire microphone array system fully considers the complexity of the in-vehicle acoustic environment, and realizes effective separation and precise positioning of sound sources in different seat zones through multi-point distributed arrangement. The array pickup units maintain appropriate spatial distance, which not only ensures the continuity of the sound field coverage, but also avoids signal interference between adjacent units. This arrangement scheme lays the foundation for subsequent partition sound pickup, sound field topology construction, and virtual microphone formation.

[0063] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not limit it; although the foregoing embodiments of the present application have been described in detail, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-passenger seat area-oriented in-vehicle wireless KTV partition sound pickup and mixing method, characterized in that, The method comprises the following steps: Step 1: control the microphone array to enter the partition pickup mode, perform echo cancellation, noise reduction and voice optimization on each pickup of the microphone array to obtain a partition pickup stream, and send the partition pickup stream and an accompaniment signal to an audio channel manager according to an uplink channel; Step 2: after receiving the partition pickup stream and the accompaniment signal, the audio channel manager performs seat area multi-shadow separation and resonance gating processes, which include: generating a seat area sound field topology based on a fixed sound path in the vehicle, and forming a virtual microphone set based on the sound field topology; generating a main singer template and a leakage template based on the virtual microphone set, and performing cross matching on the partition pickup stream to label main singer shadows, echo shadows and leakage shadows, and suppressing non-target segments according to the path order; when multiple seat areas sing at the same time, discontinuous reconstruction is performed on the occluded syllable segment based on resonance gating; aligning the reconstruction results of each seat area with the accompaniment phase anchor and the beat skeleton to obtain a seat area aligned dry sound stream, and performing spatial reverberation and voice optimization on the seat area aligned dry sound stream of each seat area and the accompaniment signal to obtain a processed audio signal; In step 2, continuous voiced candidate segments are selected on each main virtual microphone according to a frame length of 20 milliseconds, a frame shift of 10 milliseconds and a sampling rate of 48000 Hz, beat key points, energy envelope, frequency band energy ratio and transient are extracted to form a main singer template; leakage samples are collected on adjacent seat areas according to the seat area sound field topology and the arrival index, and are aligned according to the delay sampling point number, the lag, the frequency band energy ratio and the direction consistency are recorded to form a leakage template; In step 2, the main singer similarity and the leakage similarity are calculated on the partition pickup stream with a sliding window and the main singer template and the leakage template respectively, and the echo similarity is calculated by combining the accompaniment beat kick and the reserved path delay sampling point number; each window is preliminarily judged as a main singer shadow, an echo shadow or a leakage shadow according to a preset threshold, conflicts are resolved by using score priority and inheritance of the previous window, and continuous determination of 3 windows is used as the confirmation criterion; when it is determined to be a main singer shadow, the path components with an arrival index of 2 or more are subjected to graded attenuation positively correlated with the arrival index and additional attenuation for paths with a high leakage risk level; when it is determined to be an echo shadow, uniform attenuation is performed within a beat kick ± 15 milliseconds range, and additional attenuation is performed on paths with an arrival index of not less than 3; when it is determined to be a leakage shadow, the determined source path and its adjacent paths are gradedly suppressed by a fixed attenuation amount; the rise and fall edges are transitioned and the maximum total attenuation upper limit is set for all suppression operations; a timestamped shadow category labeling sequence and a suppression instruction sequence are generated, and the suppressed partition pickup stream is output. In step 2, frame-level detection is performed on the partitioned sound pickup stream of all seat areas and the virtual microphone. When two or more seat areas are confirmed to be overlapped in the same time window, the target seat area is determined by combining the envelope reference and beat key points of the main singer template and the leakage template. The occluded syllable segment of the target seat area is determined by combining the envelope reference and beat key points of the main singer template and the leakage template. The resonance peak is extracted in the neighborhood of the occluded syllable segment. The resonance gating map is formed in the frequency band and time. The side suppression mark is set at the corresponding lag position according to the leakage template of the adjacent seat area, which is used to trigger the subsequent leakage compensation. The baseline of the occluded syllable segment is recovered by taking the main singer template as the envelope reference. The unoccluded detail segment of the same seat area is re-injected in the time slice of the open state of the resonance gating map according to the level and boundary transition rules. When the side suppression mark is triggered, the leakage template matched with the side suppression mark is introduced for compensation and superposition to obtain the reconstruction result.

2. The multi-passenger seat area oriented in-vehicle wireless KTV sub-zone sound pickup and mixing method of claim 1, wherein, In step 1, after receiving the control instruction for triggering the partitioned sound pickup mode, the geometric parameters and seat area pointing parameter set of the microphone array are loaded. The sampling rate is set to 48000 Hz, the quantization bit depth is set to 24 bits, the frame length is set to 20 ms, the frame shift is set to 10 ms, and a unified time scale is established based on the main clock of the microphone array. The sampling frequency deviation of the accompaniment signal and the microphone array collection signal is controlled within ±50 parts per million by using a sampling rate fine-tuning resampling method, so that the two can be compared under the unified time scale. In order to reduce the overall time delay, the uplink channel target end-to-end delay budget of the partitioned sound pickup mode is set to not more than 80 ms.

3. The multi-passenger seat area oriented in-vehicle wireless KTV sub-zone sound pickup and mixing method of claim 2, wherein, The seat area reference coordinates are constructed and the in-vehicle fixed sound path is determined by the following process: taking the geometric center of the vehicle as the origin, the forward direction as the positive direction of the x-axis, the left side as the positive direction of the y-axis, and the upper side as the positive direction of the z-axis to construct the seat area reference coordinates; taking the center of each seat headrest as the reference, moving 70 mm along the x-axis, 120 mm along the z-axis, and the y-axis coordinate being located on the seat center line to obtain the mouth reference point coordinates; the driver seat, the front passenger seat, the rear left seat, the rear middle seat, and the rear right seat are sequentially labeled as seat area one to seat area five; the three-dimensional coordinates of each array pickup unit in the microphone array are recorded; the direct propagation path from the mouth reference point of each seat area to each array pickup unit and the reflection path generated on the reflecting surface are defined as the in-vehicle fixed sound path.

4. The multi-passenger seat area oriented in-vehicle wireless KTV sub-zone sound pickup and mixing method of claim 3, wherein, In step 2, the seat area reference coordinates, the mouth reference point coordinates of each seat area, the three-dimensional coordinates of each array pickup unit, and the spatial pointing direction are determined. The front windshield, the rear windshield, the left side window, the right side window, the instrument panel, the door trim panel, the roof lining, and the floor are determined as reflecting surfaces, and the seat backrest, the headrest, the instrument panel body, the central control channel, and the trunk partition are determined as non-penetrable bodies. The direct propagation path, the first reflection path, and the second reflection path are enumerated for each seat area and each array pickup unit. The paths are sampled at 10 mm intervals along the path. If any sampling point falls within the non-penetrable body outer boundary, the path is discarded. The paths are indexed in the order of arrival according to the geometric length from short to long, and the path type and reflecting surface sequence are recorded. The time delay is converted and aligned to an integer sampling point at a speed of 343 meters per second and a sampling rate of 48000 Hz.

5. The multi-passenger seat area oriented in-vehicle wireless KTV sub-zone sound pickup and mixing method of claim 4, wherein, In step 2, the process of generating the seat sound field topology map based on the fixed sound path in the vehicle includes: performing seat crossing detection on the reserved path, using a 50mm expansion in each direction for the non-native seat area surrounding boundary, and marking the leakage risk level as high if at least one crossing occurs; calculating the minimum distance of the path to the non-native seat area mouth reference point, and marking it as high if it is less than or equal to 0.25 meters, and marking it as medium if it is between 0.25 meters and 0.40 meters; calculating the angle between the path incident direction and the spatial pointing direction of the array pickup unit, and marking the leakage risk level as high if the non-native seat area path is less than or equal to 25 degrees, and marking the leakage risk level as medium if it is between 25 degrees and 50 degrees; additionally marking the secondary reflection path as medium in the leakage risk level; taking the mouth reference point and the array pickup unit as nodes, and taking the reserved path as edges, and adding the attributes of arrival index, path type, reflection surface sequence, delay sampling point and leakage risk level; organizing in layers within each seat area subgraph, first sorting by arrival index, then sorting by path type as direct propagation priority over one reflection, one reflection priority over two reflections, and sorting by leakage risk level from low to high within the same type, to obtain the seat sound field topology map.

6. The multi-passenger zone-oriented in-vehicle wireless KTV sub-zone sound pickup and mixing method of claim 5, wherein, In step 2, the process of forming a virtual microphone set based on the sound field topology map includes: selecting the array pickup unit with the smallest arrival index and the direct propagation path as the main channel under frame length 20 milliseconds and frame shift 10 milliseconds, and if the main channel short-time energy is less than the silence baseline by more than 6 decibel threshold, then sequentially selecting a direct propagation path with a smaller index, and if the direct propagation path does not meet the condition, then selecting a one reflection path; applying corresponding delay sampling point alignment to the selected channel, and using 5 milliseconds linear cross transition for channel switching; when the evaluated speech intelligibility meets the preset rule, selecting at most 2 one reflection paths with arrival index not greater than 3 and leakage risk level low or medium as support channels, and mixing them in with a gain of 6 decibels lower than the main channel; applying a fixed attenuation of 12 decibels to the non-native seat area incident direction with a high leakage risk level; pausing the support channel when the short-time energy is less than the silence baseline by more than 3 decibel threshold for 300 milliseconds; adding the direction to the side suppression direction set when the similarity to the adjacent seat area virtual microphone reaches 0.8 and lasts for more than 100 milliseconds; outputting at least one main virtual microphone for each seat area, and if there are two symmetric direct propagation paths with arrival index difference not greater than 2 and leakage risk level low, then additionally outputting one auxiliary virtual microphone.

Citation Information

Patent Citations

  • Vehicle-mounted multi-sound-zone audio processing system and method

    CN110475180A

  • Single-microphone multi-array pickup method and system based on sound attenuation simulation

    CN120416712A