A method for acoustic tuning of a large-screen voice front-end and related products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-08-11
AI Technical Summary
在实际使用过程中,大屏设备的内置扬声器或外接条形音响会持续播放节目声音、会议音频、背景音乐或提示音,该部分回放声会直接泄入麦克风阵列,形成较强的声学回声与串扰,同时墙面、玻璃、地面等硬质界面易产生明显混响,对远场语音采集造成明显干扰
本申请实施例提供的大屏语音前端分区声学调教方法,基于当前用户所在交互区域匹配相应的语音前端参数预设,对当前采集到的语音执行声学回声消除、波束成形与去混响处理,得到处理后的语音信号;所述语音前端参数预设包括声学回声消除参数、波束成形参数与去混响参数;所述语音前端参数预设是基于所述交互区域对应的语音前端调教描述量映射得到的;所述语音前端调教描述量是基于静态房间画像确定的;所述静态房间画像包括房间几何模型、材质声学属性参数集合、主动显示面模型、回放声源位置、麦克风阵列位置、交互区域集合以及所述交互区域集合中各交互区域对应的声学路径摘要;基于接收到的实际回放响应与静态房间画像中的声学路径摘要进行比较,根据比较结果对所述语音前端参数预设进行修正,以根据修正后的语音前端参数预设对后续采集的语音执行声学回声消除、波束成形与去混响处理。
Smart Images

Figure CN122551794A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio signal processing technology, and in particular to a method for acoustic tuning of a large-screen voice front-end and related products. Background Technology
[0002] With the widespread adoption of smart screens in living rooms, conference rooms, smart classrooms, and exhibition halls, far-field voice interaction based on microphone arrays has become the mainstream human-computer interaction method. In actual use, the built-in speakers or external soundbars of the large-screen devices continuously play program audio, conference audio, background music, or notification tones. This playback sound directly leaks into the microphone array, creating strong acoustic echoes and crosstalk. Simultaneously, hard surfaces such as walls, glass, and floors easily generate significant reverberation, causing considerable interference to far-field voice acquisition.
[0003] Existing large-screen voice front-end processing mainly relies on real-time audio signals during runtime for acoustic processing. However, this method is entirely dependent on iterative calculations of instantaneous audio signals, which is susceptible to interference from factors such as instantaneous noise and sudden changes in the sound field. This makes the voice front-end parameters extremely volatile, and the fixed parameter configuration mode is often used, which cannot adapt to the complex acoustic characteristics of different interaction areas. Problems such as insufficient voice gain, incomplete echo suppression, and severe reverberation residue are very likely to occur, resulting in poor sound pickup stability and seriously affecting the far-field voice interaction experience.
[0004] How to suppress fluctuations in voice front-end parameters, improve the stability of voice pickup, and ensure a good far-field voice interaction experience in scenarios with continuous playback, spatial reverberation, and differences in regional acoustic characteristics is an urgent technical problem to be solved. Summary of the Invention
[0005] To address the aforementioned issues, this application provides a method for acoustic tuning of a large-screen voice front-end and related products. The aim is to suppress fluctuations in voice front-end parameters, improve the stability of sound pickup, and ensure a good far-field voice interaction experience in scenarios with continuous playback, spatial reverberation, and differences in regional acoustic characteristics.
[0006] The embodiments of this application disclose the following technical solutions: The first aspect of this application provides a method for partitioned acoustic tuning of a large-screen voice front-end, the method comprising: Based on the corresponding preset voice front-end parameters matched to the current user's interaction area, acoustic echo cancellation, beamforming, and dereverberation processing are performed on the currently acquired voice signal to obtain the processed voice signal. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interaction area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, the location of the playback sound source, the location of the microphone array, a set of interaction areas, and an acoustic path summary corresponding to each interaction area in the set of interaction areas. Based on the comparison between the received actual playback response and the acoustic path summary in the static room image, the preset voice front-end parameters are corrected according to the comparison results, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
[0007] Optionally, the method for obtaining the static room image includes: A geometric model of the room is constructed based on the scanned image data of the room; the image data includes color images and depth maps. Using the image data, semantic segmentation, object recognition, and material classification are performed on non-screen static areas. If the confidence value of the material classification result exceeds a preset confidence threshold, a set of material acoustic attribute parameters is obtained through probability-weighted material acoustic mapping. If the confidence value of the material classification result does not exceed the preset confidence threshold, a rollback process is performed based on room type, object category, default material table, or user annotation. An active display surface model is established based on the interactive area where the large screen is located, and the parameters of the active display surface model are determined. The active display surface model parameters include screen plane, screen normal, screen boundary, screen mask, screen reflection properties, playback sound source position parameters, and frequency band coupling parameters. The interactive areas are divided according to the room geometry model, the set of material acoustic properties parameters, the active display surface model parameters, the playback sound source position and the microphone array position to obtain an interactive area set, and the corresponding acoustic path summary is determined based on each interactive area in the interactive area set. Based on the room geometry model, the set of material acoustic properties parameters, the parameters of the active display surface model, the location of the playback sound source, the location of the microphone array, the set of interactive areas, and the corresponding acoustic path summary, a static room profile is determined.
[0008] Optionally, constructing a room geometric model based on the scanned room image data includes: Identify multiple objects in the image data to obtain a first object set; The first set of objects is filtered based on multidimensional conditions to obtain a second set of objects; the multidimensional conditions include at least two of the following: semantic recognition, occupation persistence, motion consistency, rigid anchoring relationship or historical position consistency. Remove the movable objects from the second object set to obtain the third object set; the movable objects include people, pets, chairs, wheelchairs, and strollers. Based on the aforementioned third set of objects, a room geometric model is constructed.
[0009] Optionally, determining the corresponding acoustic path summary based on each interaction region in the set of interaction regions includes: A dual-path acoustic estimation model is used to determine the corresponding acoustic path summary based on each interaction region in the set of interaction regions. The dual-path acoustic estimation model includes a playback path model and a user voice path model. The acoustic path summary includes a playback path summary and a user voice path summary. The playback path model is used to obtain the playback path summary based on the propagation risk of the playback signal from the built-in speaker or external soundbar of the display device entering the microphone array through the active display surface model, the main reflective surface, and free space. The playback path summary includes the frequency band playback leakage risk, maximum echo delay, echo tail length, frequency band leakage mask, and null candidate direction to configure acoustic echo cancellation parameters. The user voice path model is used to obtain the user voice path summary based on the direct path, occlusion path, and reflection path of the user voice to the microphone array within the interaction region. The user voice path summary includes the direct path visibility, direct sound gain, strong reflection risk, late reverberation risk, and microphone channel reliability to configure beamforming parameters and dereverberation parameters.
[0010] Optionally, the step of comparing the received actual playback response with the acoustic path summary in the static room profile, and correcting the preset speech front-end parameters based on the comparison result, includes: Within a preset time period, the actual playback response is determined based on the playback reference signal and the actual playback response received by the microphone array. Calculate the response deviation between the actual playback response and the acoustic path summary; Based on the response deviation, the preset speech front-end parameters are boundedly corrected according to a preset step size to obtain the corrected preset speech front-end parameters; the correction includes correcting the echo cancellation filter length, correcting the residual echo suppression intensity, correcting the dereverberation intensity, or correcting the channel weight.
[0011] Optionally, the static room image further includes a room signature, and the method further includes: The real-time room signature is compared with the historical room signature. If the difference between the real-time room signature and the historical room signature is less than a first preset threshold, the preset voice front-end parameters of each interaction area are directly reused. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The room signature includes screen pose, main plane parameters, furniture surround box, material histogram, free space, playback sound source, and microphone layout. If the difference is between the first preset threshold and the second preset threshold, then a local incremental update is performed; If the difference exceeds the second preset threshold, a prompt will be made to rescan the room; the first preset threshold is less than the second preset threshold.
[0012] Optionally, the method for obtaining the current user's interaction area includes: Based on at least one of the following: visual human detection, sound source arrival direction, wake word localization, historical wake-up area, application mode, or user manual selection, determine one or more candidate interaction areas where the current user is located. If there is a candidate interaction region, then the candidate region is taken as the target interaction region; if there are multiple candidate interaction regions, then the confidence values of the multiple candidate interaction regions are fused by region confidence, and the target interaction region is obtained based on the fusion result. The target interaction area is compared with the historical interaction area. If the comparison result indicates that the target interaction area and the historical interaction area are consistent, the target interaction area is output. If the comparison result indicates that the target interaction area and the historical interaction area are inconsistent, a region switching anti-shake verification is triggered. The region switching anti-shake verification is as follows: if the continuous effective duration of the target interaction area exceeds a preset duration, the region switching is effective and the target interaction area is output. If the continuous effective duration of the target interaction area does not exceed the preset duration, the region switching is invalid and the historical interaction area is output.
[0013] Optionally, the preset mapping method for the voice front-end parameters includes: The voice front-end tuning descriptors include echo descriptors, channel reliability descriptors, and reverberation descriptors; the echo descriptors include playback energy, echo delay, and echo tail duration; the channel reliability descriptors include microphone channel signal-to-noise ratio and channel reliability index; the reverberation descriptors include late reverberation duration and speech signal-to-noise ratio; Acoustic echo cancellation parameters are generated using the echo class descriptor mapping; the acoustic echo cancellation parameters include acoustic echo cancellation filter length, residual echo suppression strength, and leakage risk mask; Beamforming parameters are generated by mapping the channel reliability class descriptor; the beamforming parameters include the number of effective microphone channels, beam weight, and beam operating frequency band. De-reverberation parameters are generated using the reverberation class descriptor mapping; the de-reverberation parameters include de-reverberation intensity, late reverberation time-domain parameters, and post-filter coefficients.
[0014] A second aspect of this application provides a large-screen voice front-end zone acoustic tuning device, the device comprising: The region-adaptive speech processing module is used to match corresponding preset speech front-end parameters based on the current user's interaction region, and perform acoustic echo cancellation, beamforming, and dereverberation processing on the currently acquired speech to obtain the processed speech signal. The preset speech front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset speech front-end parameters are obtained based on the mapping of the speech front-end tuning descriptor corresponding to the interaction region. The speech front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, the location of the playback sound source, the location of the microphone array, a set of interaction regions, and an acoustic path summary corresponding to each interaction region in the set of interaction regions. The voice front-end parameter correction module is used to compare the received actual playback response with the acoustic path summary in the static room image, and correct the preset voice front-end parameters according to the comparison result, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
[0015] A third aspect of this application provides an electronic device that includes a memory, a processor, a microphone array interface, a playback reference signal interface, and an optional image acquisition interface.
[0016] Compared with the prior art, this application has the following beneficial effects: The large-screen voice front-end partitioned acoustic tuning method provided in this application embodiment matches corresponding preset voice front-end parameters based on the current user's interaction area, and performs acoustic echo cancellation, beamforming, and dereverberation processing on the currently acquired voice signal to obtain a processed voice signal. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interaction area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, playback sound source location, microphone array location, a set of interaction areas, and an acoustic path summary corresponding to each interaction area in the set of interaction areas. The received actual playback response is compared with the acoustic path summary in the static room profile, and the preset voice front-end parameters are corrected according to the comparison result, so that acoustic echo cancellation, beamforming, and dereverberation processing are performed on the subsequently acquired voice signal based on the corrected preset voice front-end parameters.
[0017] The aforementioned method relies on a static room profile containing information on geometry, acoustic materials, microphone layout, interaction areas, and playback paths to accurately generate a descriptive quantity for the corresponding interaction area's voice front-end tuning and map it to a dedicated voice front-end parameter preset. This abandons a uniform, universal parameter model, allowing for targeted processing tailored to the actual acoustic environment of different interaction areas, adapting to different pickup conditions from the source. Based on the user's real-time interaction area, a dedicated parameter preset is dynamically matched to perform directional acoustic enhancement processing on the collected speech, effectively offsetting inherent acoustic differences in echo intensity, sound wave reflection paths, and spatial reverberation between different interaction areas, significantly reducing the pickup quality gap across the entire area. By comparing the actual playback response with the standard playback path summary in the static room profile, subtle sound field shifts are identified in real time, and the voice front-end parameter preset is dynamically corrected. This effectively compensates for acoustic characteristic deviations caused by minor environmental changes or device posture shifts, avoiding acoustic processing failures due to parameter fixation and continuously maintaining the stable operation of the voice front-end processing. By combining static partitioning and customized parameters with runtime dynamic correction, the system effectively avoids issues such as sudden changes in sound quality or fluctuations in signal-to-noise ratio during user location switching. This ensures that the voice clarity, noise reduction, and echo suppression effects are consistent across the entire interactive area, significantly improving the overall stability and scene adaptability of the voice front-end processing and effectively optimizing the far-field voice interaction experience on large-screen devices. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for partitioned acoustic tuning of a large-screen voice front-end, as provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a large-screen voice front-end partitioned acoustic tuning device provided in an embodiment of this application. Detailed Implementation
[0020] As described above, existing large-screen voice front-end processing mainly relies on real-time audio signals during runtime for acoustic processing. However, this method relies entirely on iterative calculations of instantaneous audio signals, which is susceptible to interference from factors such as instantaneous noise and sudden changes in the sound field. This makes the voice front-end parameters extremely volatile, and the fixed parameter configuration mode is often used, which cannot adapt to the complex acoustic characteristics of real-world scenarios. This can easily lead to problems such as insufficient human voice gain, incomplete echo suppression, and severe reverberation residue, resulting in poor sound pickup stability and seriously affecting the far-field voice interaction experience.
[0021] In view of the above problems, this application proposes a large-screen voice front-end zonal acoustic tuning method and related products. Based on the current user's interaction area, corresponding preset voice front-end parameters are matched, and acoustic echo cancellation, beamforming, and dereverberation processing are performed on the currently acquired voice signal to obtain a processed voice signal. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interaction area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, playback sound source location, microphone array location, a set of interaction areas, and acoustic path summaries corresponding to each interaction area in the set of interaction areas. The received actual playback response is compared with the acoustic path summaries in the static room profile, and the preset voice front-end parameters are corrected according to the comparison results. The corrected preset voice front-end parameters are then used to perform acoustic echo cancellation, beamforming, and dereverberation processing on subsequently acquired voice signals.
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0023] See Figure 1 This figure is a flowchart of a large-screen voice front-end zone acoustic tuning method provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps: S101. Based on the current user's interaction area, match the corresponding preset voice front-end parameters, and perform acoustic echo cancellation, beamforming and dereverberation processing on the currently collected voice signal to obtain the processed voice signal.
[0024] The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interactive area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, playback sound source location, microphone array location, an interactive area set, and an acoustic path summary corresponding to each interactive area in the interactive area set.
[0025] The preset voice front-end parameters are a set of configuration parameters directly used for voice processing, including acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters.
[0026] Voice front-end tuning descriptors refer to intermediate feature quantities calculated from static room profiles and used to map voice front-end parameters, including echo risk, direct path visibility, reverberation intensity, and channel reliability.
[0027] For example, the formula for the speech front-end tuning descriptor FED_k is: FED_k={V_km, E_play, k, m(b), R_speech, k(b), R_play, k(b), T_late, k(b), C_km(b), Tail_echo, k, m, N_play, k, τ_echo, k, m}; Wherein, V_km is obtained by the overlap ratio of the ray between the interaction region and the microphone with the free space voxel; E_play,k,m(b) represents the playback signal leakage risk or echo intensity estimate of the m-th microphone in frequency band b under the interaction region k; R_speech,k(b) and R_play,k(b) describe the user voice reflection risk and playback sound reflection risk, respectively; T_late,k(b) is determined based on room volume, fixed surface material distribution, reflection density and free space ratio estimation; C_km(b) reflects the occlusion of each channel, screen coupling strength, noise level and calibration stability; Tail_echo,k,m represents the tail decay time of the echo signal received by the m-th microphone under the interaction region k, used to configure echo cancellation and déreverberation intensity; N_play,k represents the null candidate direction associated with the active display surface model, external soundbar or strong reflective surface, used to configure beam null or sidelobe control.
[0028] In one feasible implementation, the preset mapping method for the voice front-end parameters includes: The voice front-end tuning descriptors include echo descriptors, channel reliability descriptors, and reverberation descriptors; the echo descriptors include playback energy, echo delay, and echo tail duration; the channel reliability descriptors include microphone channel signal-to-noise ratio and channel reliability index; the reverberation descriptors include late reverberation duration and speech signal-to-noise ratio; Acoustic echo cancellation parameters are generated using the echo class descriptor mapping; the acoustic echo cancellation parameters include acoustic echo cancellation filter length, residual echo suppression strength, and leakage risk mask; Beamforming parameters are generated by mapping the channel reliability class descriptor; the beamforming parameters include the number of effective microphone channels, beam weight, and beam operating frequency band. De-reverberation parameters are generated using the reverberation class descriptor mapping; the de-reverberation parameters include de-reverberation intensity, late reverberation time-domain parameters, and post-filter coefficients.
[0029] For example, FED_k is mapped to the acoustic front-end parameter preset P_k for the corresponding interaction area. P_k is structured into the acoustic echo cancellation parameter AECProfile_k, the beamforming parameter BeamProfile_k, and the dereverberation parameter DRVProfile_k, as shown in the following formula: AECProfile_k={L_AEC,k,μ_AEC,k,G_RES,k,H_init,k,DTD_th,k,FreezePolicy_k,LeakageMask_k(b)}; Wherein, L_AEC,k represents the length of the Acoustic Echo Cancellation (AEC) adaptive filter, which can be determined based on the maximum echo delay and safety margin; μ_AEC,k represents the AEC update step size, which can be set based on playback risk, dual-talk risk, and near-end speech confidence; G_RES,k represents the residual echo suppression strength; H_init,k represents the optional echo path initial prior, used to reduce the time for AEC to converge from zero; DTD_th,k represents the region-related dual-talk detection threshold; FreezePolicy_k represents the AEC update freeze strategy during strong dual-talk or strong playback; LeakageMask_k(b) represents the band-related playback leakage risk mask.
[0030] BeamProfile_k={θ_k,ω_k,N_k,M_k,W_k(b),SidelobeCtrl_k}; Wherein, θ_k represents the beam pointing angle, which can be calculated based on the center of the interactive area and the center of the microphone array; ω_k represents the beamwidth, which can be set according to the area angle width, multi-person coverage requirements, and direct path stability; N_k represents the null direction set, including the main sound direction of the screen, the direction of the external bar speaker, the direction of strong reflection, or the direction of continuous noise; M_k represents the set of channels participating in beamforming; W_k(b) represents the channel weights set by frequency band; SidelobeCtrl_k represents the sidelobe suppression strategy for the screen side, glass wall side, or other strong reflection directions.
[0031] DRVProfile_k={λ_DRV,k,τ_late,k,R_late,k(b),PostFilter_k}; Where λ_DRV,k represents the de-reverberation intensity; τ_late,k represents the late reverberation onset time; R_late,k(b) represents the band-dependent late reverberation risk; and PostFilter_k represents the post-filtering parameters.
[0032] For example, FED_k is mapped to the acoustic front-end parameter preset P_k of the corresponding interactive area to complete the conversion from acoustic feature quantization value to engineering configuration parameter. The mapping relationship is shown in Table 1.
[0033] Table 1
[0034] Through the above mapping process, the front-end tuning description is completely converted into preset voice front-end parameters applicable to each interactive area, enabling the voice front-end processing module to directly complete the initial configuration based on the static room profile, thereby improving the stability and consistency of far-field voice pickup.
[0035] In one feasible implementation, the method for obtaining a static room image includes: A geometric model of the room is constructed based on the scanned image data of the room; the image data includes color images and depth maps. Using the image data, semantic segmentation, object recognition, and material classification are performed on non-screen static areas. If the confidence value of the material classification result exceeds a preset confidence threshold, a set of material acoustic attribute parameters is obtained through probability-weighted material acoustic mapping. If the confidence value of the material classification result does not exceed the preset confidence threshold, a rollback process is performed based on room type, object category, default material table, or user annotation. An active display surface model is established based on the interactive area where the large screen is located, and the parameters of the active display surface model are determined. The active display surface model parameters include screen plane, screen normal, screen boundary, screen mask, screen reflection properties, playback sound source position parameters, and frequency band coupling parameters. The interactive areas are divided according to the room geometry model, the set of material acoustic properties parameters, the active display surface model parameters, the playback sound source position and the microphone array position to obtain an interactive area set, and the corresponding acoustic path summary is determined based on each interactive area in the interactive area set. Based on the room geometry model, the set of material acoustic properties parameters, the parameters of the active display surface model, the location of the playback sound source, the location of the microphone array, the set of interactive areas, and the corresponding acoustic path summary, a static room profile is determined.
[0036] For example, the set of material acoustic property parameters is as follows: Γ_i(b) = Σ_c p_i(c)·Λ(c,b); Where i represents a surface unit; b represents a frequency band; c represents a material category; p_i(c) represents the probability that surface unit i belongs to material category c; Λ(c, b) represents the acoustic property mapping table of material category c in frequency band b. Γ_i(b) includes at least the absorption coefficient α_i(b), the reflection coefficient ρ_i(b), and the scattering coefficient σ_i(b), and may also include the transmission coefficient τ_i(b) if necessary. When the confidence level of material identification is low, a fallback can be performed based on room type, object category, default material table, or user annotation.
[0037] An active display surface, unlike a typical wall or glass surface, represents the plane of the display device. It possesses three attributes: a display area, a passive reflective surface, and an active playback coupler. The large screen serves as both an interactive interface and a geometrical reflective boundary, and together with built-in speakers or external soundbars, constitutes the primary source of playback interference entering the microphone array.
[0038] For example, the active display surface model can be represented as: ADS={Π_scr, n_scr, edge_scr, mask_scr, r_scr(b), Sset_scr, η_scr, m(b), H_play, m(b, τ)}; Wherein, Π_scr represents the screen plane; n_scr represents the screen normal; edge_scr represents the screen boundary; mask_scr represents the screen mask, used to exclude ordinary material recognition; r_scr(b) represents the passive reflection property of the screen; Sset_scr represents the set of positions of the built-in speaker or external soundbar; η_scr,m(b) represents the frequency band correlation coupling parameter from the active display surface or playback sound source to the m-th microphone channel; H_play,m(b,τ) represents the optional playback path summary, used for AEC initialization, filter length, and residual echo suppression.
[0039] By actively modeling the display surface, the screen playback leakage risk, echo tail length, null candidate direction, and channel coupling strength can be calculated based on the screen plane, speaker, external sound bar location, screen reflection characteristics, microphone array location, and interaction area relationship.
[0040] To address the relatively stable user activity areas in static large-screen scenarios, an interaction region-level modeling strategy is adopted to achieve refined zoning acoustic tuning. For example, as shown in Table 2, Table 2 illustrates the interaction region divisions in various scenarios.
[0041] Table 2
[0042] By dividing the interaction areas as described above, a dedicated acoustic model and preset voice front-end parameters can be established for each interaction area, providing a foundation for subsequent adaptive acoustic processing based on user location.
[0043] The interactive area can be generated automatically, manually, or in a hybrid manner. The automatic method derives the area based on the screen plane, free space on the ground, fixed furniture boundaries, and typical standing distances; the manual method allows the user to confirm or adjust the area via the interface after scanning; the hybrid method first automatically provides suggested areas, which the user then corrects.
[0044] In one feasible implementation, a room geometric model is constructed based on scanned image data of the room, including: Identify multiple objects in the image data to obtain a first object set; The first set of objects is filtered based on multidimensional conditions to obtain a second set of objects; the multidimensional conditions include at least two of the following: semantic recognition, occupation persistence, motion consistency, rigid anchoring relationship or historical position consistency. Remove the movable objects from the second object set to obtain the third object set; the movable objects include people, pets, chairs, wheelchairs, and strollers. Based on the aforementioned third set of objects, a room geometric model is constructed.
[0045] The first object set here represents the collection of various objects within a room detected from RGB-D image data through image recognition. This includes a set of fixed structures, a set of permanently fixed equipment, a set of dynamic living entities, and a set of dynamic, uncertain objects. The set of fixed structures includes building structures such as floors, ceilings, walls, columns, and fixed partitions; the set of permanently fixed equipment includes large screens, TV cabinets, fixed conference tables, fixed sound-absorbing panels, and fixed display stands; the set of dynamic living entities includes people and pets; and the set of dynamic, uncertain objects includes movable chairs, folding chairs, office chairs with casters, portable dining chairs, portable children's chairs, suitcases, portable whiteboards, trolleys, temporary display stands, and cleaning tools. Here, RGB-D indicates that color images and depth maps were jointly acquired. The multidimensional criteria here are used to filter stable objects across multiple dimensions, including at least two of semantic recognition, occupancy persistence, motion consistency, rigid anchoring relationships, or historical position consistency. Semantic recognition provides prior knowledge of the object category; motion consistency analysis determines whether the object has moved based on optical flow, depth difference, and instance pose changes in multi-frame RGB-D sequences; occupancy persistence estimates the proportion of stable occurrences of voxels or surface units across multiple time slices, multiple viewpoints, or multiple scan cycles; rigid anchoring relationships identify whether the object has a long-term connection to the ground, walls, or fixed components; and historical consistency determines whether the object's position is stable in historical room portraits.
[0046] The second set of objects here refers to the set of objects that are retained after multi-dimensional filtering, which have high stability and are likely to belong to a fixed structure or be placed for a long time.
[0047] Movable objects here refer to objects whose location is easily changed and are not suitable for inclusion in static room modeling, including human bodies, pets, folding chairs, wheeled office chairs, suitcases, mobile whiteboards, trolleys, and temporary display stands. Even if they remain stationary during a single scan, they are by default classified as temporary / movable objects and are not used as fixed boundaries for scene reconstruction and acoustic modeling; they are only allowed to be included in the fixed structure set when they are identified as fixed-installation seats, fixed rows of seats, or explicitly marked as fixed components by the user.
[0048] The third object set here refers to the final object set that, based on the second object set, retains only the room's fixed structure and long-term fixed equipment after removing all movable objects.
[0049] For example, as shown in Table 3, Table 3 illustrates the classification of scanned indoor objects and the use of differentiated modeling and retention strategies for different categories.
[0050] Table 3
[0051] Through the above classification and modeling strategies, dynamic interference objects can be filtered out, and only long-term stable acoustic influencing elements can be retained. This constructs a static room profile that combines stability and scene adaptability, providing reliable spatial prior support for subsequent interactive area division, parameter preset mapping, and playback path verification.
[0052] The room geometry model can be represented using point clouds, 3D occupancy raster maps, voxel meshes, truncated signed distance function (TSDF) reconstruction results, or 3D mesh models. In a preferred embodiment, denoising, distortion correction, and coordinate transformation are first performed on the depth map, and then the static regions are fused into a point cloud; subsequently, a representation usable for geometric analysis and path calculation is obtained through voxelization, TSDF fusion, or mesh generation.
[0053] G_static={Π_scr, Π_floor, Π_ceiling, Π_wall_j, B_furniture_l, Mesh_static, Occ_free, Occ_block, Sset, Mset}; Wherein, Π_scr represents the screen plane; Π_floor represents the ground; Π_ceiling represents the top surface; Π_wall_j represents the set of walls; B_furniture_l represents the fixed furniture enclosure box; Mesh_static represents the static background mesh; Occ_free represents the free space voxel; Occ_block represents the fixed occlusion voxel; Sset represents the set of playback sound source locations; and Mset represents the set of microphone array element locations.
[0054] In one feasible implementation, based on each interactive region in the set of interactive regions, a corresponding acoustic path summary is determined, including: A dual-path acoustic estimation model is used to determine the corresponding acoustic path summary based on each interaction region in the set of interaction regions. The dual-path acoustic estimation model includes a playback path model and a user voice path model. The acoustic path summary includes a playback path summary and a user voice path summary. The playback path model is used to obtain the playback path summary based on the propagation risk of the playback signal from the built-in speaker or external soundbar of the display device entering the microphone array through the active display surface model, the main reflective surface, and free space. The playback path summary includes the frequency band playback leakage risk, maximum echo delay, echo tail length, frequency band leakage mask, and null candidate direction to configure acoustic echo cancellation parameters. The user voice path model is used to obtain the user voice path summary based on the direct path, occlusion path, and reflection path of the user voice to the microphone array within the interaction region. The user voice path summary includes the direct path visibility, direct sound gain, strong reflection risk, late reverberation risk, and microphone channel reliability to configure beamforming parameters and dereverberation parameters.
[0055] For example, the formula for the playback path model is defined as follows: H_play,k,m={E_play,k,m(b),τ_echo,k,m,Tail_echo,k,m,Leakage_k,m(b),N_play,k}; Wherein, E_play,k,m(b) represents the leakage risk or echo intensity estimate of the playback signal of the m-th microphone in frequency band b under the interaction region k, used to adjust residual echo suppression and channel weights; τ_echo,k,m represents the maximum propagation delay of the playback signal received by the m-th microphone under the interaction region k; Tail_echo,k,m represents the tail decay duration of the echo signal received by the m-th microphone under the interaction region k, used to configure echo cancellation and dereverberation intensity; Leakage_k,m(b) represents the leakage intensity of the playback signal of the m-th microphone in frequency band b under the interaction region k, used to generate a frequency band leakage risk mask; N_play,k represents the null candidate direction associated with the active display surface model, external bar speaker, or strong reflective surface, used to configure beam null or sidelobe control.
[0056] For example, the formula for the user voice path model is defined as follows: H_speech, k, m = {V_km, G_direct, k, m, R_speech, k, m(b), T_late, k(b), C_km(b)}; Wherein, V_km represents the direct path visibility from interaction region k to the m-th microphone, used for channel selection and beamwidth configuration; G_direct,k,m represents the direct sound estimation gain of the m-th microphone in interaction region k, used for channel weights and array gain settings; R_speech,k,m(b) represents the strong reflection risk of user speech at the m-th microphone in frequency band b in interaction region k, used for null candidate and sidelobe control; T_late,k(b) represents the local late reverberation risk in frequency band b in interaction region k, used for déreverberation intensity and post-filter configuration; C_km(b) represents the microphone channel reliability of the m-th microphone in frequency band b in interaction region k, used for channel set and frequency band weight configuration.
[0057] Through the aforementioned dual-path acoustic estimation model, the spatial and acoustic information in the static room profile is directly mapped to the configuration parameters of the acoustic echo cancellation, beamforming, and dereverberation modules, achieving precise tuning based on scene priors: on the one hand, the playback path model effectively suppresses echo and playback leakage, and on the other hand, the user voice path model enhances the target speech and suppresses reflection and reverberation. The two work together to significantly improve the clarity and stability of far-field voice interaction.
[0058] In one feasible implementation, the static room portrait further includes a room signature, and the method further includes: The real-time room signature is compared with the historical room signature. If the difference between the real-time room signature and the historical room signature is less than a first preset threshold, the preset voice front-end parameters of each interactive area are directly reused. The room signature includes screen pose, main plane parameters, furniture bounding box, material histogram, free space, playback sound source and microphone layout. If the difference is between the first preset threshold and the second preset threshold, then a local incremental update is performed; If the difference exceeds the second preset threshold, a prompt will be made to rescan the room; the first preset threshold is less than the second preset threshold.
[0059] Room signatures are feature summaries used to quickly verify whether a room has changed. These summaries include screen pose, master plane parameters, furniture bounding boxes, material histograms, free space, sound source and microphone layout, etc. At a minimum, they include screen plane pose, floor / wall / ceiling master plane parameters, main fixed furniture bounding boxes, material histograms, free space descriptions, playback sound source layout summaries, and microphone array installation summaries.
[0060] The real-time room signature here is the current room signature calculated in real time using a small number of RGB-D frames when the device is powered on or the application is started.
[0061] The historical room signatures here are the room signatures that are persistently stored after the scanning phase is completed.
[0062] If the difference between the real-time room signature and the historical room signature is less than the first preset threshold, it indicates that the room layout and acoustic environment have not changed significantly, and the preset voice front-end parameters of each interaction area are directly reused; if the difference is between the first preset threshold and the second preset threshold, it indicates that the room has only changed locally, and the system performs a local incremental update; if the difference exceeds the second preset threshold, it indicates that the room layout or acoustic environment has changed significantly, and the user is prompted to re-perform the room scan to update the static room profile and the preset voice front-end parameter library.
[0063] For example, the incremental updates are tailored to the type of change. For instance, when the display device moves slightly, the screen plane, active display surface model, playback sound source location, and related AEC / Beam parameters are updated; when the location of the conference table or fixed booth changes, the fixed furniture boundaries, interactive areas, and strong reflection path candidates are updated; when the curtains open or close or the material area changes, the local material properties, reverberation risk, and déreverberation parameters are updated; when the external soundbar is moved or replaced, the playback sound source layout, playback path summary, and AEC parameters are updated; and when the local display is adjusted, the free space, reflection candidates, and related area parameter presets are updated.
[0064] Through room signature verification and a hierarchical update mechanism, adaptive maintenance of voice front-end parameters is achieved. On the one hand, historical parameter presets can be quickly reused with a small number of RGB-D verification frames, significantly shortening application startup time and improving user experience. On the other hand, a hierarchical processing of "direct reuse - partial update - rescanning" is implemented through two-level thresholds. Targeted incremental updates are performed on local changes in the scene, avoiding unnecessary full rescanning, reducing computational overhead and user intervention costs, and providing timely prompts for recalibration when significant scene changes occur, ensuring the accuracy of acoustic processing. This mechanism effectively improves the system's robustness to dynamic changes in real-world scenes while reducing the frequency of image data acquisition, balancing efficiency, acoustic performance, and user privacy protection.
[0065] In one feasible implementation, the method for obtaining the current user's interaction area includes: Based on at least one of the following: visual human detection, sound source arrival direction, wake word localization, historical wake-up area, application mode, or user manual selection, determine one or more candidate interaction areas where the current user is located. If there is a candidate interaction region, then the candidate region is taken as the target interaction region; if there are multiple candidate interaction regions, then the confidence values of the multiple candidate interaction regions are fused by region confidence, and the target interaction region is obtained based on the fusion result. The target interaction area is compared with the historical interaction area. If the comparison result indicates that the target interaction area and the historical interaction area are consistent, the target interaction area is output. If the comparison result indicates that the target interaction area and the historical interaction area are inconsistent, a region switching anti-shake verification is triggered. The region switching anti-shake verification is as follows: if the continuous effective duration of the target interaction area exceeds a preset duration, the region switching is effective and the target interaction area is output. If the continuous effective duration of the target interaction area does not exceed the preset duration, the region switching is invalid and the historical interaction area is output.
[0066] Based on the current user's interaction area or application mode, the corresponding parameter presets are selected to perform acoustic echo cancellation, beamforming, and dereverberation processing. For acoustic echo cancellation, the playback reference signal is used as a reference, and parameters such as L_AEC,k, μ_AEC,k, G_RES,k, H_init,k, and DTD_th,k in AECProfile_k are used to achieve echo estimation, update control, and residual echo suppression. For beamforming, beam pointing, width control, null setting, and channel selection are performed based on θ_k, ω_k, N_k, M_k, and W_k(b) in BeamProfile_k. For dereverberation, local reverberation components are suppressed based on λ_DRV,k, τ_late,k, R_late,k(b), and PostFilter_k in DRVProfile_k.
[0067] S102. Based on the comparison between the received actual playback response and the acoustic path summary in the static room image, the preset voice front-end parameters are corrected according to the comparison result, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
[0068] In one feasible implementation: Within a preset time period, the actual playback response is determined based on the playback reference signal and the actual playback response received by the microphone array. Calculate the response deviation between the actual playback response and the acoustic path summary; Based on the response deviation, the preset speech front-end parameters are boundedly corrected according to a preset step size to obtain the corrected preset speech front-end parameters; the correction includes correcting the echo cancellation filter length, correcting the residual echo suppression intensity, correcting the dereverberation intensity, or correcting the channel weight.
[0069] The preset time period here includes time periods with no near-end voice or a low probability of two-way communication.
[0070] For example, during periods with no near-end speech or low probability of dual-talk, the actual playback response h_est is estimated using the playback reference signal and the actual playback response received by the microphone array. This estimate is then compared with the acoustic path summary predicted by the static room profile to make small-step adjustments to η_scr,m(b), E_play,k,m(b), L_AEC,k, G_RES,k, λ_DRV,k, or relevant channel weights. If the actual playback response deviates from the prediction result of the static room profile for an extended period, the room profile quality index is reduced, triggering a local re-examination or rescan.
[0071] The large-screen voice front-end partitioned acoustic tuning method provided in this application embodiment matches corresponding preset voice front-end parameters based on the current user's interaction area, and performs acoustic echo cancellation, beamforming, and dereverberation processing on the currently acquired voice signal to obtain a processed voice signal. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interaction area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, playback sound source location, microphone array location, a set of interaction areas, and an acoustic path summary corresponding to each interaction area in the set of interaction areas. The received actual playback response is compared with the acoustic path summary in the static room profile, and the preset voice front-end parameters are corrected according to the comparison result, so that acoustic echo cancellation, beamforming, and dereverberation processing are performed on the subsequently acquired voice signal based on the corrected preset voice front-end parameters.
[0072] The aforementioned method relies on a static room profile containing information on geometry, acoustic materials, microphone layout, interaction areas, and playback paths to accurately generate a descriptive quantity for the corresponding interaction area's voice front-end tuning and map it to a dedicated voice front-end parameter preset. This abandons a uniform, universal parameter model, allowing for targeted processing tailored to the actual acoustic environment of different interaction areas, adapting to different pickup conditions from the source. Based on the user's real-time interaction area, a dedicated parameter preset is dynamically matched to perform directional acoustic enhancement processing on the collected speech, effectively offsetting inherent acoustic differences in echo intensity, sound wave reflection paths, and spatial reverberation between different interaction areas, significantly reducing the pickup quality gap across the entire area. By comparing the actual playback response with the standard playback path summary in the static room profile, subtle sound field shifts are identified in real time, and the voice front-end parameter preset is dynamically corrected. This effectively compensates for acoustic characteristic deviations caused by minor environmental changes or device posture shifts, avoiding acoustic processing failures due to parameter fixation and continuously maintaining the stable operation of the voice front-end processing. By combining static partitioning and customized parameters with runtime dynamic correction, the system effectively avoids issues such as sudden changes in sound quality or fluctuations in signal-to-noise ratio during user location switching. This ensures that the speech clarity, noise reduction, and echo suppression effects are consistent across the entire interactive area, significantly improving the overall stability and scene adaptability of the voice front-end processing and effectively optimizing the far-field voice interaction experience in complex indoor scenarios.
[0073] The large-screen voice front-end partitioned acoustic tuning method described in the above embodiments can achieve stable far-field voice interaction in a variety of typical scenarios. The application of this application is illustrated below with three examples.
[0074] Taking the large TV screen in the exhibition hall as an example, after the exhibition hall setup is completed, a room scan is performed: the large screen displays a scan pattern to identify the screen plane and boundaries; the RGB-D acquisition device scans the space in front of the exhibition hall, eliminating guides, visitors, and temporary exhibition carts, retaining only the large screen, walls, exhibition stands, and fixed partitions, thus constructing a static room profile. The static room profile includes the room's geometric model, a set of material acoustic property parameters, an active display surface model, the location of playback sound sources, the location of microphone arrays, a set of interactive areas, and acoustic path summaries corresponding to each interactive area, such as the guide standing area, the near-field visitor area, and the far-field viewing area.
[0075] Based on a static room profile, a voice front-end tuning descriptor is generated for each interactive area. This descriptor is then mapped to corresponding preset voice front-end parameters. The preset parameters are matched to the current user's interactive area, and acoustic echo cancellation, beamforming, and dereverberation processing are applied to the acquired speech signal to obtain the processed audio signal. Subsequently, the received actual playback response is compared with the acoustic path summary in the static room profile. Based on the comparison result, the preset voice front-end parameters are corrected, and subsequent acquired speech is then processed using these corrected parameters. This method significantly improves the stability and clarity of audio pickup for explanations in scenarios with strong reverberation in exhibition halls and loud screen playback.
[0076] Taking a conference room large screen as an example, when deployed in a conference room, the system scans the large screen, conference table, walls, glass, and fixed speakers, and by default excludes office chairs with casters and attendees to construct a static room profile. This static room profile includes the room's geometric model, material acoustic properties, active display surface model, playback sound source and microphone array positions, as well as playback path summaries for interactive areas such as the main seating area, seat cluster area, and whiteboard presentation area.
[0077] Based on a static room profile, descriptive data for the voice front-end tuning of each interactive area is generated and mapped to corresponding preset voice front-end parameters. During conference speaking, the current user's interactive area is located based on visual human detection or the direction of sound source arrival. The corresponding preset voice front-end parameters are matched, and acoustic echo cancellation, beamforming, and dereverberation processing are performed to obtain the processed voice signal. Subsequently, the actual playback response is compared with the playback path summary, the preset voice front-end parameters are corrected, and acoustic echo cancellation, beamforming, and dereverberation processing are performed on subsequent voice acquisitions based on the corrected preset voice front-end parameters. This method effectively reduces the interference of changes in the position of the moving chair on the room profile and beam parameters, ensuring the quality of voice interaction in conference scenarios.
[0078] Taking a large home entertainment screen as an example, after the TV is installed, the system scans fixed or semi-fixed components such as the TV, TV cabinet, walls, curtains, carpets, and sofas, excluding family members, pets, portable chairs, and toy cars by default, to construct a static room profile. The static room profile includes the room's geometric model, material acoustic properties, active display surface model, playback sound source and microphone array positions, and playback path summaries for interactive areas such as the sofa viewing area and the standing motion-sensing area.
[0079] Based on a static room profile, a voice front-end tuning descriptor is generated for each interactive area and mapped to a corresponding preset voice front-end parameter. The preset voice front-end parameter is matched to the user's interactive area, and acoustic echo cancellation, beamforming, and dereverberation processing are performed to obtain the processed voice signal. Subsequently, the preset voice front-end parameter is corrected by comparing the actual playback response with the playback path summary. Based on the corrected preset voice front-end parameter, subsequent voice acquisitions are then processed with acoustic echo cancellation, beamforming, and dereverberation. This method simultaneously addresses echo suppression in movie-watching scenarios and voice clarity in haptic interaction scenarios, improving the overall experience of voice interaction on large home screens.
[0080] Based on the large-screen voice front-end partition acoustic tuning method described in the previous embodiments, this application also provides a large-screen voice front-end partition acoustic tuning device. Figure 2 This is a schematic diagram of the device. Figure 2 As shown, the large-screen voice front-end zone acoustic tuning device includes: The region-adaptive speech processing module 201 is used to match the corresponding preset speech front-end parameters based on the current user's interaction region, and perform acoustic echo cancellation, beamforming, and dereverberation processing on the currently acquired speech to obtain the processed speech signal. The preset speech front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset speech front-end parameters are obtained based on the mapping of the speech front-end tuning descriptor corresponding to the interaction region. The speech front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, the location of the playback sound source, the location of the microphone array, a set of interaction regions, and an acoustic path summary corresponding to each interaction region in the set of interaction regions.
[0081] The voice front-end parameter correction module 202 is used to compare the received actual playback response with the acoustic path summary in the static room image, and correct the preset voice front-end parameters according to the comparison result, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
[0082] Optionally, methods for obtaining static room images include: A geometric model of the room is constructed based on the scanned image data of the room; the image data includes color images and depth maps. Using the image data, semantic segmentation, object recognition, and material classification are performed on non-screen static areas. If the confidence value of the material classification result exceeds a preset confidence threshold, a set of material acoustic attribute parameters is obtained through probability-weighted material acoustic mapping. If the confidence value of the material classification result does not exceed the preset confidence threshold, a rollback process is performed based on room type, object category, default material table, or user annotation. An active display surface model is established based on the interactive area where the large screen is located, and the parameters of the active display surface model are determined. The active display surface model parameters include screen plane, screen normal, screen boundary, screen mask, screen reflection properties, playback sound source position parameters, and frequency band coupling parameters. The interactive areas are divided according to the room geometry model, the set of material acoustic properties parameters, the active display surface model parameters, the playback sound source position and the microphone array position to obtain an interactive area set, and the corresponding acoustic path summary is determined based on each interactive area in the interactive area set. Based on the room geometry model, the set of material acoustic properties parameters, the parameters of the active display surface model, the location of the playback sound source, the location of the microphone array, the set of interactive areas, and the corresponding acoustic path summary, a static room profile is determined.
[0083] Optionally, based on the scanned image data of the room, a geometric model of the room is constructed, including: Identify multiple objects in the image data to obtain a first object set; The first set of objects is filtered based on multidimensional conditions to obtain a second set of objects; the multidimensional conditions include at least two of the following: semantic recognition, occupation persistence, motion consistency, rigid anchoring relationship or historical position consistency. Remove the movable objects from the second object set to obtain the third object set; the movable objects include people, pets, chairs, wheelchairs, and strollers. Based on the aforementioned third set of objects, a room geometric model is constructed.
[0084] Optionally, based on each interactive region in the set of interactive regions, a corresponding acoustic path summary is determined, including: A dual-path acoustic estimation model is used to determine the corresponding acoustic path summary based on each interaction region in the set of interaction regions. The dual-path acoustic estimation model includes a playback path model and a user voice path model. The acoustic path summary includes a playback path summary and a user voice path summary. The playback path model is used to obtain the playback path summary based on the propagation risk of the playback signal from the built-in speaker or external soundbar of the display device entering the microphone array through the active display surface model, the main reflective surface, and free space. The playback path summary includes the frequency band playback leakage risk, maximum echo delay, echo tail length, frequency band leakage mask, and null candidate direction to configure acoustic echo cancellation parameters. The user voice path model is used to obtain the user voice path summary based on the direct path, occlusion path, and reflection path of the user voice to the microphone array within the interaction region. The user voice path summary includes the direct path visibility, direct sound gain, strong reflection risk, late reverberation risk, and microphone channel reliability to configure beamforming parameters and dereverberation parameters.
[0085] Optionally, the voice front-end parameter correction module is used for: Within a preset time period, the actual playback response is determined based on the playback reference signal and the actual playback response received by the microphone array. Calculate the response deviation between the actual playback response and the acoustic path summary; Based on the response deviation, the preset speech front-end parameters are boundedly corrected according to a preset step size to obtain the corrected preset speech front-end parameters; the correction includes correcting the echo cancellation filter length, correcting the residual echo suppression intensity, correcting the dereverberation intensity, or correcting the channel weight.
[0086] Optionally, the static room image also includes a room signature, and the device further includes an update module; The update module is used to compare the real-time room signature with the historical room signature. If the difference between the real-time room signature and the historical room signature is less than a first preset threshold, the preset voice front-end parameters of each interaction area are directly reused. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The room signature includes screen pose, main plane parameters, furniture surround box, material histogram, free space, playback sound source, and microphone layout. If the difference is between the first preset threshold and the second preset threshold, then a local incremental update is performed; If the difference exceeds the second preset threshold, a prompt will be made to rescan the room; the first preset threshold is less than the second preset threshold.
[0087] Optionally, the methods for obtaining the current user's interaction area include: Based on at least one of the following: visual human detection, sound source arrival direction, wake word localization, historical wake-up area, application mode, or user manual selection, determine one or more candidate interaction areas where the current user is located. If there is a candidate interaction region, then the candidate region is taken as the target interaction region; if there are multiple candidate interaction regions, then the confidence values of the multiple candidate interaction regions are fused by region confidence, and the target interaction region is obtained based on the fusion result. The target interaction area is compared with the historical interaction area. If the comparison result indicates that the target interaction area and the historical interaction area are consistent, the target interaction area is output. If the comparison result indicates that the target interaction area and the historical interaction area are inconsistent, a region switching anti-shake verification is triggered. The region switching anti-shake verification is as follows: if the continuous effective duration of the target interaction area exceeds a preset duration, the region switching is effective and the target interaction area is output. If the continuous effective duration of the target interaction area does not exceed the preset duration, the region switching is invalid and the historical interaction area is output.
[0088] Optionally, the preset mapping method for the voice front-end parameters includes: The voice front-end tuning descriptors include echo descriptors, channel reliability descriptors, and reverberation descriptors; the echo descriptors include playback energy, echo delay, and echo tail duration; the channel reliability descriptors include microphone channel signal-to-noise ratio and channel reliability index; the reverberation descriptors include late reverberation duration and speech signal-to-noise ratio; Acoustic echo cancellation parameters are generated using the echo class descriptor mapping; the acoustic echo cancellation parameters include acoustic echo cancellation filter length, residual echo suppression strength, and leakage risk mask; Beamforming parameters are generated by mapping the channel reliability class descriptor; the beamforming parameters include the number of effective microphone channels, beam weight, and beam operating frequency band. De-reverberation parameters are generated using the reverberation class descriptor mapping; the de-reverberation parameters include de-reverberation intensity, late reverberation time-domain parameters, and post-filter coefficients.
[0089] In addition, embodiments of this application also provide an electronic device, which includes a memory, a processor, a microphone array interface, a playback reference signal interface, and an optional image acquisition interface.
[0090] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment solution according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0091] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for zoned acoustic tuning of a large-screen voice front-end, characterized in that, include: Based on the current user's interaction area, corresponding preset voice front-end parameters are matched, and acoustic echo cancellation, beamforming, and dereverberation processing are performed on the currently acquired voice signal to obtain the processed voice signal. The preset voice front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset voice front-end parameters are obtained based on the voice front-end tuning descriptor mapping corresponding to the interaction area. The voice front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, the location of the playback sound source, the location of the microphone array, a set of interaction areas, and an acoustic path summary corresponding to each interaction area in the set of interaction areas. Based on the comparison between the received actual playback response and the acoustic path summary in the static room image, the preset voice front-end parameters are corrected according to the comparison results, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
2. The method according to claim 1, characterized in that, The methods for obtaining the static room image include: A geometric model of the room is constructed based on the scanned image data of the room; the image data includes color images and depth maps. Using the image data, semantic segmentation, object recognition, and material classification are performed on non-screen static areas. If the confidence value of the material classification result exceeds a preset confidence threshold, a set of material acoustic attribute parameters is obtained through probability-weighted material acoustic mapping. If the confidence value of the material classification result does not exceed the preset confidence threshold, rollback processing is performed based on room type, object category, default material table, or user annotation. An active display surface model is established based on the interactive area where the large screen is located, and the parameters of the active display surface model are determined. The active display surface model parameters include screen plane, screen normal, screen boundary, screen mask, screen reflection properties, playback sound source position parameters, and frequency band coupling parameters. The interactive areas are divided according to the room geometry model, the set of material acoustic properties parameters, the active display surface model parameters, the playback sound source position and the microphone array position to obtain an interactive area set, and the corresponding acoustic path summary is determined based on each interactive area in the interactive area set. Based on the room geometry model, the set of material acoustic properties parameters, the parameters of the active display surface model, the location of the playback sound source, the location of the microphone array, the set of interactive areas, and the corresponding acoustic path summary, a static room profile is determined.
3. The method according to claim 2, characterized in that, The construction of a room geometric model based on scanned room image data includes: Identify multiple objects in the image data to obtain a first object set; The first set of objects is filtered based on multidimensional conditions to obtain a second set of objects; the multidimensional conditions include at least two of the following: semantic recognition, occupation persistence, motion consistency, rigid anchoring relationship or historical position consistency. Remove the movable objects from the second object set to obtain the third object set; the movable objects include people, pets, chairs, wheelchairs, and strollers; Based on the aforementioned third set of objects, a room geometric model is constructed.
4. The method according to claim 2, characterized in that, The step of determining the corresponding acoustic path summary based on each interaction region in the set of interaction regions includes: A dual-path acoustic estimation model is used to determine the corresponding acoustic path summary based on each interaction region in the set of interaction regions. The dual-path acoustic estimation model includes a playback path model and a user voice path model. The acoustic path summary includes a playback path summary and a user voice path summary. The playback path model is used to obtain the playback path summary based on the propagation risk of the playback signal from the built-in speaker or external soundbar of the display device entering the microphone array through the active display surface model, the main reflective surface, and free space. The playback path summary includes the frequency band playback leakage risk, maximum echo delay, echo tail length, frequency band leakage mask, and null candidate direction to configure acoustic echo cancellation parameters. The user voice path model is used to obtain the user voice path summary based on the direct path, occlusion path, and reflection path of the user voice to the microphone array within the interaction region. The user voice path summary includes the direct path visibility, direct sound gain, strong reflection risk, late reverberation risk, and microphone channel reliability to configure beamforming parameters and dereverberation parameters.
5. The method according to claim 1, characterized in that, The process of comparing the received actual playback response with the acoustic path summary in the static room profile, and correcting the preset speech front-end parameters based on the comparison result, includes: Within a preset time period, the actual playback response is determined based on the playback reference signal and the actual playback response received by the microphone array. Calculate the response deviation between the actual playback response and the acoustic path summary; Based on the response deviation, the preset speech front-end parameters are boundedly corrected according to a preset step size to obtain the corrected preset speech front-end parameters; the correction includes correcting the echo cancellation filter length, correcting the residual echo suppression intensity, correcting the dereverberation intensity, or correcting the channel weight.
6. The method according to claim 1, characterized in that, The static room image also includes a room signature, and the method further includes: The real-time room signature is compared with the historical room signature. If the difference between the real-time room signature and the historical room signature is less than a first preset threshold, the preset voice front-end parameters of each interactive area are directly reused. The room signature includes screen pose, main plane parameters, furniture bounding box, material histogram, free space, playback sound source and microphone layout. If the difference is between the first preset threshold and the second preset threshold, then a local incremental update is performed; If the difference exceeds the second preset threshold, a prompt will be made to rescan the room; the first preset threshold is less than the second preset threshold.
7. The method according to claim 1, characterized in that, The methods for obtaining the current user's interaction area include: Based on at least one of the following: visual human detection, sound source arrival direction, wake word localization, historical wake-up area, application mode, or user manual selection, determine one or more candidate interaction areas where the current user is located. If there is a candidate interaction region, then the candidate region is taken as the target interaction region; if there are multiple candidate interaction regions, then the confidence values of the multiple candidate interaction regions are fused by region confidence, and the target interaction region is obtained based on the fusion result. The target interaction area is compared with the historical interaction area. If the comparison result indicates that the target interaction area and the historical interaction area are consistent, the target interaction area is output. If the comparison result indicates that the target interaction area and the historical interaction area are inconsistent, a region switching anti-shake verification is triggered. The region switching anti-shake verification is as follows: if the continuous effective duration of the target interaction area exceeds a preset duration, the region switching is effective and the target interaction area is output. If the continuous effective duration of the target interaction area does not exceed the preset duration, the region switching is invalid and the historical interaction area is output.
8. The method according to claim 1, characterized in that, The preset mapping method for the voice front-end parameters includes: The voice front-end tuning descriptors include echo descriptors, channel reliability descriptors, and reverberation descriptors; the echo descriptors include playback energy, echo delay, and echo tail duration; the channel reliability descriptors include microphone channel signal-to-noise ratio and channel reliability index; the reverberation descriptors include late reverberation duration and speech signal-to-noise ratio; Acoustic echo cancellation parameters are generated using the echo class descriptor mapping; the acoustic echo cancellation parameters include the acoustic echo cancellation filter length, residual echo suppression strength, and leakage risk mask; Beamforming parameters are generated by mapping the channel reliability class descriptor; the beamforming parameters include the number of effective microphone channels, beam weight, and beam operating frequency band. De-reverberation parameters are generated using the reverberation class descriptor mapping; the de-reverberation parameters include de-reverberation intensity, late reverberation time-domain parameters, and post-filter coefficients.
9. A large-screen voice front-end zone acoustic tuning device, characterized in that, include: The region-adaptive speech processing module is used to match corresponding preset speech front-end parameters based on the current user's interaction region, and perform acoustic echo cancellation, beamforming, and dereverberation processing on the currently acquired speech to obtain the processed speech signal. The preset speech front-end parameters include acoustic echo cancellation parameters, beamforming parameters, and dereverberation parameters. The preset speech front-end parameters are obtained based on the mapping of the speech front-end tuning descriptor corresponding to the interaction region. The speech front-end tuning descriptor is determined based on a static room profile. The static room profile includes a room geometric model, a set of material acoustic property parameters, an active display surface model, the location of the playback sound source, the location of the microphone array, a set of interaction regions, and an acoustic path summary corresponding to each interaction region in the set of interaction regions. The voice front-end parameter correction module is used to compare the received actual playback response with the acoustic path summary in the static room image, and correct the preset voice front-end parameters according to the comparison result, so as to perform acoustic echo cancellation, beamforming and dereverberation processing on the subsequently acquired voice based on the corrected preset voice front-end parameters.
10. An electronic device, characterized in that, The electronic device includes a memory, a processor, a microphone array interface, a playback reference signal interface, and an optional image acquisition interface.