A video and audio linkage control system and method based on multi-modal scene recognition

CN122824936APending Publication Date: 2026-09-25HEBEI JIAQIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610956452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,现有技术多停留在单一音频或单一字幕维度的独立优化,缺乏对人声活动与字幕显示之间关联场景的综合识别能力,不易于准确刻画用户实际收听环境中的人声清晰度变化,同时,现有方案在面对人声不足与低频遮蔽等多种影响因素时,通常采用统一增益调节方式,缺乏对不同干扰成因的区分与验证机制,导致调节策略针对性不足,容易引入过度补偿或失真问题,不易于实现稳定、精细化的影音联动控制效果

Benefits of technology

[0036]1、通过获取中心声道音频信号与字幕控制信号,并基于语音活动区间与字幕显示段的重叠关系确定人声字幕关联场景,使得系统能够在真实播放内容中准确锁定人声与字幕同步出现的有效片段,从而解决了现有技术中无法结合音频与字幕信息对有效语义场景进行准确识别,导致后续调节缺乏场景基础的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824936A_ABST
    Figure CN122824936A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal scene identification's audio-video linkage control system and method, it is related to audio-video linkage technical field, obtains center sound channel audio signal, subtitle control signal and microphone acquisition signal, after audio-video playing equipment enters human voice enhancement adjustment mode, by center sound channel audio signal determine speech activity interval, by subtitle control signal determine subtitle display section, and when the overlap proportion of speech activity interval and subtitle display section reaches preset proportion, determine human voice subtitle correlation scene;Delay compensation is carried out under the correlation scene and extracts human voice frequency band equivalent sound pressure level and low-frequency band equivalent sound pressure level, respectively judges human voice intelligibility insufficient state identifier and low-frequency interference enhancement state identifier;It can realize the differential identification and verification of human voice insufficient and low-frequency shielding, improve the pertinence and accuracy of sound adjustment, avoid the distortion problem caused by uniform gain adjustment, so as to improve user listening intelligibility and overall audio-visual experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio-visual linkage technology, specifically to an audio-visual linkage control system and method based on multimodal scene recognition. Background Technology

[0002] With the continuous development of intelligent audio-visual terminals and multimedia playback technologies, audio decoding, subtitle rendering, and multimodal signal fusion have gradually become important directions for improving users' audio-visual experience. Existing audio-visual playback devices can usually achieve basic volume adjustment or voice enhancement based on audio signals, while providing auxiliary understanding functions by combining subtitle information.

[0003] However, existing technologies mostly focus on independent optimization of a single audio or subtitle dimension, lacking the ability to comprehensively identify the correlation between human voice activity and subtitle display. They are not easy to accurately depict changes in the clarity of human voice in the user's actual listening environment. At the same time, when faced with multiple influencing factors such as insufficient human voice and low-frequency masking, existing solutions usually adopt a uniform gain adjustment method, lacking a mechanism to distinguish and verify different interference causes. This results in insufficient targeting of adjustment strategies, which can easily introduce overcompensation or distortion problems, making it difficult to achieve stable and refined audio-visual linkage control effects. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the technical solution of this invention is as follows:

[0005] An audio-visual linkage control method based on multimodal scene recognition includes the following steps:

[0006] S1. Acquire the center channel audio signal, subtitle control signal and microphone acquisition signal. After the audio-visual playback device enters the human voice enhancement adjustment mode, determine the voice activity range and subtitle display segment based on the center channel audio signal and subtitle control signal, and determine the human voice subtitle association scene based on the voice activity range and the subtitle display segment.

[0007] S2. In the human voice subtitle association scenario, delay compensation is performed on the microphone acquisition signal to obtain the compensated microphone acquisition signal, and the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low frequency band are extracted. Based on the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low frequency band, the status indicators of insufficient human voice clarity and enhanced low frequency interference are determined respectively.

[0008] S3. Within the human voice subtitle association scenario, determine a first verification period and a second verification period that do not overlap. Based on the first verification period, perform a temporary increase verification of the center channel gain for the determined state indicator of insufficient human voice clarity, and obtain the first verification human voice sound pressure level. Based on the second verification period, perform a temporary decrease verification of the low frequency gain for the determined state indicator of enhanced low frequency interference, and obtain the second verification low frequency sound pressure level.

[0009] S4. Determine the weight value of the voice intelligibility insufficiency status indicator based on the difference between the first verified human voice sound pressure level and the equivalent sound pressure level of the human voice frequency band; determine the weight value of the low-frequency interference enhancement status indicator based on the difference between the equivalent sound pressure level of the low-frequency frequency band and the second verified low-frequency sound pressure level; and generate an audio-visual linkage control command based on the weight values ​​of the determined candidate causes.

[0010] Furthermore, in S1, the overlap duration is determined based on the speech activity interval and the subtitle display segment, and the ratio of the overlap duration to the speech activity interval duration is the overlap ratio.

[0011] The subtitle display segment appears synchronously with the voice activity segment, and the time difference between the start time of the subtitle display segment and the start time of the voice activity segment is less than the preset synchronization deviation threshold.

[0012] Furthermore, in S2, when the equivalent sound pressure level of the human voice frequency band is lower than the preset human voice sound pressure threshold, the status indicator of insufficient human voice clarity is determined to be established.

[0013] When the equivalent sound pressure level in the low-frequency band is higher than the preset low-frequency sound pressure threshold, the low-frequency interference enhancement status is determined to be established.

[0014] Furthermore, in S3, when both the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator are valid, only the center channel gain is changed during the first verification period, and only the low-frequency gain is changed during the second verification period.

[0015] The first verification period and the second verification period do not overlap.

[0016] Furthermore, in S3, after the first verification period ends, the audio output state is restored to the audio output state before the start of the first verification period, and then the second verification period begins.

[0017] After the second verification period ends, the audio output status will be restored to the audio output status before the start of the second verification period.

[0018] Furthermore, in S4, the sound pressure level of the first verified human voice is compared with the corresponding reference sound pressure level of the human voice to obtain the human voice verification quantity;

[0019] The second verification low-frequency sound pressure level is compared with the corresponding low-frequency reference sound pressure level to obtain the low-frequency verification quantity.

[0020] A reference time period is set within the human voice subtitle association scene. The reference time period is located outside the first verification time period and the second verification time period. The reference time period includes the corresponding human voice reference sound pressure level and low frequency reference sound pressure level.

[0021] Furthermore, in S4, the voice verification quantity is compared with the preset voice enhancement threshold. When the voice verification quantity reaches the preset voice enhancement threshold, the weight value of the voice clarity insufficiency status indicator is determined based on the amount by which the voice verification quantity exceeds the preset voice enhancement threshold.

[0022] The low-frequency verification quantity is compared with the preset low-frequency reduction threshold. When the low-frequency verification quantity reaches the preset low-frequency reduction threshold, the weight value of the low-frequency interference enhancement status indicator is determined based on the amount by which the low-frequency verification quantity exceeds the preset low-frequency reduction threshold.

[0023] Furthermore, in S4, when the number of human voice verifications does not reach the preset human voice enhancement threshold, the weight value of the insufficient human voice clarity status indicator is determined as an invalid weight.

[0024] When the low-frequency verification quantity does not reach the preset low-frequency reduction threshold, the weight value of the low-frequency interference enhancement status identifier is determined as an invalid weight.

[0025] Audio-visual linkage control commands do not include audio adjustment parameters corresponding to invalid weights.

[0026] Furthermore, in S4, when both the weight values ​​of the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator are valid weights, the audio-visual linkage control commands include center channel gain adjustment parameters, low-frequency gain adjustment parameters, and subtitle enhancement parameters.

[0027] The center channel gain adjustment parameter is determined by the weight value of the insufficient vocal clarity status indicator;

[0028] The low-frequency gain adjustment parameter is determined by the weight value of the low-frequency interference enhancement status indicator;

[0029] The subtitle enhancement parameters are jointly determined by the weight values ​​of the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator.

[0030] An audio-visual linkage control system based on multimodal scene recognition includes a scene recognition module, a candidate cause determination module, a time-division verification module, and a weight control module.

[0031] The scene recognition module acquires the center channel audio signal, subtitle control signal and microphone acquisition signal. After the audio-visual playback device enters the human voice enhancement adjustment mode, the center channel audio signal determines the voice activity range, the subtitle control signal determines the subtitle display segment, and when the overlap ratio between the voice activity range and the subtitle display segment reaches a preset ratio, the human voice subtitle associated scene is determined.

[0032] The candidate cause determination module performs delay compensation on the microphone acquisition signal in the human voice subtitle association scenario to obtain the compensated microphone acquisition signal, and extracts the human voice frequency band equivalent sound pressure level and the low frequency band equivalent sound pressure level from the compensated microphone acquisition signal. Based on the human voice frequency band equivalent sound pressure level and the low frequency band equivalent sound pressure level, it determines the status flag of insufficient human voice clarity and the status flag of enhanced low frequency interference, respectively.

[0033] The time-sharing verification module determines a first verification period and a second verification period that do not overlap within the human voice and subtitle associated scene. Based on the first verification period, it verifies the temporary increase of the center channel gain for the determined state of insufficient human voice clarity and obtains the first verification human voice sound pressure level. Based on the second verification period, it verifies the temporary decrease of the low frequency gain for the determined state of enhanced low frequency interference and obtains the second verification low frequency sound pressure level.

[0034] The weight control module determines the weight value of the voice intelligibility deficiency status indicator based on the difference between the first verified voice sound pressure level and the equivalent sound pressure level of the voice frequency band, determines the weight value of the low-frequency interference enhancement status indicator based on the difference between the equivalent sound pressure level of the low-frequency frequency band and the second verified low-frequency sound pressure level, and generates audio-visual linkage control instructions including audio adjustment parameters and subtitle enhancement parameters based on the weight values ​​of the determined candidate causes.

[0035] The beneficial effects of this invention are as follows:

[0036] 1. By acquiring the center channel audio signal and subtitle control signal, and determining the human voice and subtitle association scene based on the overlap between the voice activity interval and the subtitle display segment, the system can accurately locate the effective segment in which human voice and subtitle appear synchronously in the actual playback content. This solves the problem in the existing technology that it is impossible to accurately identify effective semantic scenes by combining audio and subtitle information, resulting in a lack of scene basis for subsequent adjustment.

[0037] 2. By performing delay compensation on the microphone acquisition signal and extracting the equivalent sound pressure level of the human voice frequency band and the low frequency band, the status indicators of insufficient human voice clarity and enhanced low frequency interference are determined respectively. Combined with a non-overlapping time-division verification mechanism, single-factor verification is carried out, so that different interference factors can be distinguished, verified and evaluated independently. This solves the problem in the existing technology of single judgment of the cause of human voice clarity problem and inability to distinguish between insufficient human voice and low frequency masking leading to misadjustment. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the method steps of the present invention;

[0039] Figure 2 This is a schematic diagram illustrating the relationship between the voice activity interval and the subtitle display time segment of the present invention;

[0040] Figure 3 This invention relates to the timeline of scenes associated with human voice subtitles.

[0041] Figure 4 This is a schematic diagram illustrating the delay compensation principle of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] Example 1

[0044] Please see Figures 1-4 The present invention provides an audio-visual linkage control method based on multimodal scene recognition, which includes the following steps.

[0045] In S1, after the audio-visual playback device enters the voice enhancement adjustment mode, it acquires the center channel audio signal, subtitle control signal, and microphone acquisition signal. The voice enhancement adjustment mode can be triggered by the user through the sound settings menu, remote control buttons, voice commands, or mobile terminal control commands. Upon receiving the entry command, the audio-visual playback device writes a voice enhancement adjustment mode flag into its local control program. When the voice enhancement adjustment mode flag is active, the audio-visual playback device begins reading the center channel audio signal, subtitle control signal, and microphone acquisition signal corresponding to the current playback time.

[0046] The center channel audio signal is obtained from the audio decoding link. The audio-visual playback device demultiplexes and decodes the current media data to obtain multi-channel audio data. When the current media data includes a center channel, the sampled data corresponding to the center channel is read as the center channel audio signal. When the current media data is stereo content, the in-phase components of the left and right channel sampled data are extracted to obtain a virtual center channel audio signal, which is then used as the center channel audio signal. The subtitle control signal is obtained from the subtitle parsing link or the subtitle rendering link. The subtitle control signal includes the subtitle start time, subtitle end time, and subtitle display status. The microphone acquisition signal is obtained from the microphone installed in the television, soundbar, speakers, remote control, or mobile terminal. The microphone acquisition signal is used to reflect the actual sound field state near the user's listening position.

[0047] When determining the speech activity range based on the center channel audio signal, the audio-visual playback device divides the center channel audio signal into multiple audio frames according to a 40ms time window, with a 20ms overlap between adjacent audio frames. For each audio frame, a frequency domain transformation is performed to obtain the spectral amplitude of that audio frame. The squared spectral amplitudes within the 300Hz to 3400Hz frequency band are summed to obtain the human voice frequency band energy. The squared spectral amplitudes within the 80Hz to 8000Hz frequency band are summed to obtain the total effective energy of the center channel. The human voice frequency band energy is divided by the total effective energy of the center channel to obtain the proportion of human voice frequency band energy.

[0048] Audio frames with a center channel effective total energy below a preset silence threshold are not included in the speech activity interval determination. The preset silence threshold is determined by the noise floor of the center channel in a dialogue-free segment, or it can be written into the parameter table by the device during factory testing. In this embodiment, the preset silence threshold is -48dBFS. Audio frames with a center channel effective total energy reaching the preset silence threshold and a human voice frequency band energy proportion greater than a preset proportion threshold are marked as human voice frames. The preset proportion threshold is determined by the distribution of human voice frequency band energy proportions in dialogue samples, music samples, and ambient sound samples; in this embodiment, it is 45%. When the duration covered by consecutive human voice frames reaches a preset duration, the time period covered by consecutive human voice frames is determined as the speech activity interval. The preset duration is determined based on the sustained difference between short noises and normal dialogue; in this embodiment, it is 500ms. When the interval between two adjacent groups of human voice frames is less than 200ms, the two adjacent groups of human voice frames are merged into the same speech activity interval.

[0049] When the subtitle control signal determines the subtitle display segment, the audio-visual playback device reads the subtitle start time, subtitle end time, and subtitle display status from the subtitle entry. When the subtitle display status is "display," the time period between the subtitle start time and the subtitle end time is defined as the subtitle display segment. For external subtitles, the subtitle start time and subtitle end time are obtained from the timeline field of the subtitle file; for embedded subtitles, the subtitle start time and subtitle end time are obtained from the subtitle track timestamp in the media container; for streaming subtitles, the subtitle start time and subtitle end time are obtained from the subtitle timestamp received by the playback device.

[0050] After the voice activity interval and subtitle display segment are determined, the audio-visual playback device calculates the overlap duration between the voice activity interval and the subtitle display segment. The start time of the voice activity interval is denoted as Tv1, and the end time as Tv2. The start time of the subtitle display segment is denoted as Ts1, and the end time as Ts2. When min(Tv2, Ts2) is greater than max(Tv1, Ts1), the overlap duration is the difference between min(Tv2, Ts2) and max(Tv1, Ts1); when min(Tv2, Ts2) is not greater than max(Tv1, Ts1), the overlap duration is 0. The overlap duration is divided by the voice activity interval duration to obtain the overlap ratio. The voice activity interval duration is the difference between Tv2 and Tv1, where min is the minimum value and max is the maximum value.

[0051] The preset ratio is determined by sample statistics and the device parameter table. Specifically, during the device's factory testing phase, film clips with dialogue subtitles are selected as positive samples, and lyric subtitles, decorative subtitles, on-screen description subtitles, and voice clips without dialogue subtitles are selected as negative samples. The overlap ratio in the positive and negative samples is calculated separately. Within the candidate range of 60% to 85%, the ratio that satisfies the positive sample recognition requirements and ensures that the false triggering of negative samples meets the control requirements is selected and written into the device parameter table as the preset ratio. In this embodiment, 70% is selected as the preset ratio after sample statistics. The device can also update the preset ratio in the parameter table during the calibration phase based on the subtitle rendering delay and the source subtitle synchronization error.

[0052] A preset synchronization deviation threshold is used to determine whether the subtitle display segment and the audio activity interval appear synchronously. The preset synchronization deviation threshold is determined based on audio output delay, subtitle rendering delay, and source subtitle production errors. The device can measure the audio output delay and subtitle rendering delay during factory testing and combine them with common subtitle production errors to form a synchronization deviation parameter. In this embodiment, the preset synchronization deviation threshold is 0.8s. When the time difference between the start time of the subtitle display segment and the start time of the audio activity interval is less than the preset synchronization deviation threshold, it is determined that the subtitle display segment and the audio activity interval appear synchronously.

[0053] For example, the voice activity interval is from 120.00s to 124.00s, and the subtitle display interval is from 120.20s to 124.10s. At this time, Tv1 is 120.00s, Tv2 is 124.00s, Ts1 is 120.20s, and Ts2 is 124.10s.

[0054] min(Tv2, Ts2) is 124.00s, max(Tv1, Ts1) is 120.20s, and the overlap duration is 3.80s.

[0055] The duration of the voice activity interval is 4.00 seconds, with an overlap rate of 95.0%, reaching the preset ratio of 70%. The time difference between the start time of the subtitle display segment and the start time of the voice activity interval is 0.20 seconds, which is less than the preset synchronization deviation threshold of 0.8 seconds. Therefore, this segment is identified as a voice-to-subtitle association scene.

[0056] In S2, in the scenario of human voice and subtitle association, the audio-visual playback device performs delay compensation on the microphone acquisition signal so that the microphone acquisition signal corresponds to the center channel audio signal on the same playback time axis.

[0057] Delay compensation can be achieved through a pre-calibrated playback acquisition delay. Specifically, during factory testing or initial calibration, the audio-visual playback device plays a test tone, records the playback time of the test tone at the audio output, and records the acquisition time of the test tone by the microphone. The difference between the acquisition time and the playback time is used as the playback acquisition delay and written into the device parameter table. During operation, the audio-visual playback device performs time shifting on the microphone acquisition signal according to the playback acquisition delay to obtain the compensated microphone acquisition signal.

[0058] In another implementation, within each time window, the center channel audio signal and the microphone acquisition signal are discretely sampled to obtain the corresponding discrete signal sequence.

[0059] The discrete signal sequence is bandpass filtered according to a preset frequency band range, which is used to extract the frequency components corresponding to the human voice frequency band or low-frequency band. After bandpass filtering, the filtered discrete signal sequence is squared point by point to obtain the instantaneous energy sequence. The instantaneous energy sequence is accumulated within the current time window and divided by the number of sampling points in the current time window to obtain the average energy value within that time window. The average energy value is used as the band energy value of the corresponding frequency band to characterize the sound energy intensity of the corresponding frequency band within the current time window.

[0060] The human voice frequency band corresponds to a frequency range of 300Hz to 3400Hz, and the low frequency band corresponds to a frequency range of 20Hz to 250Hz.

[0061] After obtaining the frequency band energy value, the frequency band energy value is processed by sound pressure mapping in combination with the microphone sensitivity calibration parameters, so as to convert it into an equivalent sound pressure level corresponding to the user's listening position.

[0062] For example, if the peak energy of the human voice in the center channel audio signal occurs at 122.00s, and the corresponding peak value in the microphone acquisition signal occurs at 122.11s, then the delay compensation is 110ms. The audio-visual playback device advances the microphone acquisition signal by 110ms to obtain the compensated microphone acquisition signal.

[0063] During the equipment factory calibration phase, weighted mapping relationships were constructed for the status indicators of insufficient human voice clarity and enhanced low-frequency interference, respectively.

[0064] First, test audio samples containing speech and low-frequency background components are played in a standard acoustic testing environment, and a speech-subtitle synchronization scenario is constructed by the synchronization relationship between the human voice activity range and the subtitle presentation range.

[0065] In this scenario, the center channel gain boost and low-frequency band attenuation operations are executed independently, and the corresponding verification output results are obtained in accordance with the time-sharing verification method consistent with the operation phase.

[0066] The calibration process for the status indicator of insufficient human voice clarity is as follows: The difference before and after the center channel gain is boosted is used as the basic verification quantity. The difference is calculated from the change in sound pressure level within the human voice verification interval, and the sound pressure reference value of the corresponding reference time period is subtracted to form the human voice verification difference ΔVvoice.

[0067] Using ΔVvoice as the horizontal axis variable and the candidate cause weights as the vertical axis variable, repeated sampling was performed under various audio samples, playback intensities, and acoustic environments to construct a statistical mapping relationship between ΔVvoice and the weight response. A human voice weight calibration curve was then generated through regression fitting. For the low-frequency interference enhancement status indicator, the difference in sound pressure level before and after the low-frequency attenuation operation was used as the low-frequency verification difference ΔVlow. Similarly, the corresponding low-frequency reference value for the baseline period was subtracted to form a low-frequency verification difference sequence. A low-frequency weight calibration curve was constructed based on the mapping relationship between ΔVlow and its corresponding weights, and then discretized to form a parameter table. During the runtime phase, based on the real-time calculated ΔVvoice and ΔVlow, the weight values ​​of the corresponding candidate causes were obtained by looking up values ​​in the parameter table or by interpolation.

[0068] After the compensated microphone acquires the signal, the audio-visual playback device extracts the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low-frequency frequency band within the sampling window corresponding to the speech activity interval. The sampling window can be set according to the length of the speech activity interval; in this embodiment, the sampling window length is 1 second.

[0069] The vocal and low-frequency bands can be preset according to the device parameter table. The vocal band is used to cover the frequency range in white voices that contributes significantly to clarity; in this embodiment, it is set to 300Hz to 3400Hz. The low-frequency band is used to cover low-frequency components that can easily create a booming effect and affect the perception of dialogue; in this embodiment, it is set to 20Hz to 250Hz. The above frequency band ranges can be written into the parameter table during the device's factory testing phase based on dialogue samples, low-frequency sound effect samples, and speaker frequency response results from different sources, or they can be adjusted during the device calibration phase based on the low-frequency response of the user's room.

[0070] The audio-visual playback device performs vocal band filtering and low-frequency band filtering on the compensated microphone signals to obtain vocal band sound pressure level (SPL) signals and low-frequency band sound pressure level (SPL) signals. For any sampling window, the audio-visual playback device first calculates the average square of the corresponding frequency band SPL signal within that sampling window, and then converts it to decibels based on the ratio between the average square and the square of the reference SPL value to obtain the equivalent continuous SPL level for the corresponding frequency band. The reference SPL can be 20 μPa, commonly used in acoustic measurements. Before the SPL conversion, the audio-visual playback device converts the digital sampled values ​​into SPL values ​​based on the microphone sensitivity calibration value, so that the calculated result corresponds to the actual SPL level near the user's listening position.

[0071] The equivalent sound pressure level (SPL) of the human voice band is used to characterize the actual arrival sound pressure of the white voice at the user's location. The equivalent sound pressure level of the low-frequency band is used to characterize the actual arrival sound pressure of low-frequency sounds at the user's location. To reduce the impact of transient noise, this embodiment can take the median value of the equivalent sound pressure level of the human voice band and the equivalent sound pressure level of the low-frequency band from three consecutive sampling windows as the sound pressure level judgment value for the current human voice caption association scene.

[0072] To prevent the microphone signal from being affected by user speech, external noise, or changes in microphone position, the audio-visual playback device performs validity screening on the compensated microphone signal before extracting the equivalent sound pressure level. Specifically, the audio-visual playback device can compare the correlation between the current playback reference signal and the compensated microphone signal. When the correlation value is lower than a preset correlation threshold, the microphone input level suddenly increases and is not related to the playback reference signal, or the microphone signal is clipped, the corresponding sampling window is marked as an invalid sampling window. Invalid sampling windows are not involved in the determination of the equivalent sound pressure level in the human voice band, the equivalent sound pressure level in the low frequency band, and subsequent weight values.

[0073] A preset human voice sound pressure level threshold is used to determine whether the current dialogue voice may be too low. This threshold can be determined based on the ambient background sound pressure level, the user's viewing distance, and the device's volume level. Specifically, the audio-visual playback device measures the ambient background sound pressure level during a brief non-dialogue segment before the speech activity interval, and adds a preset human voice margin to the ambient background sound pressure level as the preset human voice sound pressure level threshold. The preset human voice margin can be set according to the dialogue intelligibility requirements in a typical indoor viewing scenario; in this embodiment, it is set to 18dB. If the ambient background sound pressure level is 41dB, then the preset human voice sound pressure level threshold is 59dB. The device can also use 58dB to 65dB as an optional threshold range and select a specific threshold within this range based on the ambient background sound pressure level and the user's volume level.

[0074] A preset low-frequency sound pressure threshold is used to determine whether the current low-frequency sound pressure may cause blockage. This threshold can be determined based on the ambient background sound pressure, the device's low-frequency output capability, and the room's low-frequency response. Specifically, the audio-visual playback device can play a low-frequency test signal during the device calibration phase to obtain the low-frequency response at the user's listening position; then, it can combine this with the low-frequency comfort range at normal viewing volume to determine the preset low-frequency sound pressure threshold. In this embodiment, the selectable range of the preset low-frequency sound pressure threshold is 68dB to 75dB. If the current device has a strong low-frequency response in the living room environment, the preset low-frequency sound pressure threshold is set to 70dB.

[0075] In a preferred implementation, to improve the reliability of the low-frequency interference enhancement status indicator, the audio-visual playback device, based on the fact that the equivalent sound pressure level in the low-frequency band is higher than a preset low-frequency sound pressure threshold, further calculates the sound pressure difference between the equivalent sound pressure level in the low-frequency band and the equivalent sound pressure level in the human voice band. When the sound pressure difference reaches a preset masking difference threshold, the low-frequency interference enhancement status indicator is used as a candidate reason for subsequent verification. The preset masking difference threshold can be 8dB to 15dB and is written into the parameter table according to the device calibration data.

[0076] In a specific example, the audio / video playback device obtains the compensated microphone signal within a scene where human voice and subtitles are associated. After band filtering and sound pressure level calculation, the equivalent sound pressure levels for the human voice band in three consecutive sampling windows are 56.1dB, 55.4dB, and 55.8dB, respectively, with a median of 55.8dB. The equivalent sound pressure levels for the low-frequency band are 72.6dB, 71.9dB, and 73.2dB, respectively, with a median of 72.6dB. The current ambient background sound pressure level is 41dB, and the preset human voice margin is 18dB; therefore, the preset human voice sound pressure threshold is 59dB. The preset low-frequency sound pressure threshold in the current device parameter table is 70dB.

[0077] Because the equivalent sound pressure level (SPL) in the human voice band (55.8 dB) is lower than the preset human voice SPL threshold of 59 dB, the audio-visual playback device determines that the human voice clarity is insufficient. Because the equivalent SPL in the low-frequency band (72.6 dB) is higher than the preset low-frequency SPL threshold of 70 dB, the audio-visual playback device determines that the low-frequency interference is enhanced. After both candidate reasons are established, the audio-visual playback device proceeds to the subsequent time-sharing verification process.

[0078] In S3, S1 has already disclosed the methods for determining the voice activity range, subtitle display segment, and human voice-subtitle association scenarios. S2 has already disclosed the method for performing delay compensation on the microphone acquisition signal in the human voice-subtitle association scenario, and extracting the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low-frequency frequency band based on the compensated microphone acquisition signal. S2 also disclosed that: when the equivalent sound pressure level of the human voice frequency band is lower than a preset human voice sound pressure threshold, a status indicator indicating insufficient human voice clarity is established; when the equivalent sound pressure level of the low-frequency frequency band is higher than a preset low-frequency sound pressure threshold, a status indicator indicating enhanced low-frequency interference is established.

[0079] In this embodiment, the insufficient voice clarity status flag indicates that in the current voice-to-subtitle association scenario, the equivalent sound pressure level of the voice frequency band at the user's listening location is lower than the preset voice sound pressure threshold corresponding to the current scenario, and the current dialogue voice may have an excessively low sound pressure. The enhanced low-frequency interference status flag indicates that in the current voice-to-subtitle association scenario, the equivalent sound pressure level of the low-frequency frequency band at the user's listening location is higher than the preset low-frequency sound pressure threshold corresponding to the current scenario, and the current low-frequency sound field may affect the perception of dialogue. The above candidate reasons serve as status markers for entering the verification process, and are subsequently verified through the first verification period and the second verification period.

[0080] It should be noted that the status indicators for insufficient voice clarity and enhanced low-frequency interference are initial status markers before entering the verification process, and are not equivalent to the final control cause. Only after the candidate cause has obtained effective verification quantity and formed effective weight after passing the corresponding verification period will it participate in the generation of audio-visual linkage control instructions.

[0081] When both the insufficient voice clarity status indicator and the enhanced low-frequency interference status indicator are active, the audio-visual playback device determines a first verification period and a second verification period within the voice-subtitle association scenario. The first verification period is used to verify a temporary increase in center channel gain in response to the insufficient voice clarity status indicator, and the second verification period is used to verify a temporary decrease in low-frequency gain in response to the enhanced low-frequency interference status indicator.

[0082] Both the first and second verification periods are determined within the context of the voice-to-subtitle association scenario. The audio-visual playback device first determines the overlapping region between the voice activity region and the subtitle display segment based on the already determined voice activity region and subtitle display segment. Specifically, the later of the start time of the voice activity region and the start time of the subtitle display segment is taken as the start time of the overlapping region; the earlier of the end time of the voice activity region and the end time of the subtitle display segment is taken as the end time of the overlapping region.

[0083] After obtaining the overlapping interval, the audio / video playback device subtracts the boundary transition time from the beginning and end of the overlapping interval to obtain the stable verification interval. The boundary transition time is determined based on the audio gain gradation time, subtitle rendering refresh time, and the pre-stabilization time required for microphone sound pressure level analysis. It can be the maximum value of the above times or a time slightly greater than the maximum value. In this embodiment, the gradation time of the center channel gain and low-frequency gain is 100ms, the subtitle rendering refresh time is approximately 40ms, and the pre-stabilization time for microphone sound pressure level analysis is approximately 160ms. Therefore, the boundary transition time is taken as 200ms.

[0084] The audio-visual playback device determines the effective statistical length of a single verification period based on the sampling time required for sound pressure level extraction. In S2, sound pressure level extraction needs to be completed within a continuous sampling window. In this embodiment, the effective statistical length is set to 1.00s. Since temporary gain changes require a gradual ingress and recovery, a single verification period includes the initial gain gradual change time, the middle effective statistical time, and the subsequent gain recovery time. In this embodiment, the initial gain gradual change time is 100ms, the middle effective statistical time is 1.00s, and the subsequent gain recovery time is 100ms; therefore, the length of a single verification period is 1.20s.

[0085] A recovery interval is set between the first and second verification periods. The recovery interval is determined based on the audio output state recovery time and the microphone sound pressure level smooth update time. In this embodiment, the audio output state recovery time is 100ms, and the microphone sound pressure level smooth update time is approximately 200ms; therefore, the recovery interval is set to 300ms. The recovery interval is used to ensure that after the temporary increase control of the center channel gain during the first verification period ends, the sound field state returns to its state before the start of the first verification period, before entering the second verification period.

[0086] To reduce the impact of temporary verification on the user's viewing experience, the audio-visual playback device prioritizes setting the first and second verification periods within segments in the stable verification range where energy changes are gradual, subtitles are not switched, and audio transient amplitude is below the preset transient threshold. The temporary increase in center channel gain and the temporary decrease in low-frequency gain are limited to the preset verification amplitude range, so that the temporary control is only used to obtain measurable sound pressure changes and does not directly form formal listening adjustment.

[0087] When the duration of the stable verification interval is greater than or equal to the sum of the lengths of two individual verification periods and the length of a recovery interval, the audio-visual playback device determines the first verification period, the recovery interval, and the second verification period in chronological order. Specifically, the start time of the stable verification interval is taken as the start time of the first verification period; 1.20 seconds after the start time of the first verification period is taken as the end time of the first verification period; 300ms after the end time of the first verification period is taken as the start time of the second verification period; and 1.20 seconds after the start time of the second verification period is taken as the end time of the second verification period. The first and second verification periods determined in this way do not overlap on the time axis.

[0088] When the duration of the stable verification interval is less than the sum of the lengths of two individual verification periods and a recovery interval, the audio-visual playback device will not perform both verifications simultaneously within the current voice-to-subtitle association scenario, but will postpone the composite verification to the next voice-to-subtitle association scenario that meets the duration condition. This process ensures that both the first and second verification periods have sufficient sound pressure level statistics time.

[0089] In a specific segment, the voice activity period is from 120.00s to 124.00s, and the subtitle display period is from 120.20s to 124.10s. The overlap between the voice activity period and the subtitle display period is from 120.20s to 124.00s. After deducting the first 200ms and the last 200ms of the overlap period, the stable verification period is from 120.40s to 123.80s, with a duration of 3.40s. The total length of the two individual verification periods is 2.40s, the recovery interval is 0.30s, and the required total duration is 2.70s. The stable verification period meets the duration requirement. Therefore, the audio-visual playback device determines 120.40s to 121.60s as the first verification period, 121.60s to 121.90s as the recovery interval, and 121.90s to 123.10s as the second verification period.

[0090] Before the start of the first verification period, the audio / video playback device records the current audio output status. The current audio output status includes center channel gain, low-frequency gain, overall volume level, left and right channel gains, and the basic gains of other channels. During the first verification period, the audio / video playback device only changes the center channel gain, temporarily increasing it. The amount of this temporary increase can be determined by the device parameter table; in this embodiment, it is set to 3dB. During the first verification period, the low-frequency gain, overall volume level, left and right channel gains, and the basic gains of other channels remain at their pre-verification period state.

[0091] To reduce the impact of sudden gain changes on sound pressure level statistics, the temporary increase in center channel gain during the first verification period can be achieved using a gradual increase and recovery method. Specifically, in the first 100ms of the first verification period, the center channel gain is gradually increased from its original value to the target temporary gain; the target temporary gain is maintained for the middle 1.00s; and in the last 100ms, the center channel gain is gradually restored to the gain value before the start of the first verification period. The audio-visual playback device acquires the microphone signal during the effective statistical time of the middle 1.00s and obtains the first verification human voice sound pressure level according to the delay compensation and human voice frequency band equivalent sound pressure level extraction method in S2.

[0092] After the first verification period ends, the audio-visual playback device restores the center channel gain to the gain value recorded before the start of the first verification period, based on the audio output status recorded before the start of the first verification period. The low-frequency gain, overall volume level, left and right channel gains, and other channel base gains remain unchanged. After the restoration interval ends, the audio-visual playback device enters the second verification period.

[0093] Before the start of the second verification period, the audio-visual playback device records the current audio output status. During the second verification period, the audio-visual playback device only changes the low-frequency gain, temporarily reducing it. The amount of temporary reduction in low-frequency gain can be determined by the device parameter table; in this embodiment, it is set to 4dB. During the second verification period, the center channel gain, overall volume level, left and right channel gains, and other channel base gains remain at their states before the start of the second verification period.

[0094] The temporary reduction of low-frequency gain during the second verification period can also be achieved using a gradual reduction and recovery method. Specifically, in the first 100ms of the second verification period, the low-frequency gain is gradually reduced from its original value to the target temporary attenuation; the target temporary attenuation is maintained for the middle 1.00s; and in the last 100ms, the low-frequency gain is gradually restored to the gain value before the start of the second verification period. The audio-visual playback device acquires the microphone signal during the effective statistical time of the middle 1.00s and obtains the second verification low-frequency sound pressure level according to the delay compensation and low-frequency band equivalent sound pressure level extraction method in S2.

[0095] During the second verification period, the audio-visual playback device can also simultaneously extract the second verification human voice auxiliary sound pressure level. The second verification human voice auxiliary sound pressure level is used to determine whether the sound pressure difference between the low frequency and the human voice has narrowed after the low frequency gain is temporarily reduced, thereby avoiding the determination of the low frequency interference enhancement status as an effective cause simply because the low frequency sound pressure drops.

[0096] After the second verification period ends, the audio-visual playback device restores the low-frequency gain to the gain value before the start of the second verification period based on the audio output status recorded before the start of the second verification period. The center channel gain, total volume level, left and right channel gains and other channel basic gains remain unchanged.

[0097] In the data example of this embodiment, the equivalent sound pressure level (SPL) of the human voice band before the start of the first verification period is 55.8 dB. During the first verification period, the audio-visual playback device only temporarily increases the center channel gain by 3 dB, resulting in a first verification human voice SPL of 59.2 dB within the effective statistical time. The equivalent SPL of the low-frequency band before the start of the second verification period is 72.6 dB. During the second verification period, the audio-visual playback device only temporarily decreases the low-frequency gain by 4 dB, resulting in a second verification low-frequency SPL of 68.9 dB within the effective statistical time. A recovery interval is set between the first and second verification periods, and the temporary controls corresponding to the two verification periods do not exist simultaneously.

[0098] Through the above implementation process, the first verification period is used to obtain the first verification human voice sound pressure level after the center channel gain is temporarily increased, and the second verification period is used to obtain the second verification low-frequency sound pressure level after the low-frequency gain is temporarily decreased. The two verification periods do not overlap, and only one gain parameter is changed within each verification period. After each verification period ends, the audio output state is restored to the state before the start of the corresponding verification period. The first verification human voice sound pressure level and the second verification low-frequency sound pressure level can be used by S4 to determine the weight values ​​of the insufficient human voice intelligibility status flag and the low-frequency interference enhancement status flag.

[0099] In S4, within the voice-to-text association scenario, the audio-visual playback device sets a reference time period. This reference time period is located outside the first and second verification time periods. During this period, no temporary increase control of the center channel gain or temporary decrease control of the low-frequency gain is applied. The reference time period is used to obtain the sound pressure level state without temporary verification control. The reference time period can be selected from the sampling window used to determine candidate causes in S2, or from a stable segment still within the voice-to-text association scenario before the start of the first verification time period. The length of the reference time period is consistent with the sound pressure level sampling window in S2; in this embodiment, it is 1.00 seconds.

[0100] It should be noted that: the difference between the first verified human voice sound pressure level and the equivalent sound pressure level of the human voice frequency band referred to in S4 refers to the difference between the first verified human voice sound pressure level and the equivalent sound pressure level of the human voice frequency band obtained without applying temporary control to the center channel gain. This equivalent sound pressure level of the human voice frequency band without applying temporary control can be obtained from the reference time period. The difference between the equivalent sound pressure level of the low frequency band and the second verified low frequency sound pressure level referred to in S4 refers to the difference between the equivalent sound pressure level of the low frequency band obtained without applying temporary control to the low frequency gain and the second verified low frequency sound pressure level. This equivalent sound pressure level of the low frequency band without applying temporary control can be obtained from the reference time period.

[0101] The baseline time period can be determined as follows: The audio-visual playback device first determines the overlapping area between the voice activity range and the subtitle display range, then excludes the first verification time period, the second verification time period, and the recovery interval. Subsequently, from the remaining time period, a time period with a continuous duration of not less than 1.00s, without subtitle switching, total volume changes, or temporary channel gain control is selected as the baseline time period. If the remaining time period is less than 1.00s, the uncontrolled sampling window where candidate cause judgment has been completed in S2 is used as the baseline time period. Through this process, both the baseline time period and the verification time period correspond to the same human voice subtitle association scenario, which can reduce the sound pressure difference caused by different playback content.

[0102] Within the reference time period, the audio-visual playback device obtains the corresponding human voice reference sound pressure level and the corresponding low-frequency reference sound pressure level according to the delay compensation and equivalent sound pressure level extraction method in S2. The corresponding human voice reference sound pressure level represents the equivalent sound pressure level of the human voice frequency band at the user's listening position when no temporary increase control of the center channel gain is performed. The corresponding low-frequency reference sound pressure level represents the equivalent sound pressure level of the low-frequency frequency band at the user's listening position when no temporary decrease control of the low-frequency gain is performed.

[0103] The audio-visual playback device compares the first verified human voice sound pressure level with the corresponding human voice reference sound pressure level to obtain the human voice verification quantity. The human voice verification quantity represents the increase in the equivalent sound pressure level of the human voice frequency band relative to the reference time period after a temporary increase in the center channel gain. If the first verified human voice sound pressure level is higher than the corresponding human voice reference sound pressure level, the difference between the two is taken as the human voice verification quantity; if the first verified human voice sound pressure level is not higher than the corresponding human voice reference sound pressure level, the human voice verification quantity is recorded as 0.

[0104] The audio-visual playback device compares the second verified low-frequency sound pressure level (SPL) with the corresponding low-frequency reference SPL to obtain the low-frequency verification quantity. The low-frequency verification quantity represents the reduction in the equivalent SPL of the low-frequency band relative to the reference time period after a temporary reduction in low-frequency gain. If the second verified low-frequency SPL is lower than the corresponding low-frequency reference SPL, the difference between the corresponding low-frequency reference SPL and the second verified low-frequency SPL is used as the low-frequency verification quantity; if the second verified low-frequency SPL is not lower than the corresponding low-frequency reference SPL, the low-frequency verification quantity is recorded as 0.

[0105] In one implementation, the audio-visual playback device uses the difference between the corresponding low-frequency reference sound pressure level and the corresponding human voice reference sound pressure level as the reference masking difference, the difference between the second verified low-frequency sound pressure level and the second verified human voice auxiliary sound pressure level as the verification masking difference, and the difference between the reference masking difference and the verification masking difference as the low-frequency masking improvement indicator.

[0106] A preset voice boost threshold is used to determine whether the voice sound pressure level increase at the user's listening position reaches an effective level after a temporary increase in the center channel gain. This threshold can be determined by the microphone sound pressure level measurement fluctuation, the short-term fluctuation of the indoor sound field, and the amount of temporary gain in the center channel. During the device's factory testing phase, the audio-visual playback device plays dialogue samples multiple times in a standard living room environment, recording the natural fluctuation range of the equivalent sound pressure level of the voice frequency band without changing the gain; and recording the range of voice sound pressure level increase that can be stably measured at the user's listening position after a temporary increase in the center channel gain. The device writes the minimum boost amount that is significantly higher than the natural fluctuation range and reflects the actual arrival of the temporary gain in the center channel at the user's position into the parameter table as the preset voice boost threshold. In this embodiment, the microphone measurement fluctuation is approximately 0.5dB, the short-term indoor fluctuation is approximately 0.6dB, and the boost amount that can be stably measured at the user's position after a 3dB temporary increase in the center channel is usually greater than 1.5dB. Therefore, the preset voice boost threshold is set to 1.5dB.

[0107] A preset low-frequency reduction threshold is used to determine whether the reduction in low-frequency sound pressure at the user's listening position is effective after a temporary reduction in low-frequency gain. This threshold can be determined by microphone measurement fluctuations, room low-frequency response fluctuations, and the amount of temporary low-frequency attenuation. During the device's factory testing or room calibration phase, the audio-visual playback device plays a low-frequency test segment and a dialogue segment containing low-frequency background noise, recording the natural fluctuation range of the equivalent sound pressure level in the low-frequency band without changing the low-frequency gain; then, it records the range of low-frequency sound pressure reduction that can be stably measured at the user's listening position after a temporary reduction in low-frequency gain. The device writes the minimum reduction amount that is significantly higher than the natural fluctuation range and reflects the actual effect of the temporary low-frequency attenuation on the user's position into a parameter table as the preset low-frequency reduction threshold. In this embodiment, the short-term low-frequency fluctuation in the room is approximately 0.8 dB, and the amount of reduction that can be stably measured at the user's position after a 4 dB temporary reduction in low-frequency gain is usually greater than 2.0 dB; therefore, the preset low-frequency reduction threshold is set to 2.0 dB.

[0108] The audio-visual playback device compares the voice verification value with a preset voice enhancement threshold. When the voice verification value reaches the preset voice enhancement threshold, the insufficient voice clarity status indicator passes verification, and the weight value of the insufficient voice clarity status indicator is determined based on the amount by which the voice verification value exceeds the preset voice enhancement threshold. When the voice verification value does not reach the preset voice enhancement threshold, it is determined that the temporary increase in center channel gain has not resulted in an effective voice sound pressure level boost, and the weight value of the insufficient voice clarity status indicator is determined to be invalid. An invalid weight can be recorded as 0, or as an invalid flag in the parameter table.

[0109] The audio-visual playback device compares a low-frequency verification value with a preset low-frequency reduction threshold. When the low-frequency verification value reaches the preset low-frequency reduction threshold, the low-frequency interference enhancement status indicator is deemed to have passed verification, and its weight value is determined based on the amount by which the low-frequency verification value exceeds the preset low-frequency reduction threshold. When the low-frequency verification value does not reach the preset low-frequency reduction threshold, it is determined that the temporary reduction in low-frequency gain has not resulted in an effective low-frequency sound pressure reduction, and the weight value of the low-frequency interference enhancement status indicator is determined to be invalid.

[0110] When the low-frequency verification quantity reaches the preset low-frequency reduction threshold but the low-frequency masking improvement indication quantity does not reach the preset masking improvement threshold, the audio-visual playback device can determine the weight value of the low-frequency interference enhancement status indicator as an invalid weight or reduce the weight value to avoid misjudging the low-frequency sound pressure reduction capability as dialogue masking improvement capability.

[0111] The weight values ​​are determined based on the effectiveness of the verification quantity exceeding the corresponding preset threshold. The audio-visual playback device first determines the amount by which the human voice verification quantity exceeds the preset human voice enhancement threshold, and the amount by which the low-frequency verification quantity exceeds the preset low-frequency reduction threshold. The larger the excess, the more significant the improvement in sound pressure at the user's listening position by the temporary verification control, and the higher the credibility of the corresponding candidate cause.

[0112] During the factory testing phase, calibration data is established for different device models, speaker configurations, and playback modes. For the insufficient vocal clarity status indicator, the device repeatedly performs temporary center channel gain increase verification in a standard listening environment. The distribution range of vocal verification exceeding a preset vocal boost threshold is recorded, and the upper limit of this distribution range that can be stably repeated is written into a parameter table as a reference value for vocal weight normalization. For the enhanced low-frequency interference status indicator, the device repeatedly performs temporary low-frequency gain decrease verification in a standard listening environment and room samples with different low-frequency responses. The distribution range of low-frequency verification exceeding a preset low-frequency decrease threshold is recorded, and the upper limit of this distribution range that can be stably repeated is written into a parameter table as a reference value for low-frequency weight normalization.

[0113] During operation, when the voice verification quantity reaches the preset voice enhancement threshold, the audio-visual playback device determines the weight value of the insufficient voice clarity status indicator based on the relative position of the excess amount of the voice verification quantity within the voice weight normalization reference value. The closer the excess amount of the voice verification quantity is to the voice weight normalization reference value, the closer the weight value of the insufficient voice clarity status indicator is to the highest effective weight. When the voice verification quantity does not reach the preset voice enhancement threshold, the weight value of the insufficient voice clarity status indicator is determined to be an invalid weight.

[0114] When the low-frequency verification quantity reaches the preset low-frequency reduction threshold, the audio-visual playback device determines the weight value of the low-frequency interference enhancement status indicator based on the relative position of the excess amount of the low-frequency verification quantity within the low-frequency weight normalization reference value. The closer the excess amount of the low-frequency verification quantity is to the low-frequency weight normalization reference value, the closer the weight value of the low-frequency interference enhancement status indicator is to the highest effective weight. When the low-frequency verification quantity does not reach the preset low-frequency reduction threshold, the weight value of the low-frequency interference enhancement status indicator is determined to be an invalid weight.

[0115] Specifically, the audio-visual playback device can determine the weight value of the voice clarity insufficiency status indicator according to Wvoice=min{1, max{0, (Vvoice-Tvoice) / Rvoice}}, and determine the weight value of the low-frequency interference enhancement status indicator according to Wlow=min{1, max{0, (Vlow-Tlow) / Rlow}}. Wherein, Wvoice is the weight value of the voice clarity insufficiency status indicator, Vvoice is the voice verification quantity, Tvoice is the preset voice enhancement threshold, Rvoice is the voice weight normalization reference value, Wlow is the weight value of the low-frequency interference enhancement status indicator, Vlow is the low-frequency verification quantity, Tlow is the preset low-frequency reduction threshold, and Rlow is the low-frequency weight normalization reference value.

[0116] To facilitate control, audio-visual playback devices can convert continuously obtained weight values ​​into control levels. These control levels are stored in a device parameter table, and the level boundaries in the parameter table are formed by the aforementioned calibration data. For example, when the normalized reference value for human voice weight is 2.5dB, a human voice verification quantity exceeding the preset human voice enhancement threshold by 0.8dB corresponds to a lower effective weight; an exceedance of 1.7dB corresponds to a medium effective weight; and an exceedance close to or reaching 2.5dB corresponds to a higher effective weight. The weight value of the low-frequency interference enhancement status indicator can be determined in the same way by the exceedance amount of the low-frequency verification quantity.

[0117] In a specific example, the reference sound pressure level for human voices during the baseline time period is 55.8 dB, and the reference sound pressure level for low frequencies is 72.6 dB. The first verification human voice sound pressure level during the first verification time period is 59.2 dB, and the second verification low-frequency sound pressure level during the second verification time period is 68.9 dB. Therefore, the human voice verification quantity is 3.4 dB, and the low-frequency verification quantity is 3.7 dB. The preset human voice enhancement threshold is 1.5 dB, and the preset low-frequency reduction threshold is 2.0 dB. The human voice verification quantity exceeds the preset human voice enhancement threshold by 1.9 dB, corresponding to a human voice intelligibility insufficiency status flag weight of 0.7. The low-frequency verification quantity exceeds the preset low-frequency reduction threshold by 1.7 dB, corresponding to a low-frequency interference enhancement status flag weight of 0.7.

[0118] In another specific example, the corresponding human voice reference sound pressure level is 56.0 dB, the first verification human voice sound pressure level is 56.8 dB, and the human voice verification quantity is 0.8 dB. Since this does not reach the preset human voice enhancement threshold of 1.5 dB, the weight value of the insufficient human voice clarity status indicator is determined to be an invalid weight. The corresponding low-frequency reference sound pressure level is 72.0 dB, the second verification low-frequency sound pressure level is 68.8 dB, and the low-frequency verification quantity is 3.2 dB. This reaches the preset low-frequency reduction threshold of 2.0 dB, so the weight value of the low-frequency interference enhancement status indicator is determined to be a valid weight. In this case, the audio-visual linkage control command does not include the center channel gain adjustment parameter corresponding to the insufficient human voice clarity status indicator, but it may include the low-frequency gain adjustment parameter corresponding to the low-frequency interference enhancement status indicator.

[0119] The audio-visual playback device generates audio-visual linkage control commands based on the weight values ​​of the determined candidate causes. When the weight value of the insufficient voice clarity status indicator is an effective weight, the audio-visual linkage control command includes a center channel gain adjustment parameter. The center channel gain adjustment parameter can be determined by looking up a table using the weight value of the insufficient voice clarity status indicator. For example, when the weight value is 0.4, the center channel gain adjustment is +1dB; when the weight value is 0.7, the center channel gain adjustment is +2dB; and when the weight value is 1.0, the center channel gain adjustment is +3dB.

[0120] When the weight value of the low-frequency interference enhancement status indicator is an effective weight, the audio-visual linkage control command includes low-frequency gain adjustment parameters. The low-frequency gain adjustment parameters can be determined by looking up the weight value of the low-frequency interference enhancement status indicator in a table. For example, when the weight value is 0.4, the low-frequency gain adjustment is -1dB; when the weight value is 0.7, the low-frequency gain adjustment is -2dB; and when the weight value is 1.0, the low-frequency gain adjustment is -3dB.

[0121] The audio-visual linkage control instructions also include subtitle enhancement parameters. These parameters can include one or more of the following: subtitle brightness enhancement, subtitle outline enhancement, and subtitle background transparency adjustment. The subtitle enhancement parameters are determined by weight values ​​of pre-defined candidate causes. Specifically, when the weight value of the insufficient voice clarity status indicator or the weight value of the low-frequency interference enhancement status indicator is between 0.4 and 0.7, the subtitle brightness enhancement is set to a first brightness increment, and the subtitle outline width is set to a first outline width. When either weight value is greater than 0.7, the subtitle brightness enhancement is set to a second brightness increment, and the subtitle outline width is set to a second outline width, with the second brightness increment being greater than the first brightness increment and the second outline width being greater than the first outline width.

[0122] When both the weight values ​​for the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator are valid, the audio-visual linkage control command includes center channel gain adjustment parameters, low-frequency gain adjustment parameters, and subtitle enhancement parameters. The center channel gain adjustment parameter is determined by the weight value of the insufficient voice clarity status indicator, the low-frequency gain adjustment parameter is determined by the weight value of the low-frequency interference enhancement status indicator, and the subtitle enhancement parameter is jointly determined by the weight values ​​of both the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator. When either the insufficient voice clarity status indicator or the low-frequency interference enhancement status indicator has an invalid weight, the audio-visual linkage control command does not include the audio adjustment parameter corresponding to the invalid weight.

[0123] Example 2

[0124] Please refer to Figures 1-4 Specifically: This embodiment illustrates that even when only a status indicator indicating insufficient human voice clarity exists, the method of the present invention can still achieve a relatively good control effect.

[0125] The audio-visual playback device identifies the scene where voice and subtitles are associated in the voice enhancement adjustment mode. In the current segment, the voice activity range is from 210.00s to 214.50s, and the subtitle display range is from 210.10s to 214.70s. When the overlap ratio between the voice activity range and the subtitle display range reaches a preset ratio, the current segment is identified as a scene where voice and subtitles are associated.

[0126] In the scenario of human voice subtitle association, the device performs delay compensation on the microphone signal and extracts the equivalent sound pressure level (SPL) of the human voice frequency band and the equivalent SPL of the low-frequency band. The equivalent SPL of the human voice frequency band is 53dB, which is lower than the preset human voice SPL threshold of 60dB. The equivalent SPL of the low-frequency band is 64dB, which does not reach the preset low-frequency SPL threshold of 70dB. Therefore, the device determines that the insufficient human voice clarity status is valid, while the low-frequency interference enhancement status is invalid.

[0127] The device determines the first verification period within the human voice captioning scenario. During the first verification period, the device temporarily increases the center channel gain by 3dB, while keeping the low-frequency gain unchanged. After the first verification period ends, the device obtains the first verified human voice sound pressure level. If the corresponding human voice reference sound pressure level is 53dB and the first verified human voice sound pressure level is 57dB, then the human voice verification quantity is 4dB. If the human voice verification quantity reaches the preset human voice enhancement threshold of 2dB, the insufficient human voice clarity status is determined as a valid cause, and a weight value is determined based on the amount by which the human voice verification quantity exceeds the preset human voice enhancement threshold.

[0128] In this embodiment, since the low-frequency interference enhancement status flag is not valid, the device does not generate low-frequency gain adjustment parameters. If the control process needs to reserve a second verification period in the human voice subtitle association scene, this second verification period can be used as an uncontrolled observation period, where the low-frequency gain remains unchanged, and the sound pressure level during this observation period is not used to determine the effective weight of the low-frequency interference enhancement status flag. The audio-visual linkage control command includes center channel gain adjustment parameters and subtitle enhancement parameters. The center channel gain adjustment parameters can officially increase the gain of the center channel by the amount corresponding to the weight value. The subtitle enhancement parameters can improve the visibility of subtitles, allowing users to obtain visual assistance in weak dialogue segments.

[0129] In this embodiment, if a normal overall volume boosting method is used, background music, ambient sounds, and sound effects will also increase, potentially causing auditory discomfort for the user. This invention generates center channel gain adjustment parameters only for verified and effective voice intelligibility deficiency indicators, making the control object more focused. Therefore, even in non-composite scenarios where only voice intelligibility is present, this invention can still achieve more targeted dialogue enhancement through verification quantities and weight values.

[0130] Example 3

[0131] Please refer to Figures 1-4 Specifically: This embodiment illustrates that even when only a low-frequency interference enhancement status indicator exists, the method of the present invention can still achieve a better control effect.

[0132] The audio-visual playback device identifies a scene associated with human voice subtitles in human voice enhancement adjustment mode. In this scene, the equivalent sound pressure level of the human voice band is 62dB, which is not lower than the preset human voice sound pressure threshold of 60dB. The equivalent sound pressure level of the low-frequency band is 78dB, which is higher than the preset low-frequency sound pressure threshold of 70dB. Therefore, the device determines that the low-frequency interference enhancement status indicator is valid, while the insufficient human voice clarity status indicator is invalid.

[0133] The device determines a second verification period within the human voice subtitle association scenario. During this second verification period, the device temporarily reduces the low-frequency gain by 4dB, while keeping the center channel gain unchanged. After the second verification period ends, the device obtains the second verification low-frequency sound pressure level and simultaneously obtains the second verification human voice auxiliary sound pressure level. If the corresponding low-frequency reference sound pressure level is 78dB and the second verification low-frequency sound pressure level is 73dB, then the low-frequency verification quantity is 5dB. If the corresponding human voice reference sound pressure level is 62dB and the second verification human voice auxiliary sound pressure level is 62.5dB, then the reference masking difference is 16dB, the verification masking difference is 10.5dB, and the low-frequency masking improvement indication quantity is 5.5dB. If the low-frequency verification quantity reaches the preset low-frequency reduction threshold of 2dB and the low-frequency masking improvement indication quantity reaches the preset masking improvement threshold, the low-frequency interference enhancement status indicator is determined as a valid cause, and a weight value is determined based on the amount by which the low-frequency verification quantity exceeds the preset low-frequency reduction threshold.

[0134] In this embodiment, since the insufficient voice clarity status indicator is invalid, the device does not generate center channel gain adjustment parameters. If the control process requires reserving a first verification period, this first verification period can be used as an uncontrolled observation period, where the center channel gain remains unchanged, and the sound pressure level during this observation period is not used to determine the effective weight of the insufficient voice clarity status indicator. The audio-visual linkage control instructions include low-frequency gain adjustment parameters and subtitle enhancement parameters. The low-frequency gain adjustment parameters are used to reduce the low-frequency band output and reduce the impact of low-frequency energy on dialogue clarity. The subtitle enhancement parameters are determined based on the weight value of the low-frequency interference enhancement status indicator and are used to enhance visual assistance in dialogue scenes where low-frequency obstruction is significant.

[0135] In this embodiment, since the insufficient voice clarity status indicator is invalid, the device does not perform a temporary increase verification of the center channel gain, nor does it generate center channel gain adjustment parameters. The audio-visual linkage control commands include low-frequency gain adjustment parameters and subtitle enhancement parameters. The low-frequency gain adjustment parameters are used to reduce the low-frequency band output, minimizing the impact of low-frequency energy on dialogue clarity. The subtitle enhancement parameters are determined based on the weight value of the low-frequency interference enhancement status indicator and are used to enhance visual assistance in dialogue scenes where low-frequency obstruction is significant.

[0136] In this embodiment, the human voice itself is not low, but the low-frequency sound field is strong. If traditional human voice enhancement methods are used to directly increase the center channel, the dialogue may become louder, but low-frequency masking still exists, or even cause an increase in overall loudness. This invention verifies and confirms the cause of low-frequency masking by temporarily reducing the low-frequency gain, and then generates low-frequency gain adjustment parameters, which can more accurately solve the problem of low-frequency masking of dialogue.

[0137] Example 4

[0138] Please refer to Figures 1-4 Specifically: This embodiment illustrates that when two candidate causes are initially established, but one of the candidate causes is found to be invalid after verification, the method of the present invention can avoid error control.

[0139] In a scenario involving the association of human voice with subtitles, the device extracts the equivalent sound pressure level (SPL) of the human voice frequency band and the equivalent SPL of the low-frequency band. The equivalent SPL of the human voice frequency band is 55 dB, which is lower than the preset human voice SPL threshold of 60 dB. The equivalent SPL of the low-frequency band is 72 dB, which is higher than the preset low-frequency SPL threshold of 70 dB. The device initially determines that both the insufficient human voice clarity status indicator and the enhanced low-frequency interference status indicator are valid.

[0140] The device first determines the first verification period. During this period, the center channel gain is temporarily increased only, and the first verification voice sound pressure level is obtained. If the corresponding baseline voice sound pressure level is 55dB and the first verification voice sound pressure level is 59dB, then the voice verification amount is 4dB. When the voice verification amount reaches the preset voice enhancement threshold of 2dB, the weight value of the voice intelligibility insufficient status indicator is determined as the effective weight.

[0141] Subsequently, during the second verification period, the device only temporarily reduced the low-frequency gain and obtained the second verified low-frequency sound pressure level. If the corresponding low-frequency reference sound pressure level is 72dB and the second verified low-frequency sound pressure level is 71dB, then the low-frequency verification quantity is 1dB. Since the low-frequency verification quantity did not reach the preset low-frequency reduction threshold of 2dB, the weight value of the low-frequency interference enhancement status indicator was determined to be an invalid weight.

[0142] In this embodiment, the audio-visual linkage control command does not include the low-frequency gain adjustment parameter corresponding to the invalid weight. The audio-visual linkage control command includes the center channel gain adjustment parameter corresponding to the valid weight, and may include subtitle enhancement parameters. Therefore, the device will not directly reduce the low-frequency gain simply because the equivalent sound pressure level in the low-frequency band exceeds the threshold in the initial judgment.

[0143] This embodiment illustrates that the candidate causes of the present invention do not directly determine the formal control. Candidate causes need to undergo time-division verification and weight determination before they can participate in audio-visual linkage control. This method avoids misinterpreting a short-term low-frequency rise as continuous low-frequency masking, reducing sound effect loss caused by erroneous low-frequency reduction.

[0144] Example 5

[0145] Please refer to Figures 1-4 Specifically: This embodiment illustrates that when both candidate causes are valid and composite control is required, the method of the present invention can produce a better linkage control result.

[0146] In a scenario involving the association of human voice with subtitles, the device detected an equivalent sound pressure level of 52dB in the human voice frequency band, which is lower than the preset human voice sound pressure level threshold of 60dB. The equivalent sound pressure level in the low-frequency band was 79dB, which is higher than the preset low-frequency sound pressure level threshold of 70dB. The device determined that both the insufficient human voice clarity status indicator and the enhanced low-frequency interference status indicator were valid.

[0147] The device determines a first verification period and a second verification period. During the first verification period, only the center channel gain is changed; during the second verification period, only the low-frequency gain is changed. The first and second verification periods do not overlap. After the first verification period ends, the device restores the audio output state to its state before the start of the first verification period, and then enters the second verification period. After the second verification period ends, the device restores the audio output state to its state before the start of the second verification period.

[0148] The device obtains the reference sound pressure level (SPL) for human voice and the reference SPL for low frequency during a reference time period. This reference time period is outside the first and second verification time periods. If the reference SPL for human voice is 52 dB and the first verification human voice SPL is 57 dB, then the human voice verification quantity is 5 dB. If the reference SPL for low frequency is 79 dB and the second verification low frequency SPL is 73 dB, then the low frequency verification quantity is 6 dB. Both the human voice verification quantity and the low frequency verification quantity reach their corresponding preset thresholds. The device determines the weight values ​​for the insufficient human voice clarity status indicator and the enhanced low frequency interference status indicator based on the corresponding exceedance amounts.

[0149] In this embodiment, the audio-visual linkage control commands include center channel gain adjustment parameters, low-frequency gain adjustment parameters, and subtitle enhancement parameters. The center channel gain adjustment parameter is determined by the weight value of the insufficient voice clarity status indicator. The low-frequency gain adjustment parameter is determined by the weight value of the low-frequency interference enhancement status indicator. The subtitle enhancement parameter is jointly determined by the weight values ​​of the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator. For example, the device can officially increase the center channel by 3dB, officially decrease the low frequency by 3dB, and simultaneously enhance the subtitle contrast.

[0150] In this embodiment, if only the center channel is enhanced, low-frequency masking may still exist. If only the low frequencies are reduced, the problem of insufficient vocal quality may still exist. This invention obtains the weight values ​​of the two candidate causes through time-division verification, and then generates a composite audio-visual linkage control command, so that center channel adjustment, low-frequency adjustment, and subtitle enhancement are driven by the same verification result. This method can achieve a more stable improvement in dialogue clarity under complex acoustic problems.

[0151] Example 6

[0152] Please refer to Figures 1-4 An audio-visual linkage control system based on multimodal scene recognition includes: a scene recognition module, a candidate cause determination module, a time-division verification module, and a weight control module.

[0153] The scene recognition module is used to acquire the center channel audio signal, subtitle control signal, and microphone acquisition signal. After the audio-visual playback device enters the human voice enhancement adjustment mode, the scene recognition module determines the voice activity range from the center channel audio signal and the subtitle display segment from the subtitle control signal. When the overlap ratio between the voice activity range and the subtitle display segment reaches a preset ratio, the scene associated with the human voice and subtitle is determined.

[0154] The candidate cause determination module is used to perform delay compensation on the microphone acquisition signal in the context of human voice subtitle association, obtaining the compensated microphone acquisition signal. The module extracts the equivalent sound pressure level (SPL) of the human voice frequency band and the equivalent sound pressure level of the low-frequency frequency band from the compensated microphone acquisition signal, and based on these SPLs, determines the status indicators for insufficient human voice clarity and enhanced low-frequency interference, respectively.

[0155] The time-division verification module is used to determine a first and second verification period that do not overlap within the context of a voice-subtitle association scene. For a confirmed state indicating insufficient voice clarity, the time-division verification module performs a temporary increase in center channel gain during the first verification period to obtain the first verified voice sound pressure level. For a confirmed state indicating increased low-frequency interference, the time-division verification module performs a temporary decrease in low-frequency gain during the second verification period to obtain the second verified low-frequency sound pressure level.

[0156] The weight control module is used to obtain a voice verification quantity based on the first verified human voice sound pressure level and the corresponding human voice reference sound pressure level, and to obtain a low-frequency verification quantity based on the second verified low-frequency sound pressure level and the corresponding low-frequency reference sound pressure level. The weight control module determines the weight value of the insufficient voice clarity status indicator based on the voice verification quantity, determines the weight value of the low-frequency interference enhancement status indicator based on the low-frequency verification quantity, and generates audio-visual linkage control commands based on the weight values ​​of the determined candidate causes.

[0157] In this embodiment, the scene recognition module, candidate cause determination module, time-sharing verification module, and weight control module can be implemented by the processor of the audio-visual playback device executing corresponding programs, or they can be implemented collaboratively by the audio processing chip, display controller, and microphone signal processing unit. Through the cooperation of these modules, the system can first identify the scene associated with human voice and subtitles, and then perform candidate cause judgment, time-sharing verification, and weight control. This allows the audio-visual linkage control to not rely on a single volume judgment, thereby improving the accuracy and stability of dialogue clarity adjustment.

[0158] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

Claims

1. A method for audio-visual linkage control based on multimodal scene recognition, characterized in that, Includes the following steps: S1. Acquire the center channel audio signal, subtitle control signal and microphone acquisition signal. After the audio-visual playback device enters the human voice enhancement adjustment mode, determine the voice activity range and subtitle display segment based on the center channel audio signal and subtitle control signal, and determine the human voice subtitle association scene based on the voice activity range and the subtitle display segment. S2. In the human voice subtitle association scenario, delay compensation is performed on the microphone acquisition signal to obtain the compensated microphone acquisition signal, and the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low frequency band are extracted. Based on the equivalent sound pressure level of the human voice frequency band and the equivalent sound pressure level of the low frequency band, the status indicators of insufficient human voice clarity and enhanced low frequency interference are determined respectively. S3. Within the human voice subtitle association scenario, determine a first verification period and a second verification period that do not overlap. Based on the first verification period, perform a temporary increase verification of the center channel gain for the determined state indicator of insufficient human voice clarity, and obtain the first verification human voice sound pressure level. Based on the second verification period, perform a temporary decrease verification of the low frequency gain for the determined state indicator of enhanced low frequency interference, and obtain the second verification low frequency sound pressure level. S4. Based on the difference between the first verified human voice sound pressure level and the equivalent sound pressure level of the human voice frequency band, determine the weight value of the human voice insufficiency status indicator; based on the difference between the equivalent sound pressure level of the low frequency band and the second verified low frequency sound pressure level, determine the weight value of the low frequency interference enhancement status indicator; and based on the weight values ​​of the determined candidate causes, generate an audio-visual linkage control command.

2. The audio-visual linkage control method based on multimodal scene recognition according to claim 1, characterized in that, In S1, the overlap duration is determined based on the voice activity interval and the subtitle display segment, and the overlap duration is the ratio of the voice activity interval duration to the overlap ratio. When the overlap ratio between the voice activity interval and the subtitle display segment reaches a preset ratio, the human voice subtitle association scene is determined; The subtitle display segment appears synchronously with the voice activity interval, and the time difference between the start time of the subtitle display segment and the start time of the voice activity interval is less than a preset synchronization deviation threshold.

3. The audio-visual linkage control method based on multimodal scene recognition according to claim 1, characterized in that, In S2, when the equivalent sound pressure level of the human voice frequency band is lower than the preset human voice sound pressure threshold, it is determined that the human voice intelligibility is insufficient. When the equivalent sound pressure level in the low-frequency band is higher than the preset low-frequency sound pressure threshold, the low-frequency interference enhancement status indicator is determined to be valid.

4. The audio-visual linkage control method based on multimodal scene recognition according to claim 1, characterized in that, In S3, when both the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator are established, only the center channel gain is changed during the first verification period, and only the low-frequency gain is changed during the second verification period. The first verification period and the second verification period do not overlap.

5. The audio-visual linkage control method based on multimodal scene recognition according to claim 4, characterized in that, In S3, after the first verification period ends, the audio output state is restored to the audio output state before the start of the first verification period, and then the second verification period begins. After the second verification period ends, the audio output state will be restored to the audio output state before the start of the second verification period.

6. The audio-visual linkage control method based on multimodal scene recognition according to claim 1, characterized in that, In S4, the sound pressure level of the first verified human voice is compared with the corresponding human voice reference sound pressure level to obtain the human voice verification quantity; The second verification low-frequency sound pressure level is compared with the corresponding low-frequency reference sound pressure level to obtain the low-frequency verification quantity. The human voice subtitle association scene is set with a reference time period, which is located outside the first verification time period and the second verification time period. The reference time period includes the corresponding human voice reference sound pressure level and low frequency reference sound pressure level.

7. The audio-visual linkage control method based on multimodal scene recognition according to claim 6, characterized in that, In S4, the voice verification quantity is compared with a preset voice enhancement threshold. When the voice verification quantity reaches the preset voice enhancement threshold, the weight value of the voice clarity insufficiency status indicator is determined based on the amount by which the voice verification quantity exceeds the preset voice enhancement threshold. The low-frequency verification quantity is compared with a preset low-frequency reduction threshold. When the low-frequency verification quantity reaches the preset low-frequency reduction threshold, the weight value of the low-frequency interference enhancement status indicator is determined based on the amount by which the low-frequency verification quantity exceeds the preset low-frequency reduction threshold.

8. The audio-visual linkage control method based on multimodal scene recognition according to claim 7, characterized in that, In S4, when the number of human voice verifications does not reach the preset human voice enhancement threshold, the weight value of the human voice insufficiency status identifier is determined as an invalid weight. When the low-frequency verification quantity does not reach the preset low-frequency reduction threshold, the weight value of the low-frequency interference enhancement status identifier is determined to be an invalid weight. The audio-visual linkage control instructions do not include audio adjustment parameters corresponding to invalid weights.

9. The audio-visual linkage control method based on multimodal scene recognition according to claim 7, characterized in that, In S4, when both the weight value of the insufficient voice clarity status indicator and the weight value of the low-frequency interference enhancement status indicator are valid weights, the audio-visual linkage control command includes the center channel gain adjustment parameter, the low-frequency gain adjustment parameter, and the subtitle enhancement parameter. The center channel gain adjustment parameter is determined by the weight value of the insufficient voice clarity status indicator; The low-frequency gain adjustment parameter is determined by the weight value of the low-frequency interference enhancement status indicator; The subtitle enhancement parameters are jointly determined by the weight values ​​of the insufficient voice clarity status indicator and the low-frequency interference enhancement status indicator.

10. An audio-visual linkage control system based on multimodal scene recognition, based on the audio-visual linkage control method based on multimodal scene recognition as described in any one of claims 1-9, characterized in that, It includes a scene recognition module, a candidate cause determination module, a time-sharing verification module, and a weight control module; The scene recognition module acquires the center channel audio signal, subtitle control signal, and microphone acquisition signal. After the audio-visual playback device enters the human voice enhancement adjustment mode, the central channel audio signal determines the voice activity range, the subtitle control signal determines the subtitle display segment, and when the overlap ratio between the voice activity range and the subtitle display segment reaches a preset ratio, the human voice subtitle associated scene is determined. The candidate cause determination module performs delay compensation on the microphone acquisition signal in the human voice subtitle association scenario to obtain a compensated microphone acquisition signal, and extracts the human voice frequency band equivalent sound pressure level and the low frequency band equivalent sound pressure level from the compensated microphone acquisition signal. Based on the human voice frequency band equivalent sound pressure level and the low frequency band equivalent sound pressure level, it determines the human voice intelligibility insufficiency status indicator and the low frequency interference enhancement status indicator, respectively. The time-division verification module determines a first verification period and a second verification period that do not overlap within the human voice subtitle association scene. Based on the first verification period, it performs a temporary increase verification of the center channel gain for the determined state of insufficient human voice clarity and obtains the first verification human voice sound pressure level. Based on the second verification period, it performs a temporary decrease verification of the low frequency gain for the determined state of enhanced low frequency interference and obtains the second verification low frequency sound pressure level. The weight control module determines the weight value of the voice intelligibility insufficiency status indicator based on the difference between the first verified human voice sound pressure level and the equivalent sound pressure level of the human voice frequency band, determines the weight value of the low-frequency interference enhancement status indicator based on the difference between the equivalent sound pressure level of the low-frequency frequency band and the second verified low-frequency sound pressure level, and generates an audio-visual linkage control command including audio adjustment parameters and subtitle enhancement parameters based on the weight values ​​of the determined candidate causes.