Intelligent television audio and video playing control system and method based on voice recognition

CN121711528BActive Publication Date: 2026-08-28HUIXIN ELECTRIC APPLIANCE (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511966346.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-08-28
Estimated Expiration
2045-12-24

AI Technical Summary

Technical Problem

[0003]然而,现有基于语音识别的智能电视音视频播放控制同步获取并利用包含用户面部特征的视频流信息,使得无法通过用户视觉运动特征,导致易将电视自身播放音频、家庭环境背景噪音误判为有效用户语音指令,或遗漏用户低音量发声但伴随明确视觉动作的有效指令,从而造成指令识别准确率低,播放更新请求与实际用户操控意图不匹配,导致智能电视音视频播放控制的可靠性不足

Benefits of technology

同步获取语音识别的智能电视在视听环境中的音频流信息与包含用户面部特征的视频流信息;对所述音频流信息和所述视频流信息进行协同播放查验,得到视听环境中已播放时刻的音频能量特征和用户的视觉运动特征,基于所述音频能量特征和所述视觉运动特征判定当前是否存在有效的用户语音指令输入时段,并生成该用户语音指令输入时段对应的语音活动标识;在判定为有效语音指令输入时段内,对所述音频流信息进行进度分离,得到目标用户对智能电视进行音视频播放控制时的播放重建指令,将所述播放重建指令转换为智能电视音视频播放信息对应的播放更新请求,进而通过所述播放更新请求确定智能电视在音视频的当前播放进度中进行实时控制的同步播放条件;依据所述语音活动标识和所述同步播放条件确定智能电视在音视频播放时的播放控制需求,并根据所述播放控制需求进行音视频播放。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121711528B_ABST
    Figure CN121711528B_ABST
Patent Text Reader

Abstract

The application provides a voice recognition-based intelligent television audio and video playing control system and method, relates to the technical field of audio and video control, and cooperatively plays and checks audio stream information and video stream information, determines whether there is an effective user voice instruction input period based on audio energy features and visual motion features, and generates a voice activity identifier corresponding to the user voice instruction input period; determines a playing reconstruction instruction of a target user when the target user controls the intelligent television to play audio and video, determines a synchronous playing condition of the intelligent television in real-time control of the current playing progress of the audio and video through a playing update request; determines a playing control requirement of the intelligent television when the intelligent television plays audio and video according to the voice activity identifier and the synchronous playing condition, and plays the audio and video. The application can dynamically adapt the voice instruction recognition and playing control of the intelligent television in a complex audio and visual environment, so as to improve the accuracy of voice instruction playing control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio and video control technology, and more specifically, to a smart TV audio and video playback control system and method based on speech recognition. Background Technology

[0002] Audio and video control is widely used in the daily interaction scenarios of smart TVs, covering core operations such as playback progress adjustment, playback status switching, and volume adjustment. It is a key function to ensure the user's viewing experience. The real-time performance and accuracy of audio and video control directly affect the user's interactive experience. It needs to rely on multimodal data collaborative perception and efficient judgment logic to achieve accurate recognition of command validity and seamless execution of operations. This is an important foundation for solving existing control pain points and meeting the intelligent interaction needs of smart TVs.

[0003] However, existing voice recognition-based smart TV audio and video playback control methods simultaneously acquire and utilize video stream information including user facial features. This makes it impossible to recognize user visual motion characteristics, leading to misinterpretations of the TV's own audio playback and ambient noise as valid user voice commands, or the omission of valid commands delivered at low volume but accompanied by clear visual movements. This results in low command recognition accuracy, a mismatch between playback update requests and actual user intentions, and insufficient reliability of smart TV audio and video playback control. Therefore, how to dynamically adapt voice command recognition and playback control for smart TVs in complex audiovisual environments to improve the accuracy of voice command playback control is a challenge facing the industry. Summary of the Invention

[0004] This application provides a voice recognition-based intelligent TV audio and video playback control system and method, which can dynamically adapt the voice command recognition and playback control of intelligent TV in complex audio-visual environments to improve the accuracy of voice command playback control.

[0005] In a first aspect, this application provides a smart TV audio and video playback control method based on voice recognition, the control method comprising the following steps: Simultaneously acquire audio stream information and video stream information containing user facial features from smart TVs with voice recognition in the audiovisual environment; The audio stream information and the video stream information are collaboratively played and verified to obtain the audio energy characteristics and the user's visual motion characteristics at the time of playback in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it is determined whether there is a valid user voice command input period and a voice activity identifier corresponding to the user voice command input period is generated. During the period when a valid voice command is input, the audio stream information is separated into different stages to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV. The playback reconstruction command is then converted into a playback update request corresponding to the audio and video playback information of the smart TV. The synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video are then determined through the playback update request. The playback control requirements of the smart TV during audio and video playback are determined based on the voice activity identifier and the synchronous playback conditions, and the audio and video playback is performed according to the playback control requirements.

[0006] In this embodiment, the collaborative playback verification of the audio stream information and the video stream information to obtain the audio energy characteristics and user visual motion characteristics at the time of playback in the audiovisual environment specifically includes: Align the audio stream information with the video stream information to obtain a synchronized audio frame sequence and video frame sequence; The audio energy characteristics of the played moments are determined based on the audio frame sequence. The user's visual motion features are obtained by detecting the user's screen area and extracting motion vectors from the video frame sequence.

[0007] In this embodiment, the visual motion features refer to the features used to determine whether the user exhibits an active movement trend that matches the voice commands.

[0008] In this embodiment, the voice activity identifier refers to key information about the effective user voice command input period after merging and de-silencing processing.

[0009] In this embodiment, during the period determined to be a valid voice command input, the audio stream information is subjected to progress separation to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV. Specifically, this includes: Extract the user voice command audio segment from the audio stream information within the valid voice command input period; The audio segment of the user's voice command is sent to a speech recognition service cluster deployed in the cloud, and converted into a corresponding text command sequence; The text instruction sequence is matched with a preset playback control instruction keyword library to generate playback reconstruction instructions for the target user to control audio and video playback on the smart TV.

[0010] In this embodiment, converting the playback reconstruction instruction into a playback update request corresponding to the smart TV audio and video playback information specifically includes: The playback reconstruction instruction and the current audio and video playback status information are packaged into a playback instruction queue and pushed to the playback control service instance of the smart TV, waiting for the service instance to receive and confirm. The playback control service instance parses the received playback reconstruction instruction into playback control operation parameters according to the predefined instruction-action mapping rules; The playback control operation parameters are submitted to the main playback engine of the smart TV for verification, and a playback update request corresponding to the audio and video playback information of the smart TV is generated.

[0011] In this embodiment, the playback update request refers to a request message that conforms to the execution format of the main playback engine.

[0012] In this embodiment, the synchronous playback condition refers to the global prerequisite that the smart TV must meet to perform real-time playback control.

[0013] In this embodiment, determining the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronous playback conditions specifically includes: Extract user commands and intent classifications from the voice activity identifiers, and extract playback progress status and set of executable operations from the synchronous playback conditions; The intent classification and the playback progress status are submitted to a matching engine based on playback strategy rules to obtain the control adaptability of the user command in the current playback context. The playback control requirements of the smart TV during audio and video playback are determined based on the control adaptability and the set of executable operations.

[0014] Secondly, this application provides a voice recognition-based intelligent TV audio and video playback control system for executing a voice recognition-based intelligent TV audio and video playback control method, the control system comprising: The information acquisition module is used to simultaneously acquire audio stream information and video stream information containing user facial features from a smart TV with voice recognition in the audiovisual environment. The playback verification module is used to perform collaborative playback verification on the audio stream information and the video stream information to obtain the audio energy characteristics and the user's visual motion characteristics at the time of playback in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it determines whether there is a valid user voice command input period and generates a voice activity identifier corresponding to the user voice command input period. The progress separation module is used to separate the audio stream information during the period when it is determined to be a valid voice command input period, to obtain the playback reconstruction instruction when the target user controls the audio and video playback of the smart TV, to convert the playback reconstruction instruction into a playback update request corresponding to the audio and video playback information of the smart TV, and then to determine the synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video through the playback update request. The playback control module is used to determine the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronous playback conditions, and to play audio and video according to the playback control requirements.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The system simultaneously acquires audio stream information and video stream information containing user facial features from a smart TV with voice recognition in an audiovisual environment. It performs collaborative playback verification on the audio and video stream information to obtain audio energy features and user visual motion features at the time of playback in the audiovisual environment. Based on the audio energy features and visual motion features, it determines whether there is a valid user voice command input period and generates a voice activity identifier corresponding to that period. Within the determined valid voice command input period, it performs progress separation on the audio stream information to obtain the playback reconstruction instruction when the target user controls the audio and video playback of the smart TV. The playback reconstruction instruction is converted into a playback update request corresponding to the smart TV's audio and video playback information. The playback update request then determines the synchronous playback conditions for real-time control of the smart TV during the current playback progress of the audio and video. Based on the voice activity identifier and the synchronous playback conditions, it determines the smart TV's playback control requirements during audio and video playback and performs audio and video playback according to these requirements.

[0016] Therefore, this application demonstrates an improvement in the adaptability of voice-controlled smart TV audio and video playback. Specifically, by simultaneously acquiring audio stream information and video stream information containing user facial features from the audiovisual environment, it can eliminate spatiotemporal misalignment of audio and video based on timestamp alignment, overcoming the limitations of single-modal data acquisition. Through collaborative playback verification of audio and video stream information and extraction of audio energy features and visual motion features, invalid scenarios with only background noise or no voice action are eliminated, accurately locating the true effective command period and generating voice activity identifiers. Within the effective period, progress separation of the audio stream information yields playback reconstruction instructions, which are then converted into playback update requests and synchronous playback conditions are determined. This process removes TV-broadcast audio and environmental noise to obtain pure user voice, accurately parsing it into structured instructions. These instructions, combined with the current playback progress, decoding status, and buffer queue information, generate synchronous playback conditions adapted to the hardware capabilities, preventing command parameters from exceeding the hardware processing range. Based on voice activity identifiers and synchronous playback conditions, playback control requirements are determined and playback is executed. The system can combine the validity of the command with the hardware readiness status to filter highly adaptable control requirements, ensuring that the executed operation matches both the user's true intent and the hardware's real-time capabilities, avoiding invalid control or conflicting operations, and ultimately improving the accuracy of voice command recognition.

[0017] In summary, the technical solution adopted in this application can dynamically adapt to the voice command recognition and playback control of smart TVs in complex audiovisual environments, thereby improving the accuracy of voice command playback control. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is an exemplary flowchart of a smart TV audio and video playback control method based on speech recognition provided in this application; Figure 2 This is a flowchart illustrating the process of determining a voice activity identifier as provided in this application; Figure 3 This is a flowchart illustrating the process for determining synchronized playback conditions as provided in this application; Figure 4 This is a module structure diagram of a smart TV audio and video playback control system based on voice recognition, provided in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0021] This application provides a speech recognition-based smart TV audio and video playback control system and method. Its core is to simultaneously acquire audio stream information and video stream information containing user facial features from a smart TV in an audiovisual environment; perform collaborative playback verification on the audio stream information and the video stream information to obtain audio energy features and user visual motion features at the playback time in the audiovisual environment; determine whether there is a valid user voice command input period based on the audio energy features and the visual motion features, and generate a voice activity identifier corresponding to the user voice command input period; within the determined valid voice command input period, perform progress separation on the audio stream information to obtain the playback reconstruction instruction when the target user controls the audio and video playback of the smart TV; convert the playback reconstruction instruction into a playback update request corresponding to the smart TV audio and video playback information; and then determine the synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video through the playback update request; determine the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronous playback conditions, and perform audio and video playback according to the playback control requirements.

[0022] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown in the figure, this is an exemplary flowchart of a smart TV audio and video playback control method based on speech recognition according to this embodiment of the present application. The control method includes the following steps: In step S1, the audio stream information of the smart TV with voice recognition in the audiovisual environment and the video stream information containing the user's facial features are acquired simultaneously.

[0023] In specific implementation, firstly, a 4-microphone linear array can be integrated in the center of the bottom of the smart TV screen to collect ambient sound at a sampling frequency of 16kHz and a sampling depth of 16bit. One audio frame is generated every 20ms and timestamped with the system clock at the millisecond level. Secondly, a 1080P wide-angle camera (60° shooting angle) is embedded in the center of the top of the screen, capturing images at a frame rate of 30 frames per second. Each frame is timestamped in the same format after the face region is extracted using a Haar-like algorithm. Thirdly, the system clock is calibrated using NTP, and the program compares the audio and video frame timestamps. If they match, they form a synchronized frame pair; if the difference exceeds 5ms, the pair is discarded. The comparison result is used as the audio stream information and video stream information containing the user's facial features in the audio-visual environment for the smart TV. In other embodiments, other methods can also be used to obtain the audio stream information and video stream information containing the user's facial features in the audio-visual environment for the smart TV, which will not be elaborated here.

[0024] It should be noted that, in this application, the audiovisual environment refers to the physical space where the smart TV is located, including background noise, TV playback sound, and environmental elements related to user activities; audio stream information refers to the continuous digital signal of all sounds in the audiovisual environment collected by the microphone array; user facial features refer to the key facial features that can be used to determine vocal behavior; and video stream information refers to the continuous image sequence containing dynamic images of the user's face captured by the camera.

[0025] In step S2, the audio stream information and the video stream information are subjected to collaborative playback verification to obtain the audio energy characteristics and user visual motion characteristics of the time already played in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it is determined whether there is a valid user voice command input period, and a voice activity identifier corresponding to the user voice command input period is generated. In this embodiment, the collaborative playback verification of the audio stream information and the video stream information to obtain the audio energy characteristics and user visual motion characteristics at the playback moment in the audiovisual environment can be achieved through the following steps: Align the audio stream information with the video stream information to obtain a synchronized audio frame sequence and video frame sequence; The audio energy characteristics of the played moments are determined based on the audio frame sequence. The user's visual motion features are obtained by detecting the user's screen area and extracting motion vectors from the video frame sequence.

[0026] In practice, firstly, the smart TV system clock can be calibrated using a network time protocol, synchronizing the standard time every 10 seconds to control the system clock error within ±1 millisecond, ensuring consistency between the audio and video acquisition time bases. Then, the acquired audio stream information is split into 20 milliseconds / frames, generating audio frames and assigning a timestamp based on the calibrated clock to each audio frame, forming an initial audio frame sequence. Simultaneously, the video stream information is split into 30 frames / seconds, assigning a timestamp of the same format to each video frame, forming an initial video frame sequence. Subsequently, the program iterates through the two initial frame sequences, comparing the timestamps of each audio and video frame. If the timestamp difference is ≤3 milliseconds, they are directly marked as a "synchronized frame pair." If the difference is between 3 and 10 milliseconds, linear interpolation is used to compensate the video frame timestamp before marking it as a synchronized frame pair. If the difference is >10 milliseconds, it is determined to be an invalid frame and discarded. Finally, all synchronized frame pairs are arranged in timestamp order to obtain the synchronized audio and video frame sequences. Then, a Short-Time Fourier Transform (STFT) is performed on each audio frame in the synchronized audio frame sequence: a Hanning window is selected as the window function (to reduce spectral leakage, theoretically based on the fact that the sidelobe attenuation rate of the Hanning window is higher than that of the rectangular window, which can reduce interference from adjacent frequency components), the window length is set to 25 milliseconds (covering the duration of basic human speech syllables, which conforms to the time-domain characteristics of speech signals), and the window overlap rate is 50% (to avoid signal breakage between frames); each audio frame is converted from a time-domain signal to a frequency-domain signal through STFT to obtain the spectral amplitude matrix of that frame; the audio energy is calculated: the amplitude value of each frequency point in the spectral amplitude matrix is ​​squared, and then the squared values ​​of all frequency points are summed to obtain the total energy of the audio frame; considering that the energy of a single frame may be affected by instantaneous noise, the average of the total energy of three consecutive adjacent audio frames is taken as the audio energy value of the "played moment" corresponding to the intermediate frame; then the timestamp of each played moment is associated with the corresponding audio energy value and stored to form the audio energy feature of the played moment.Finally, user screen region detection is performed on each frame of the synchronized video frame sequence: a Haar-like feature detection algorithm is used, and a "face contour + eye" combined feature template is selected; the sliding window size is set to 20×20 pixels, the step size is 5 pixels, the entire frame image is traversed, and the matching degree between the features in the window and the template is calculated. If the matching degree is greater than the preset threshold, the window region is marked as the user screen region; then, motion vectors are extracted: for the user screen regions of two adjacent video frames, the Lucas-Kanade optical flow method is used. First, the corner points (such as lip edges and eye corners) in the user region of the previous frame are extracted using the Shi-Tomasi corner detection algorithm, and then the position coordinates of the corresponding corner points in the next frame are calculated. The motion vector (including horizontal and vertical components) of each corner point is obtained through the coordinate difference; finally, visual motion features are generated: for all corner point motion vectors corresponding to each video frame, the average value of the vector magnitude and the consistency of the vector direction are calculated. These two parameters are associated with the playback time of the corresponding video frame to form the user's visual motion features.

[0027] It should be noted that, in this application, collaborative playback verification refers to the correlation analysis of synchronized audio stream information and video stream information, combining audio energy features extracted from the audio stream and user visual motion features extracted from the video stream to perform dual verification at each moment during the smart TV audio and video playback process, in order to determine whether there is a valid user voice command input processing process at that moment; audio frame sequence refers to an ordered set of audio data where each audio frame has a unique time identifier; video frame sequence refers to an ordered set of video data where each video frame has a unique time identifier; audio energy feature refers to the parameter of the strength of the audio signal at the time of playback; user screen area refers to the image area containing key parts of the user's body; motion vector refers to the parameter of the change in pixel position within the user screen area between adjacent video frames; visual motion feature refers to the feature that determines whether the user has an active movement trend matching the voice command.

[0028] Preferably, in this embodiment, based on the audio energy features and the visual motion features, it is determined whether there is a valid user voice command input period, and a voice activity identifier corresponding to the user voice command input period is generated, with reference to... Figure 2 As shown in the figure, this is a flowchart illustrating the process of determining a voice activity identifier in some embodiments of this application. In this embodiment, determining the voice activity identifier can be achieved using the following steps: In step S21, playback detection is performed on the audio energy features, and activity detection is performed on the visual motion features to generate preliminary audio activity segments and user visual attention concentration segments, respectively. In step S22, the preliminary audio activity segment and the user visual attention concentration segment are aligned to filter out candidate user voice command input time periods; In step S23, the rules for merging adjacent time periods and the silence interval rules within each candidate user's voice command input time period are determined; In step S24, the voice activity identifier corresponding to the user's voice command input time period is determined according to all adjacent time period merging rules and all silence interval rules.

[0029] In specific implementation, firstly, audio energy feature playback detection: 100 sets of home scene sample data (including normal user speech, TV playback sound, and environmental noise) are collected. The minimum audio energy value of the user's speech in each sample is calculated, and the average of all minimum values ​​is taken as the audio energy threshold. The synchronized audio frame sequence is traversed. If the energy value of 3 consecutive audio frames is greater than the threshold, the timestamp of the first frame is recorded as the start time. This continues until 2 consecutive frames have an energy value less than the threshold, at which point the time is recorded as the end time. This time interval is a preliminary audio activity segment. This process is repeated to generate all preliminary audio activity segments. Visual motion feature activity detection: The synchronized video frame sequence is traversed. If the "average magnitude of lip movement vector" in the visual motion features of a frame is greater than a preset value (determined through 50 sets of user vocalization tests, which can distinguish between lip micro-movement and no movement) and the "head orientation angle is between -15° and 15°", the frame is determined to be a valid frame. The start and end times of the time interval consisting of 5 consecutive valid frames are recorded, which is a user visual attention concentration segment. Next, a timestamp mapping table is established between the initial audio activity segments and the user's visual attention concentration segments, including the start time (T1) and end time (T2) of each segment. All initial audio activity segments are traversed, and for each audio segment, all visual segments are matched one by one to calculate the time overlap between the two: if audio segment T1 ≤ visual segment T2 and audio segment T2 ≥ visual segment T1, it is determined that there is an overlap, and the larger of T1 and the smaller of T2 in the overlapping part are taken as the overlapping interval. The duration of the overlapping interval is counted. If the duration is ≥ the preset minimum duration, then the overlapping interval is the candidate user voice command input period. This process is repeated to filter out all user voice command input periods. Next, the adjacent time period merging rule is determined: 50 sets of continuous user voice command samples (such as "fast forward, then turn up the volume") are collected, and the maximum interval between adjacent commands in each set of samples is counted. The average of all maximum values ​​is taken as the merging interval threshold. If the time interval between two adjacent candidate time periods is ≤ the merging interval threshold, and the audio energy characteristics of the two time periods are consistent (the difference between the energy of the starting frame of the later time period and the energy of the ending frame of the previous time period is ≤ 10%), then the two time periods are merged. The T1 of the merged time period is the previous time period T1, and the T2 is the later time period T2, thus obtaining the adjacent time period merging rule. The silence interval rule is determined: The maximum silence duration within the commands in the above 50 sets of samples is counted, and the average is taken as the silence interval threshold. Each candidate time period is traversed. If there is an interval (silent interval) within the time period where the continuous audio frame energy value is ≤ the audio energy threshold, and the duration of this interval is > the silence interval threshold, then the candidate time period is split into two time periods. The splitting points are the start and end times of the silence interval, thus obtaining the silence interval rule.Finally, for all candidate user voice command input time periods, the process is traversed and merged according to the adjacent time period merging rule: for the sorted candidate time period list, starting from the first time period, it is sequentially judged whether it meets the merging condition with the next time period. If it does, it is merged; otherwise, the original time period is retained, until all time periods are traversed to obtain the merged time period set. Then, for each merged time period, it is silenced according to the silence interval rule, and the time periods containing excessively long silences are split to obtain the final effective user voice command input time periods. For each final time period, its voice activity identification information is calculated: time period start time, time period end time, audio confidence (the proportion of frames with audio frame energy > threshold within the time period), and visual confidence (the proportion of effective visual frames within the time period). The comprehensive confidence is (audio confidence × 0.6 + visual confidence × 0.4) (the weights are determined based on the contribution of audio and visual to the command judgment in 100 sets of tests). This information is integrated to generate the voice activity identifier corresponding to the user's voice command input time period.

[0030] It should be noted that, in this application, the user voice command input period refers to a continuous time interval that simultaneously satisfies both "the presence of user voice energy" and "the presence of user voice-related actions"; the preliminary audio activity period refers to a continuous time interval containing user voice after the audio energy feature playback detection; the user visual attention concentration period refers to a continuous time interval in which the user performs voice-related actions and their attention is directed towards the television; the adjacent time interval merging rule refers to the standard for determining whether two adjacent candidate user voice command input periods need to be merged into a single continuous time interval; the silence interval rule refers to the standard for determining whether the silence interval within a candidate user voice command input period needs to be retained; and the voice activity identifier refers to the key information of the effective user voice command input period after merging and de-silencing processing.

[0031] In step S3, during the period when the voice command input is determined to be valid, the audio stream information is separated into different stages to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV. The playback reconstruction command is then converted into a playback update request corresponding to the audio and video playback information of the smart TV. The synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video are then determined through the playback update request.

[0032] In this embodiment, the following steps can be used to separate the audio stream information during the period determined to be a valid voice command input to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV: Extract the user voice command audio segment from the audio stream information within the valid voice command input period; The audio segment of the user's voice command is sent to a speech recognition service cluster deployed in the cloud, and converted into a corresponding text command sequence; The text instruction sequence is matched with a preset playback control instruction keyword library to generate playback reconstruction instructions for the target user to control audio and video playback on the smart TV.

[0033] In practice, firstly, all audio frames within the valid voice command input period are extracted from the synchronized audio frame sequence to form a dedicated audio dataset for that period. The dataset is then processed using an echo cancellation (AEC) algorithm: the original digital signal of the TV's self-broadcast audio within that period is obtained from the smart TV playback system as a reference signal. The propagation delay of the reference signal from the speaker to the microphone is estimated using the AEC algorithm (based on 100 sets of home scenario tests, the delay range is typically 5-20 milliseconds). The timestamp of the reference signal is then adjusted according to the delay to synchronize the reference signal with the TV's self-broadcast audio interference signal in the dataset. Subsequently, the dataset and the adjusted reference signal are compared using a difference operation to remove self-broadcast audio interference. Next, a voice endpoint detection (VAD) algorithm is used: the interference-free audio is split into 20-millisecond frames, and the short-time energy (sum of squares of all sample points within the frame) and zero-crossing rate (number of symbol changes at sample points within the frame) of each frame are calculated. These are compared to a preset threshold (determined based on 50 sets of clean user voice samples to ensure coverage of over 95% of valid voice frames), and consecutive valid voice frames are selected and spliced ​​together to form the user voice command audio segment. Then, the audio segments of user voice commands are preprocessed: the audio segments are uniformly converted to 16kHz sampling rate, 16-bit mono WAV format, and the difference in audio amplitude is eliminated by a mean normalization algorithm (calculate the mean of all audio sampling points, subtract the mean from each sampling point to ensure uniform amplitude range); then, an encrypted connection is established with the cloud-based voice recognition service cluster through the network module (Wi-Fi or Ethernet) of the smart TV, and the preprocessed audio segments are uploaded to the cluster; the cluster adopts a distributed architecture, containing more than 10 recognition nodes, and each node deploys a CNN-LSTM hybrid recognition model: first, the spectral features of the audio are extracted through the CNN layer (the audio is converted into a Mel spectrogram, and local texture features are extracted through 3 layers of convolutional kernels), then the temporal relationship of the features is captured through the LSTM layer (2 layers of LSTM units, memory stride set to 5, adapted to the syllable length of the voice command), and finally, the probability distribution of the text is output through a fully connected layer and a Softmax activation function, and the text with the highest probability is selected to form the text; all nodes recognize the same audio segment in parallel, and the recognition result consistent with the majority of nodes is taken as the final text command sequence.Finally, a pre-defined keyword library for playback control commands was constructed: categorized and stored by "control type," including progress control (keywords: fast forward, rewind, skip, corresponding parameters: duration, target time), status control (keywords: pause, play, stop, no parameters), and volume control (keywords: increase, decrease, mute, corresponding parameters: number of bars, on / off). Each keyword is associated with 3-5 synonyms (e.g., "fast forward" is associated with "fast forward" and "accelerate playback"). The library's data is generated based on statistical analysis of 1000 sets of commonly used user control command samples; subsequently... The text command sequence is processed by word segmentation (using the Jieba word segmentation algorithm to split words according to Chinese semantics, such as "fast forward five minutes" being split into "fast forward" and "five minutes"); the word segmentation results are traversed and compared with keywords and synonyms in the keyword library. If a control type keyword (such as "fast forward") is matched, the parameter (such as "five minutes") is further extracted. If no parameter is matched, the default value (such as "fast forward" defaults to 5 seconds) is used to supplement it; finally, the combination of "control type + parameter" is used as the playback reconstruction command when the target user controls the audio and video playback of the smart TV.

[0034] It should be noted that in this application, progress separation refers to the process of extracting a pure audio segment containing only the user's voice command within the determined valid voice command input period; audio and video playback control refers to the various adjustment operations performed by the smart TV on the currently playing audio and video content based on the playback reconstruction command generated by the user; user voice command audio segment refers to a pure user voice signal segment after removing the TV's self-playing audio and environmental noise; voice recognition service cluster refers to providing distributed speech-to-text processing capabilities to meet the user voice command recognition needs under different accents and speaking speeds; text command sequence refers to the text sequence that the smart TV can recognize after receiving playback commands; the preset playback control command keyword library refers to all executable playback control commands of the smart TV and their corresponding keywords; playback reconstruction command refers to the structured commands in the playback control operation process of the smart TV.

[0035] In this embodiment, converting the playback reconstruction instruction into a playback update request corresponding to the smart TV audio and video playback information can be achieved through the following steps: The playback reconstruction instruction and the current audio and video playback status information are packaged into a playback instruction queue and pushed to the playback control service instance of the smart TV, waiting for the service instance to receive and confirm. The playback control service instance parses the received playback reconstruction instruction into playback control operation parameters according to the predefined instruction-action mapping rules; The playback control operation parameters are submitted to the main playback engine of the smart TV for verification, and a playback update request corresponding to the audio and video playback information of the smart TV is generated.

[0036] In practice, the system first obtains the current audio and video playback status information through the status query interface of the smart TV's main playback engine. This information includes the current playback progress (formatted as "hour:minute:second"), playback status (play / pause / stop), current volume (0-100), and total audio and video duration. The obtained status information and playback reconstruction instructions are then encapsulated into queue elements in the format of "instruction content + status data + timestamp," with the timestamp being the system time at the time of encapsulation (accurate to milliseconds). These elements are added to the playback instruction queue according to the "first-come, first-served" principle. If the instruction is urgent (e.g., pause, stop), the queue element is prioritized "high" and inserted at the head of the queue (non-urgent instructions are prioritized "low" and appended to the tail of the queue). The playback instruction queue is then pushed to the playback control service instance via the inter-process communication (IPC) protocol. After pushing, a 3-second timeout waiting mechanism is initiated: if a "successful reception" confirmation message (containing the unique identifier of the queue element) is received from the service instance within 3 seconds, the push is considered successful; if no "successful reception" message is received within the timeout or a "failed reception" message is received, the message is re-pushed, with a maximum of 2 retries. Then, a predefined instruction-action mapping rule library is constructed, defining mapping relationships according to control type: progress control (e.g., "fast forward" maps to "progress increase, parameter format: seconds") and "jump". The mapping is as follows: Mapping is used for "target progress, parameter format: hour:minute:second"); status control (e.g., "pause" is mapped to "status switch, parameter: pause identifier"); and volume control (e.g., "increase" is mapped to "volume increase, parameter format: integer number of bars"). Each mapping rule includes a parameter value range (e.g., fast forward duration ≤ 300 seconds, volume bars 1-10). After receiving the playback instruction queue, the playback control service instance extracts the playback reconstruction instructions from the queue elements, first identifying the control type; then parsing the parameters according to the mapping rules: if the instruction contains explicit parameters, it is converted into a quantized value according to the rules (five minutes = 300 seconds); if the instruction does not have explicit parameters (e.g., "fast forward"), it is filled with default parameters according to the rules (e.g., default 3 seconds); after parsing, it verifies whether the parameters conform to the value range. If they conform, it generates playback control operation parameters (e.g., "operation type: progress increase, parameter: 300 seconds"); if they do not conform (e.g., "fast forward ten minutes", exceeding the 300-second limit), it is corrected to the maximum value (300 seconds) and the correction log is recorded, thus obtaining the playback control operation parameters.Finally, the playback control service instance encapsulates the playback control operation parameters into a verification request in the format of "operation type + target parameter + current state snapshot" and submits it to the parameter verification interface of the main playback engine. After receiving the request, the main playback engine initiates dual verification: the first is parameter validity verification (e.g., whether the target progress is ≤ the total audio and video duration, and whether the target volume is within the range of 0-100), and the second is state compatibility verification (e.g., the "progress increase" operation needs to verify whether the current state is "play" or "pause," and if it is "stopped," it is considered incompatible). If both verifications pass, a playback update request is generated based on the playback control operation parameters: progress control requests include "operation type + target progress + execution timing" (execution timing is "immediate"), status control requests include "operation type + target status," and volume control requests include "operation type + target volume." If the verification fails (e.g., the target progress exceeds the total duration), a "verification failed" message is returned, including the reason for the failure (e.g., "target progress exceeds the total audio and video duration"). After receiving the message, the playback control service instance prompts the user through the TV screen, thus obtaining the playback update request corresponding to the smart TV's audio and video playback information.

[0037] It should be noted that, in this application, "smart TV audio and video playback information" refers to the data set that the smart TV generates and stores in real time during the playback of audio and video content, reflecting the current playback status; "current audio and video playback status information" refers to the real-time playback data of the smart TV's current audio and video; "playback instruction queue" refers to the playback reconstruction instructions and corresponding playback status information stored in the order of reception; "playback control service instance" refers to the intermediate scheduling unit that receives the playback instruction queue, verifies the validity of the instructions, and triggers the parsing process; "predefined instruction-action mapping rule" refers to the mapping relationship between the control type of the playback reconstruction instruction and the corresponding playback engine operation parameters; "playback control operation parameters" refers to the quantified parameters of the control intent that the smart TV can recognize; "main playback engine" refers to the execution unit of playback control during smart TV audio and video playback; and "playback update request" refers to a request message that conforms to the execution format of the main playback engine.

[0038] Preferably, in this embodiment, the synchronous playback conditions for real-time control of the smart TV in the current playback progress of audio and video are determined by the playback update request, with reference to... Figure 3 As shown in the figure, this is a flowchart illustrating the process of determining synchronized playback conditions in some embodiments of this application. In this embodiment, determining the synchronized playback conditions can be achieved using the following steps: In step S31, the control action type and target playback parameters of the smart TV in the current playback progress of the audio and video are determined according to the playback update request, and a preliminary playback status synchronization instruction is generated by combining the current audio and video decoding status and buffer queue information. In step S32, the playback status synchronization command is distributed to the audio and video playback pipeline of the smart TV, and the system waits for each processing unit to return a ready confirmation. In step S33, each processing unit adjusts its local state in parallel according to the playback state synchronization instruction, and generates a local synchronization ready signal based on the adjusted state. In step S34, the playback control service verifies the local synchronization ready signals of each processing unit and generates synchronized playback conditions for the smart TV to perform real-time control in the current playback progress of audio and video.

[0039] In practice, the process begins by extracting core information from the playback update request: determining the control action type through keyword matching and identifying the target playback parameters through a parameter extraction algorithm. Then, the current decoding status is obtained through the decoder interface of the smart TV's main playback engine, including the timestamp of the current decoded frame and the number of decoded but unplayed frames. Buffer queue information, including the number of audio / video buffered frames and the timestamp range of the buffered frames, is obtained through the playback pipeline's cache module interface. Subsequently, instruction generation rules are constructed: for progress adjustments, the instruction must include "target progress timestamp + decoding frame positioning requirements"; for volume control, the instruction must include "target volume value + audio amplifier power adjustment requirements"; and for state switching, the instruction must include "target state + decoder start / stop requirements." Finally, the control action type, target playback parameters, decoding status, and buffer queue information are integrated to generate preliminary playback status synchronization instructions. Next, the list of processing units for the audio and video playback pipeline is defined, including decoders, audio amplifiers, video renderers, and cache modules. The "inter-module communication protocol" within the smart TV (optimized based on TCP protocol to reduce latency and adapt to the short-distance data transmission requirements within the playback pipeline) is adopted. The initial playback status synchronization command is encapsulated into a distribution data packet in the format of "unit ID + command content," and sent individually to each processing unit. After distribution, a ready confirmation waiting mechanism is initiated: the waiting timeout is set to 200 milliseconds (based on the results of 100 pipeline response tests). If a ready confirmation signal is received from a processing unit before the timeout, that unit is marked as "confirmed." If no ready confirmation signal is received before the timeout or a "reception failure" signal is received, the command is re-distributed, with a maximum of two retries. Then, the decoder: if the instruction is a progress adjustment type, adjusts the decoding frame positioning using the frame index positioning algorithm, and records the timestamp of the current positioning frame after completion; the audio amplifier: if the instruction is a volume control type, adjusts the output power using the power adjustment algorithm, and records the current power value after completion; the video renderer: if the instruction is a state switching type, starts the rendering cache cleanup, and records the rendering ready flag after completion; the cache module: if the instruction is marked "prioritize cache replenishment", starts the cache to accelerate download, and records the number of cached frames after completion; after each unit is adjusted, a local synchronization ready signal is generated, including "unit ID, adjustment result (success / failure), and ready parameters (such as frame timestamp, power value)", and is sent to the playback control service through the internal communication protocol.Finally, check if signals have been received from all core processing units in the audio / video playback pipeline. If any unit has not received a signal, the verification is deemed a failure, the missing unit ID is recorded, and "readjustment required" is reported. If the integrity verification passes, the ready parameters in each signal are extracted for cross-validation—for example, in the progress adjustment scenario, verify whether the timestamp of the decoder's positioning frame matches the timestamp of the video renderer's cached frame, and verify whether the number of cached frames in the cache module is greater than or equal to the preset safety value; in the volume control scenario, verify whether the actual power value of the audio amplifier matches the target volume value. If all verifications pass, integrate the verification results to generate synchronous playback conditions, including "readiness status of all processing units (all successful), ready parameters of each unit (such as frame timestamp, power value), execution timing (immediate execution, no delay)". If the verification fails, generate "verification failure reason" and trigger the readjustment process.

[0040] It should be noted that, in this application, "current playback progress" refers to the real-time position of the smart TV during audio and video content playback, which corresponds to the total duration of the audio and video content; "control action type" refers to the specific operation category corresponding to the playback update request; "target playback parameters" refers to the quantitative parameters of the execution target corresponding to the control action type; "buffer queue information" refers to the set of key data of audio and video frames stored in real time within the cache module of the smart TV's audio and video playback pipeline; "playback status synchronization instruction" refers to the initial instruction that integrates control requirements and current playback basic information; "audio and video playback pipeline" refers to the set of processing modules that carry out the entire process of audio and video playback on the smart TV; "local synchronization ready signal" refers to the status signal fed back to the playback control service by a single processing unit in the audio and video playback pipeline after completing its own local state adjustment; and "synchronization playback conditions" refers to the global prerequisites that the smart TV must meet to execute real-time playback control.

[0041] In step S4, the playback control requirements of the smart TV during audio and video playback are determined based on the voice activity identifier and the synchronous playback conditions, and audio and video playback is performed according to the playback control requirements.

[0042] In this embodiment, determining the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronization playback conditions can be achieved through the following steps: Extract user commands and intent classifications from the voice activity identifiers, and extract playback progress status and set of executable operations from the synchronous playback conditions; The intent classification and the playback progress status are submitted to a matching engine based on playback strategy rules to obtain the control adaptability of the user command in the current playback context. The playback control requirements of the smart TV during audio and video playback are determined based on the control adaptability and the set of executable operations.

[0043] In practice, firstly, voice activity identifiers are read from the instruction identifier database, user instructions are extracted through keyword matching, and then the intent classification is determined according to preset classification rules (based on 1000 sets of user instruction scenario samples), thus obtaining the user instruction and intent classification; the playback progress status is extracted from the synchronous playback conditions, including the current playback time, total duration, and progress percentage (current time / total duration), and the data source is the real-time status interface of the autonomous playback engine; when extracting the set of executable operations, the ready parameters of each processing unit in the synchronous playback conditions are traversed, and if the unit's ready parameters meet the operation requirements, the corresponding operation is added to the set, thus obtaining the playback progress status and the set of executable operations. Then, a playback strategy rule base is constructed: rules can be formulated based on 500 typical playback scenarios (such as "fast forward when the progress is low" and "rewind when the progress is high"), and each rule contains an adaptation factor (such as in the progress control category, if the target progress has not exceeded the total time, the intent matching factor is high); after receiving the intent category and playback progress status, the matching engine first matches the corresponding rule set, and then calculates the control adaptation: the control adaptation can be determined using a weighted summation model, with the weights determined based on 1000 sets of user feedback data, and the intent and progress matching degree accounting for 0.6 and hardware support accounting for 0.4 in the determination process. Finally, a preset adaptation threshold (determined based on 200 sets of control scenario tests) is set. If the control adaptation is greater than or equal to the adaptation threshold, operations matching the intent category are selected from the set of executable operations, and specific control requirements are generated by combining user command parameters. If the adaptation is less than the adaptation threshold, the reasons are analyzed, and similar operations are recommended from the set of executable operations to generate alternative requirements. The final main requirements and alternative requirements are bound to the execution constraints, and the binding result is used as the playback control requirements of the smart TV when playing audio and video.

[0044] It should be noted that, in this application, user instructions refer to the specific requests made by the user to the smart TV via voice to control audio and video playback; intent classification refers to categorizing user instructions according to control objectives to provide a clear direction for the matching engine; playback progress status refers to the relationship between the current playback time position of the audio and video and the overall progress; the set of executable operations refers to the operations supported by all processing units under synchronized playback conditions; the matching engine based on playback strategy rules refers to the calculation module of the playback control scene rules built into the smart TV; control adaptability refers to the degree of matching between the user instruction intent and the current playback progress status and hardware capabilities; playback control requirements refer to the specific control operations and execution constraints that the smart TV needs to perform.

[0045] In specific implementation, audio and video playback based on the playback control requirements can be achieved in the following way: First, the playback control service of the smart TV reads the determined playback control requirements and extracts the specific control operation type, execution parameters, and execution constraints. Then, the playback control service breaks down the requirements into dedicated tasks for each processing unit of the audio and video playback pipeline. If it is a progress adjustment requirement, an instruction to "locate to the frame corresponding to the target progress" is sent to the audio and video decoder, and an instruction to "synchronously load and output the target frame" is sent to the video renderer, and an instruction to "maintain the current volume and synchronously output the target audio frame" is sent to the audio amplifier. If it is a state switching requirement, an instruction to "stop decoding the current frame" is sent to the decoder. The system sends a "freeze current screen" command to the renderer and a "pause audio output" command to the audio amplifier. For volume control requests, it only sends a "adjust output power to the target volume value" command to the audio amplifier. During the execution of tasks by each processing unit, the playback control service receives real-time feedback on the execution status from the units. If all units report "execution completed," it determines that the audio and video playback has been adjusted according to the control requirements, and simultaneously displays the execution result through a pop-up window on the TV screen. If a unit reports "execution failed," the playback control service immediately triggers a retry mechanism. If the retry still fails, it prompts the user on the screen with "Current operation failed, please check the command parameters," ensuring that the playback effect matches the user's requirements.

[0046] Therefore, this application demonstrates an improvement in the adaptability of voice-controlled smart TV audio and video playback. Specifically, by simultaneously acquiring audio stream information and video stream information containing user facial features from the audiovisual environment, it can eliminate spatiotemporal misalignment of audio and video based on timestamp alignment, overcoming the limitations of single-modal data acquisition. Through collaborative playback verification of audio and video stream information and extraction of audio energy features and visual motion features, invalid scenarios with only background noise or no voice action are eliminated, accurately locating the true effective command period and generating voice activity identifiers. Within the effective period, progress separation of the audio stream information yields playback reconstruction instructions, which are then converted into playback update requests and synchronous playback conditions are determined. This process removes TV-broadcast audio and environmental noise to obtain pure user voice, accurately parsing it into structured instructions. These instructions, combined with the current playback progress, decoding status, and buffer queue information, generate synchronous playback conditions adapted to the hardware capabilities, preventing command parameters from exceeding the hardware processing range. Based on voice activity identifiers and synchronous playback conditions, playback control requirements are determined and playback is executed. The system can combine the validity of the command with the hardware readiness status to filter highly adaptable control requirements, ensuring that the executed operation matches both the user's true intent and the hardware's real-time capabilities, avoiding invalid control or conflicting operations, and ultimately improving the accuracy of voice command recognition.

[0047] In summary, the technical solution adopted in this application can dynamically adapt to the voice command recognition and playback control of smart TVs in complex audiovisual environments, thereby improving the accuracy of voice command playback control.

[0048] Example 2: This application provides a smart TV audio and video playback control system based on voice recognition, referencing... Figure 4 As shown in the figure, this is a modular structure diagram of a speech recognition-based intelligent TV audio and video playback control system according to this embodiment of the present application. The control system includes: The information acquisition module 100 is used to simultaneously acquire audio stream information and video stream information containing user facial features of the smart TV with voice recognition in the audio-visual environment. The playback verification module 200 is used to perform collaborative playback verification on the audio stream information and the video stream information to obtain the audio energy characteristics and the user's visual motion characteristics at the time of playback in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it determines whether there is a valid user voice command input period and generates a voice activity identifier corresponding to the user voice command input period. The progress separation module 300 is used to separate the progress of the audio stream information during the period when it is determined to be a valid voice command input period, to obtain the playback reconstruction instruction when the target user controls the audio and video playback of the smart TV, to convert the playback reconstruction instruction into a playback update request corresponding to the audio and video playback information of the smart TV, and then to determine the synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video through the playback update request. The playback control module 400 is used to determine the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronous playback conditions, and to perform audio and video playback according to the playback control requirements.

[0049] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0050] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0051] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A method for controlling audio and video playback on a smart TV based on speech recognition, characterized in that, The control method includes the following steps: Simultaneously acquire audio stream information and video stream information containing user facial features from smart TVs with voice recognition in the audiovisual environment; The audio stream information and the video stream information are collaboratively played and verified to obtain the audio energy characteristics and the user's visual motion characteristics at the time of playback in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it is determined whether there is a valid user voice command input period and a voice activity identifier corresponding to the user voice command input period is generated. Specifically, determining whether a valid user voice command input period exists based on the audio energy features and the visual motion features, and generating a voice activity identifier corresponding to that user voice command input period, includes: Playback detection is performed on the audio energy features, and activity detection is performed on the visual motion features to generate preliminary audio activity segments and user visual attention concentration segments, respectively. Align the initial audio activity segment and the user visual attention concentration segment to filter out candidate user voice command input time periods; Determine the rules for merging adjacent time periods and the silence interval rules within the voice command input time period of each candidate user; The voice activity identifier corresponding to the user's voice command input time period is determined based on all adjacent time period merging rules and all silent interval rules. The aforementioned voice activity identifier refers to key information about the effective user voice command input period after merging and de-silencing processing; During the period when a valid voice command is input, the audio stream information is separated into different stages to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV. The playback reconstruction command is then converted into a playback update request corresponding to the audio and video playback information of the smart TV. The synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video are then determined through the playback update request. Specifically, the synchronous playback conditions for determining the real-time playback progress of the smart TV in the current audio and video playback progress through the playback update request include: Based on the playback update request, determine the control action type and target playback parameters of the smart TV in the current playback progress of the audio and video, and generate a preliminary playback status synchronization instruction by combining the current audio and video decoding status and buffer queue information. The playback status synchronization command is distributed to the audio and video playback pipeline of the smart TV, and the system waits for each processing unit to return a readiness confirmation. Each processing unit adjusts its local state in parallel according to the playback state synchronization command, and generates a local synchronization ready signal based on the adjusted state. The playback control service verifies the local synchronization ready signals of each processing unit and generates synchronous playback conditions for the smart TV to perform real-time control in the current playback progress of audio and video. The synchronous playback condition mentioned above refers to the global prerequisites that a smart TV must meet to perform real-time playback control. The playback control requirements of the smart TV during audio and video playback are determined based on the voice activity identifier and the synchronous playback conditions, and the audio and video playback is performed according to the playback control requirements.

2. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, The collaborative playback verification of the audio stream information and the video stream information to obtain the audio energy characteristics and user visual motion characteristics at the time of playback in the audiovisual environment specifically includes: Align the audio stream information with the video stream information to obtain a synchronized audio frame sequence and video frame sequence; The audio energy characteristics of the played moments are determined based on the audio frame sequence. The user's visual motion features are obtained by detecting the user's screen area and extracting motion vectors from the video frame sequence.

3. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, The aforementioned visual motion features refer to the features used to determine whether a user exhibits an active movement trend that matches the voice commands.

4. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, During the period determined to be a valid voice command input, the audio stream information is subjected to progress separation to obtain the playback reconstruction command when the target user controls the audio and video playback of the smart TV. Specifically, this includes: Extract the user voice command audio segment from the audio stream information within the valid voice command input period; The audio segment of the user's voice command is sent to a speech recognition service cluster deployed in the cloud, and converted into a corresponding text command sequence; The text instruction sequence is matched with a preset playback control instruction keyword library to generate playback reconstruction instructions for the target user to control audio and video playback on the smart TV.

5. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, Converting the playback reconstruction instruction into a playback update request corresponding to smart TV audio and video playback information specifically includes: The playback reconstruction instruction and the current audio and video playback status information are packaged into a playback instruction queue and pushed to the playback control service instance of the smart TV, waiting for the service instance to receive and confirm. The playback control service instance parses the received playback reconstruction instruction into playback control operation parameters according to the predefined instruction-action mapping rules; The playback control operation parameters are submitted to the main playback engine of the smart TV for verification, and a playback update request corresponding to the audio and video playback information of the smart TV is generated.

6. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, The playback update request refers to a request message that conforms to the execution format of the main playback engine.

7. The intelligent TV audio and video playback control method based on speech recognition as described in claim 1, characterized in that, Determining the playback control requirements of a smart TV during audio and video playback based on the voice activity identifier and the synchronized playback conditions specifically includes: Extract user commands and intent classifications from the voice activity identifiers, and extract playback progress status and set of executable operations from the synchronous playback conditions; The intent classification and the playback progress status are submitted to a matching engine based on playback strategy rules to obtain the control adaptability of the user command in the current playback context; The playback control requirements of the smart TV during audio and video playback are determined based on the control adaptability and the set of executable operations.

8. A speech recognition-based intelligent television audio and video playback control system, used to execute a speech recognition-based intelligent television audio and video playback control method as described in any one of claims 1 to 7, characterized in that, The control system includes: The information acquisition module is used to simultaneously acquire audio stream information and video stream information containing user facial features from a smart TV with voice recognition in the audiovisual environment. The playback verification module is used to perform collaborative playback verification on the audio stream information and the video stream information to obtain the audio energy characteristics and the user's visual motion characteristics at the time of playback in the audiovisual environment. Based on the audio energy characteristics and the visual motion characteristics, it determines whether there is a valid user voice command input period and generates a voice activity identifier corresponding to the user voice command input period. The progress separation module is used to separate the audio stream information during the period when it is determined to be a valid voice command input period, to obtain the playback reconstruction instruction when the target user controls the audio and video playback of the smart TV, to convert the playback reconstruction instruction into a playback update request corresponding to the audio and video playback information of the smart TV, and then to determine the synchronous playback conditions for real-time control of the smart TV in the current playback progress of the audio and video through the playback update request. The playback control module is used to determine the playback control requirements of the smart TV during audio and video playback based on the voice activity identifier and the synchronous playback conditions, and to play audio and video according to the playback control requirements.

Citation Information

Patent Citations

  • Intelligent sound box voice processing method and system based on artificial intelligence

    CN120544551A

  • Audio and video player control method based on voice instruction

    CN121053987A