Digital photo frame display control method and system based on voice recognition interaction

CN122551771APending Publication Date: 2026-08-11深圳市钜弘技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-24
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

其中,传统基于语音识别交互的数码相框显示控制方法是指通过语音识别技术与数码相框进行交互,从而实现对其显示内容的控制,方法通过在数码相框内嵌入语音识别模块,用户可通过语音命令控制相框的显示切换、图片浏览、播放设置等功能,传统的数码相框依赖物理按键或触摸屏进行操作,用户需要手动进行繁琐的操作

Benefits of technology

本发明中,通过结合语音指令与图像帧状态生成同步识别数据,实现语音信号与图像播放状态的精准匹配,有效提升语音指令解析的上下文关联能力,进而在存在时序错位的指令片段中识别异常区域,并通过提取缓存深度与加载延迟特征,动态分析播放过程中的资源协调性,强化了对异常帧段的判断精度,通过对停留时长与播放节奏的偏离程度进行等级评估,精准标定异常等级并生成补偿指令,提升图像播放稳定性与响应一致性,最终在控制序列中匹配关键帧段编号与缓存策略,优化缓存调度流程,增强整体播放逻辑的自适应协同能力,确保语音控制响应的时效性与显示内容的同步性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551771A_ABST
    Figure CN122551771A_ABST
Patent Text Reader

Abstract

This invention relates to the field of voice interaction control technology, specifically to a digital photo frame display control method and system based on voice recognition interaction. The method includes the following steps: capturing voice and combining it with the display status; analyzing the timing matching degree to filter abnormal regions; extracting cache features and evaluating voice response performance; outputting control commands based on frame dwell deviation; adjusting the cache allocation strategy by comparing scheduling priorities; and outputting a digital photo frame display logic collaborative control sequence. In this invention, by combining voice commands and image frame states to generate synchronized recognition data, accurate matching of voice signals and image playback status is achieved, improving the contextual association capability of voice parsing, identifying abnormal regions with misaligned commands, dynamically analyzing the coordination of playback resources, enhancing the accuracy of abnormal frame segment judgment, evaluating the degree of deviation between dwell time and playback rhythm, matching frame segment numbers and cache strategies, enhancing the collaborative capability of playback logic, and ensuring the timeliness of voice response and display synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interaction control technology, and in particular to a digital photo frame display control method and system based on voice recognition interaction. Background Technology

[0002] The field of voice interaction control technology primarily involves interacting with users through voice recognition technology to control various operations of electronic devices. Core aspects of this field include voice input recognition, voice command parsing and execution, and the integration of device control systems with voice recognition modules. Specific application scenarios cover home automation, smart homes, in-vehicle systems, wearable devices, and many others. With the continuous advancement of voice recognition technology, voice interaction control has gradually become an important method of human-computer interaction, its main goal being to achieve device control via voice commands, providing a more natural and convenient user experience. Traditional voice recognition-based digital photo frame display control methods involve interacting with the digital photo frame through voice recognition technology to control its displayed content. This method embeds a voice recognition module within the digital photo frame, allowing users to control functions such as display switching, image browsing, and playback settings via voice commands. Traditional digital photo frames rely on physical buttons or touchscreens for operation, requiring users to perform cumbersome manual operations.

[0003] Traditional digital photo frames rely on voice recognition for interactive control. However, when capturing user voice commands, they lack a synchronous analysis mechanism for the image display status. The voice recognition results cannot effectively correlate the current image frame with the playback rhythm, leading to frequent response deviations when the system executes control commands. This is especially true when image switching is frequent or the playback mode changes dynamically, easily causing misalignment between voice commands and image display actions. Furthermore, this method fails to dynamically evaluate voice response performance and playback coordination, and cannot promptly identify buffer overflows or lagging frame segments, resulting in image playback delays, unsatisfactory operation, and other user experience issues. In addition, the lack of an adaptive adjustment strategy based on display logic means that the overall playback process lacks the ability to identify and compensate for abnormal states, limiting the practicality and stability of voice interactive control. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the existing technology and to propose a digital photo frame display control method and system based on voice recognition interaction.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a digital photo frame display control method based on voice recognition interaction, comprising the following steps: S1: Based on the preset voice command lexicon and the current display status of the digital photo frame, capture the voice signal emitted by the user, extract the start and end points of the voice in the time domain, and combine the current image frame number and playback mode identifier to generate a voice and display synchronization status recognition dataset. S2: Based on the voice and display synchronization state recognition dataset, extract the voice command keywords and the current image switching direction vector, analyze the temporal matching degree of the two at the image sequence nodes, filter the image frame intervals where there is misalignment between commands and actions and the voice duration exceeds the preset threshold, and form the display control abnormal area recognition result. S3: Based on the results of the abnormal area identification of the display control, extract the cache queue depth and image loading delay features of the corresponding image frame interval, analyze the coordination between the cache filling rate and the image switching frequency, filter the frame segments where cache overflow or loading lag occurs continuously, and obtain the voice response performance evaluation data table. S4: Based on the voice response performance evaluation data table, analyze the distribution trend of the dwell time of the corresponding image frame in the display buffer, evaluate the degree of deviation from the standard playback rhythm curve, mark the abnormality level of the frame segment according to the deviation magnitude, and output the real-time frame display abnormality adjustment compensation instruction set.

[0006] As a further aspect of the present invention, the voice and display synchronization status recognition dataset includes a voice start frame number, a command keyword matching failure marker, a current playback mode identifier, and a voice end timestamp. The display control abnormal area recognition result includes a command and action misalignment frame segment identifier, a voice continuous over-limit interval, and an image sequence node mismatch block. The voice response performance evaluation data table includes a cache depth abnormal continuous frame segment, an image loading delay abrupt change area, a cache fill rate mismatch area, and a response failure frame number. The real-time frame display abnormal adjustment and compensation command set includes a risk level label, a frame response deviation value, a cache overflow frequency index, and a playback rhythm offset level.

[0007] As a further aspect of the present invention, the step of obtaining the voice and display synchronization state recognition dataset specifically includes: S111: Based on the preset voice command dictionary and the current display status of the digital photo frame, capture the user's voice signal and perform endpoint detection, extract the start time and end time of the voice signal, and record the current image frame number and playback mode type to construct a voice and image time-series aligned sequence. S112: Based on the aforementioned speech and image time-series alignment sequence, the following formula is used: ; Calculate the time-series ratio feature value of speech and image, identify the time window when the ratio exceeds the preset threshold, and simultaneously determine whether the instruction keywords are successfully matched in the dictionary, and generate a joint judgment set of instruction validity and duration; in, The feature value representing the time-series ratio of speech to image is... Representing the The duration of a speech segment Representative and the Image switching interval for segmented speech alignment The arithmetic mean of the image switching intervals. This represents the total number of segments where speech and images are aligned. S113: For the joint determination set of the validity and duration of the instruction, extract the voice events that fail to match and exceed the duration limit, bind the corresponding image frame number and playback mode identifier, and generate a voice and display synchronization status recognition dataset.

[0008] As a further aspect of the present invention, the step of obtaining the display control abnormal area identification result specifically includes: S211: Based on the voice and display synchronization state recognition dataset, extract the expected image switching direction vector corresponding to the voice command keywords, compare it with the real-time image frame jump direction, identify the symbol consistency between the two at the image sequence nodes, and generate a command and action matching map. S212: Based on the instruction and action matching map, filter the node regions where the symbols are inconsistent and the voice duration is greater than the preset benchmark value. Combined with the continuity constraint of the image sequence, identify the misaligned concentrated frame segments belonging to the same playback cycle to obtain the display control abnormal frame segment division. S213: Invoke the display control abnormal frame segment division, perform integrated analysis on instruction and action matching degree, voice persistence dispersion, cache queue stability and image loading delay, calculate display control abnormality sensitivity, identify the sensitivity layer and mark the response block for frame segment matching, and form display control abnormality area identification result.

[0009] As a further aspect of the present invention, the steps for obtaining the voice response performance evaluation data table are as follows: S311: Based on the display control abnormal area identification results, extract the cache queue depth sequence and image loading delay timestamp of the numbered frame segment in the layer, perform time alignment on the data in the frame segment, identify the cache depth fluctuation amplitude and loading delay mutation amount, and obtain the local response abnormal feature set. S312: Based on the local response anomaly feature set, perform coupled analysis on the buffer fill rate and image switching frequency, construct a buffer and switching coordination index based on the normalized difference between the image switching frequency and the buffer fill rate, filter out frame segments whose coordination index exceeds a preset threshold, and establish a buffer and switching coordination response distribution map. S313: Call the cache and switching collaborative response distribution map, cluster the frames that continuously exceed the preset benchmark value in the collaborative index layer, mark the frame number and time coordinate corresponding to the abnormal area, and obtain the voice response performance evaluation data table.

[0010] As a further aspect of the present invention, the step of obtaining the real-time photo frame display anomaly adjustment compensation instruction set specifically includes: S411: Based on the frame segment number of the voice response performance evaluation data table, extract the sequence of the dwell time of the image frame under the specified number in the display buffer, perform time normalization processing, calculate the change rate of the dwell time per unit frame, and obtain the set of abnormal change rates of the dwell time of the frame. S412: Based on the set of abnormal change rates of frame dwell time, retrieve the standard playback rhythm curve as a reference, compare the Euclidean distance between the current dwell time sequence and the reference curve, define the deviation, identify the frame segments whose deviation exceeds the warning threshold, and obtain the set of rhythm change frame segments. S413: Based on the rhythm mutation frame segment set, bind the deviation value of each frame segment to the frame number in the digital photo frame display logic diagram, arrange them in descending order of deviation, and output the real-time photo frame display abnormal adjustment compensation instruction set.

[0011] As a further aspect of the present invention, the method further includes step S5: S5: Call the real-time photo frame display abnormal adjustment and compensation instruction set, identify the corresponding number of the risk frame segment in the digital photo frame display logic diagram, extract the image cache rescheduling unit list, compare the cache scheduling priority with the playback protection sequence, filter the frame segment number that needs to adjust the cache allocation strategy, and output the digital photo frame display logic collaborative control sequence. The digital photo frame display logic collaborative control sequence includes adjusting the target frame segment number, cache depth adjustment parameters, playback protection priority comparison items, and image preloading trigger type.

[0012] As a further aspect of the present invention, the step of obtaining the digital photo frame display logic collaborative control sequence specifically includes: S511: Call the real-time photo frame display abnormality control and compensation instruction set, extract the number of the risk frame segment in the digital photo frame display logic diagram, map the risk level of the frame segment to the boundary of the display buffer, identify the frame segment information corresponding to the playback protection level, and generate a display task risk distribution map. S512: Based on the display task risk distribution map, extract the cache rescheduling unit number and scheduling priority, match the frame segment risk level with the cache scheduling level, identify the frame segment number with insufficient scheduling coverage, and obtain the display task scheduling risk disconnect list. S513: Based on the display task scheduling risk disconnect list, according to the level number in the playback protection priority sequence, extract the key frame segment number that needs to increase the cache allocation weight, output the adjustment control parameters linked with the original cache unit in sequence, and output the digital photo frame display logic collaborative management sequence.

[0013] The voice recognition-based digital photo frame display control system is used to execute the aforementioned voice recognition-based digital photo frame display control method. The system includes: The voice synchronization monitoring module captures the user's voice signal and identifies the start and end time points based on the preset voice command lexicon and the current display status of the digital photo frame. It compares the current image frame number with the playback mode, filters the voice and display synchronization failure intervals, extracts the frame number and timestamp, and generates a voice and display synchronization status recognition dataset. Based on the voice and display synchronization status recognition dataset, the command localization module identifies the temporal matching relationship between voice command keywords and image switching direction vectors, calibrates the image sequence node numbers, matches the digital photo frame display logic diagram, extracts the range of misaligned command and action frames, and establishes the display control abnormal area recognition result. Based on the results of the abnormal area identification of the display control, the cache linkage module retrieves the continuous data of the cache queue depth sequence and image loading delay features in the frame segment, judges the coordination between the cache fill rate and the switching frequency, marks the frame segment number that meets the coordination failure threshold, and outputs the voice response performance evaluation data table. The rhythm warning module analyzes the distribution trend of the dwell time of the corresponding image frame and the degree of deviation of the standard playback rhythm curve based on the frame segment number of the voice response performance evaluation data table, extracts the rhythm deviation frame segment number, completes the level identification according to the risk classification standard, and generates a real-time frame display abnormal adjustment compensation instruction set. The cache optimization module, based on the real-time photo frame display anomaly control and compensation instruction set, finds the corresponding position number of the risky frame segment in the digital photo frame display logic diagram, retrieves the current cache rescheduling unit configuration list, compares the playback protection priority with the current scheduling level, filters the frame segment numbers that need to update the cache allocation strategy, and outputs the digital photo frame display logic collaborative control sequence.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by combining voice commands and image frame states to generate synchronized recognition data, accurate matching of voice signals and image playback states is achieved, effectively improving the contextual association capability of voice command parsing. This allows for the identification of abnormal regions in command segments with temporal misalignments. Furthermore, by extracting cache depth and loading delay features, the resource coordination during playback is dynamically analyzed, enhancing the accuracy of judging abnormal frame segments. By evaluating the degree of deviation between dwell time and playback rhythm, the abnormality level is accurately labeled and compensation commands are generated, improving image playback stability and response consistency. Finally, by matching key frame segment numbers and caching strategies in the control sequence, the cache scheduling process is optimized, enhancing the adaptive and collaborative capabilities of the overall playback logic and ensuring the timeliness of voice control response and the synchronization of displayed content. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a flowchart illustrating the acquisition of the voice and display synchronization status recognition dataset in this invention. Figure 3 This is a flowchart illustrating the process of obtaining the identification results of abnormal control areas in this invention; Figure 4 This is a flowchart illustrating the process of obtaining the voice response performance evaluation data table in this invention. Figure 5 This is a flowchart illustrating the acquisition of the real-time frame display anomaly control and compensation instruction set in this invention. Figure 6 This is a flowchart illustrating the acquisition process of the digital photo frame display logic collaborative control sequence in this invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0018] Example 1 Please see Figure 1 This invention provides a technical solution for a digital photo frame display control method based on voice recognition interaction, comprising the following steps: S1: Based on the preset voice command lexicon and the current display status of the digital photo frame, capture the voice signal emitted by the user, extract the start and end points of the voice in the time domain, and combine the current image frame number and playback mode identifier to generate a voice and display synchronization status recognition dataset. S2: Based on the voice and display synchronization status recognition dataset, extract the voice command keywords and the current image switching direction vector, analyze the temporal matching degree of the two at the image sequence nodes, filter the image frame intervals where there is misalignment between commands and actions and the voice duration exceeds the preset threshold, and form the display control abnormal area recognition result. S3: Based on the results of abnormal area identification in display control, extract the cache queue depth and image loading delay features of the corresponding image frame interval, analyze the coordination between cache filling rate and image switching frequency, filter the frame segments where cache overflow or loading lag occurs continuously, and obtain the speech response performance evaluation data table. S4: Based on the voice response performance evaluation data table, analyze the distribution trend of the dwell time of the corresponding image frame in the display buffer, evaluate the degree of deviation from the standard playback rhythm curve, mark the abnormality level of the frame segment according to the deviation magnitude, and output the real-time frame display abnormality adjustment compensation instruction set. S5: Call the real-time photo frame display anomaly control and compensation instruction set, identify the corresponding number of the risky frame segment in the digital photo frame display logic diagram, extract the image cache rescheduling unit list, compare the cache scheduling priority with the playback protection sequence, filter the frame segment numbers that need to adjust the cache allocation strategy, and output the digital photo frame display logic collaborative control sequence.

[0019] The voice and display synchronization status recognition dataset includes the voice start frame number, instruction keyword matching failure marker, current playback mode identifier, and voice end timestamp. The display control abnormal area recognition results include instruction and action misalignment frame segment identifiers, voice continuous over-limit intervals, and image sequence node mismatch blocks. The voice response performance evaluation data table includes continuous frame segments with abnormal cache depth, image loading delay abrupt change areas, cache fill rate mismatch areas, and response failure frame numbers. The real-time photo frame display abnormal adjustment and compensation instruction set includes risk level labels, frame response deviation values, cache overflow frequency indicators, and playback rhythm offset levels. The digital photo frame display logic collaborative control sequence includes adjusting target frame segment numbers, cache depth adjustment parameters, playback protection priority comparison items, and image preloading trigger types.

[0020] Please see Figure 2 The specific steps for obtaining the voice and display synchronization status recognition dataset are as follows: S111: Based on the preset voice command dictionary and the current display status of the digital photo frame, capture the user's voice signal and perform endpoint detection, extract the start time and end time of the voice signal, and record the current image frame number and playback mode type to construct a voice and image time-series aligned sequence. The system calls upon a pre-stored voice command dictionary and the current display status register data of the digital photo frame in a non-volatile storage medium. It captures analog sound wave signals from the environment via an integrated MEMS omnidirectional microphone array and converts them into a PCM digital audio stream via an analog-to-digital converter at a 16kHz sampling rate and 16-bit quantization precision. The digital signal processing module performs framing operations on the audio stream, setting the frame length to 25 milliseconds and the frame shift to 10 milliseconds. A Hamming window function is applied to each frame to suppress spectral leakage. The computing unit extracts short-time energy features and zero-crossing rate features frame by frame. The short-time energy value is compared in real-time with a high threshold of 4 times the average energy of the background noise and a low threshold of 1.5 times. If five consecutive frames are detected with energy exceeding the low threshold, including at least one frame exceeding the high threshold, the logic determines it as the start point of the speech and marks the start timestamp. If ten consecutive frames have energy below the low threshold, it is determined as the end point of the speech and marks the end timestamp. The parallel-executing display monitoring thread captures the vertical synchronization interrupt signal of the display controller, reads the image frame sequence number and playback mode status register of the current rendering buffer, and samples the image frame sequence number and mode identifier at a high frequency with a period of 5 milliseconds within the detected voice start and end time window. The data processing logic performs time-series correlation on the captured voice waveform segments, start and end timestamps, image frame sequences within the corresponding time periods, and playback mode encodings to construct a voice and image time-series aligned sequence data structure with precise time indexes, which is stored in a high-speed cache for subsequent analysis and retrieval.

[0021] S112: Based on the temporal alignment sequence of speech and image, the formula is used: ; Calculate the time-series ratio feature value of speech and image, identify the time window when the ratio exceeds the preset threshold, and simultaneously determine whether the instruction keywords are successfully matched in the dictionary, and generate a joint judgment set of instruction validity and duration; in, The feature value representing the time-series ratio of speech to image is... Representing the The duration of a speech segment Representative and the Image switching interval for segmented speech alignment The arithmetic mean of the image switching intervals. This represents the total number of segments where speech and images are aligned. Based on the speech and image temporal alignment sequence, the processor calls the constructed temporal alignment sequence above to perform feature decomposition, parameters Representing the The duration of the speech segment, which is obtained from endpoint detection. The result is calculated, and the unit is milliseconds (ms). Representative and the The image switching interval for segmented speech alignment is obtained by reading the time difference of the image frame sequence number change during the occurrence of the segmented speech. The arithmetic mean of the image switching interval is obtained by summing and averaging the data of the 50 most recent image switching intervals. This represents the total number of segments where speech and images are aligned. To verify the calculation logic and technical effectiveness of the formula, a set of actual collected experimental data was selected for illustration. Parameter acquisition process: Suppose that during a certain interaction, the following was detected: (That is, the current instruction is being analyzed); The duration of the speech segment was measured using a clock counter. (i.e., 1.2 seconds, indicating a relatively slow user speaking speed); simultaneously, the image switching interval was monitored when the command was issued. (Currently in 5-second carousel mode); Historical statistics show that the average image switching interval is... ; Formula calculation process: Substitute the above values ​​into the formula: Calculate the geometric mean term: ; Calculate the absolute value of the difference: ; Calculate the denominator: ; Final eigenvalue calculation: ; Results Analysis and Processing: Calculated speech-to-image temporal ratio feature value The processor compares this value with a preset warning threshold. Comparison, The settings are based on the statistical distribution of a large number of user behavior samples. Experimental data shows that under normal interaction rhythm... Values ​​distributed in Within the interval, when At this time, it is significant that the duration of the voice command is severely "disconnected" from the rhythm of the image, or that the user hesitates. In this example, The processor determines that the time window is a ratio exceeding the limit. Simultaneously, the processor performs keyword recognition on the speech data segment, using a lightweight convolutional neural network (CNN) as the recognition model. The input is the Mel-frequency cepstral coefficients (MFCC, extracting 13-dimensional static coefficients + 13-dimensional first-order differences + 13-dimensional second-order differences, for a total of 39-dimensional features). The model structure includes 3 convolutional layers (3×3 kernel size, ReLU activation) and 2 fully connected layers. The output layer uses the Softmax function to calculate the probability of each instruction word belonging to the preset vocabulary. If the maximum probability value is less than 0.85, the matching is determined to be a failure. The processor combines the ratio exceeding the limit determination result (True / False) with the keyword matching result (Success / Fail) to generate a joint determination set of instruction validity and duration.

[0022] S113: For the joint determination set of instruction validity and duration, extract voice events that fail to match and exceed the duration limit, bind the corresponding image frame number and playback mode identifier, and generate a voice and display synchronization status recognition dataset. The processor iterates through the joint determination set of instruction validity and duration, performs conditional filtering, and extracts event records that simultaneously satisfy both the keyword matching failure flag and the speech-to-image time-series ratio feature value exceeding the limit flag. These records represent non-instructional speech interference or hesitant, prolonged speech behavior by the user under specific display content. Based on the filtered event index value, the processor backtracks to the original speech-to-image time-series alignment sequence, reads the specific image frame number displayed on the digital photo frame screen at the time of the event, and the playback mode register value at that time. The data encapsulation engine defines structured data entries containing a unique event identifier, a speech feature vector data block, a speech duration value, an image frame number integer value, a playback mode enumeration value, and an anomaly type label string. The storage controller serializes the generated structured entries into binary format and writes them to a dedicated anomaly analysis database partition, forming a speech-to-display synchronization state recognition dataset. Taking 1000 interaction tests as an example, the filtering logic accurately extracts 53 anomaly records that meet the conditions from the original logs, providing a precise negative sample data foundation for subsequent display control anomaly area identification, ensuring that the analysis process focuses on the core interference sources that cause a decline in the interactive experience.

[0023] Please see Figure 3 The specific steps for obtaining the identification results of abnormal control areas are as follows: S211: Based on the voice and display synchronization state recognition dataset, extract the expected image switching direction vector corresponding to the voice command keywords, compare it with the real-time image frame jump direction, identify the symbol consistency between the two at the image sequence nodes, and generate a command and action matching map. The semantic parsing engine reads voice command data from the voice and display synchronization status recognition dataset. Based on a preset mapping protocol, it converts natural language commands into expected image switching direction vectors in mathematical form. The next command is mapped to a positive vector, the previous command to a negative vector, and the pause command to a zero vector. Unparseable noise is marked as a non-numerical constant. The motion vector calculation module synchronously reads the real-time image frame jump log and calculates the actual direction vector by comparing the current frame number with the previous frame number. The incrementing sequence number is recorded as 1, the decrement as -1, and the unchanged as 0. The consistency check logic performs a sign function product operation on the expected direction vector and the actual direction vector at each time node. A product of 1 indicates that the directions are consistent, a product of -1 indicates that the directions conflict, and a product of 0 and a non-zero expected vector indicates no response or false response. The graph construction unit uses the time axis as the horizontal axis and the matching state enumeration value as the node attribute. It instantiates the instruction and action matching graph data structure in memory and fills the graph nodes with thousands of millisecond-level time nodes and their corresponding matching state values. The data fragments shown in Table 1 are the key node state records extracted from the graph, which accurately record the logical consistency between the instruction intent and the hardware execution action at each moment. As shown in Table 1, some key node matching data are displayed. Table 1: Data Fragments of Nodes in the Command and Action Matching Graph As shown in Table 1, the times T=3200 and T=4800 are marked as anomalous nodes, and these nodes will constitute the key analysis objects in the graph.

[0024] S212: Based on the instruction and action matching map, filter the node regions where the symbols are inconsistent and the speech duration is greater than the preset benchmark value. Combined with the continuity constraint of the image sequence, identify the misaligned concentrated frame segments belonging to the same playback cycle to obtain the display control abnormal frame segment division. The system traverses the instruction and action matching graph, identifies abnormal nodes whose symbol matching results are not equal to 1, reads the speech duration data associated with the abnormal nodes, removes short-term noise interference nodes with a duration of less than 500ms, and retains long-term abnormal nodes, the sequence analysis unit checks the continuity of the image frame sequence, counts the number of abnormal nodes within a playback period window containing 10 consecutive image frames, and when the counter counts and finds that there are more than 3 discrete abnormal nodes or more than 2 consecutive abnormal nodes within the window, the logic judgment unit determines that there is a control misalignment in the playback period, the region delineation unit locks the frame number interval that meets the conditions, marks the frame number intervals F100 to F110 as the display control abnormal frame segment division, records the start frame number and end frame number of the division, and generates a division index list and stores it in the register file.

[0025] S213: Call the display control abnormal frame segment division, perform integrated analysis on instruction and action matching degree, voice persistence dispersion, buffer queue stability and image loading delay, calculate display control abnormality sensitivity, identify sensitivity layer and marked response block for frame segment matching, and form display control abnormal area identification result; The system calls upon the display control anomaly frame segment division to calculate four quantitative indicators: instruction-action matching degree, speech persistence dispersion, buffer queue stability, and image loading latency. The matching degree calculation logic counts the proportion of nodes with a positive matching state within the segment relative to the total number of nodes. The dispersion calculation logic calculates the standard deviation of all speech instruction durations within the segment and normalizes it. The stability calculation logic reads the remaining frame sequence in the frame buffer and calculates the variance. The latency calculation logic counts the average time difference between instruction issuance and frame buffer update. The weighted calculation unit, based on the weight coefficients determined by the analytic hierarchy process (AHP), multiplies the four normalized indicators by coefficients of 0.4, 0.2, 0.2, and 0.2 respectively, and then sums them to obtain the display control anomaly sensitivity value. The comparator compares the calculated sensitivity with a threshold of 0.6. If the result is greater than the threshold, the region is marked as a high-sensitivity anomaly block. The region identification logic writes the start frame number 500 and end frame number 520 corresponding to the high-sensitivity block and its sensitivity value 0.74 into the final list of abnormal region identification results for display control, completing the dimensionality reduction mapping from multi-dimensional features to a single risk level, and providing a quantitative basis for resource scheduling.

[0026] Please see Figure 4 The specific steps for obtaining the voice response performance evaluation data table are as follows: S311: Based on the results of abnormal region identification in display control, extract the cache queue depth sequence and image loading delay timestamp of the numbered frame segment in the layer, perform time alignment on the data in the frame segment, identify the cache depth fluctuation amplitude and loading delay mutation amount, and obtain the local response anomaly feature set; Based on the frame number in the display control anomaly region identification results, the processor delves into the underlying driver to extract detailed performance logs, including a cache queue depth sequence sampled at 10ms intervals and a timestamp of the image loading delay from decoding completion and submission to the buffer to actual screen display for each frame. The processor performs time alignment, mapping the two sequences to a unified system monotonic time axis to eliminate time base differences. Next, the processor calculates the cache depth fluctuation amplitude, i.e., the difference between the maximum and minimum depths in the sequence, and simultaneously calculates the loading delay abrupt change, i.e., the maximum value of the first derivative of the loading delay. For example, during an anomaly frame segment, if the maximum value of the sampled cache depth sequence is 3 and the minimum value is 0, then the fluctuation amplitude is 3; if the loading delay sequence abruptly changes from 30ms to 150ms, detecting a significant peak, the processor encapsulates the specific numerical features into a local response anomaly feature set. This feature set directly reflects the physical bottlenecks in the underlying data flow, revealing the micro-mechanisms leading to upper-layer display anomalies, such as buffer underloading or rendering pipeline congestion.

[0027] S312: Based on the local response anomaly feature set, perform coupled analysis on buffer fill rate and image switching frequency, construct buffer and switching coordination index based on the normalized difference between image switching frequency and buffer fill rate, filter frame segments with coordination index exceeding preset threshold, and establish buffer and switching coordination response distribution map. Based on the local response anomaly feature set, the processor performs coupled analysis on the cache fill rate and image switching frequency to assess the supply and demand balance between production and consumption. The processor calculates the cache fill rate of the number of frames submitted to the buffer by the decoder per unit time, and the image switching frequency of the screen actually refreshing to display new content. A cache-switching coordination index is constructed based on the normalized difference between the image switching frequency and the cache fill rate. This index is the absolute value of the difference between the ratio of the two and 1, reflecting the degree of supply and demand deviation. Ideally, the two are equal, and the index is 0. When the fill rate is lower than the switching frequency, it leads to cache exhaustion; conversely, it increases latency. The processor sets a preset threshold of 0.5. When the coordination index exceeds this threshold, it indicates that the supply and demand difference is too large, and visually perceptible stuttering or frame skipping is significant. The processor traverses all frame segments, filters out frame segments with excessive coordination indices, and establishes a cache-switching coordination response distribution map with the frame number as the horizontal axis and the coordination index value as the vertical axis. This map visually displays the evolution process and distribution density of the temporal supply and demand imbalance.

[0028] S313: Call the cache and switching collaborative response distribution map, cluster the frames that continuously exceed the preset benchmark value in the collaborative index layer, mark the frame number and time coordinate of the abnormal area, and obtain the speech response performance evaluation data table. The processor invokes the cache and switching coordination response distribution map and applies a density-based clustering algorithm to cluster the frames in the coordination index layer that continuously exceed the preset benchmark value. The neighborhood radius is set to 2 frames and the minimum number of points is 3. If there are regions with more than 3 consecutive frames and a frame interval of no more than 2 frames, and their coordination index values ​​all exceed the benchmark value of 0.5, they are clustered into an independent low-performance event. For each cluster, the processor extracts its start frame number, end frame number, duration, and average coordination index, and outputs the information in a structured manner as a speech response performance evaluation data table, as shown in Table 2. This table records the specific parameters of each abnormal event. Events E001 and E003 show extremely high coordination indices, indicating that within the time period, the speech response performance is limited by the coordination problem between cache and display, clarifying the key optimization targets and urgency of subsequent scheduling. Table 2: Voice Response Performance Evaluation Data Table As shown in Table 2, events E001 and E003 exhibit extremely high coordination indices, indicating that during these time periods, the system's voice response performance is limited by the coordination issues between caching and display, making them key targets for subsequent scheduling optimization.

[0029] Please see Figure 5 The specific steps for obtaining the real-time photo frame display abnormal adjustment compensation instruction set are as follows: S411: Based on the frame segment number in the speech response performance evaluation data table, extract the duration sequence of the image frame under the specified number in the display buffer, perform time normalization processing, and use the formula: ; Calculate the rate of change of unit frame dwell time and obtain the set of abnormal frame dwell time change rates; in, The rate of change of the duration of the unit frame. Representative number is The duration for which an image frame remains in the display buffer. Representative number is The duration for which an image frame remains in the display buffer. This represents the arithmetic mean of the dwell times of all image frames within the sequence. This represents the total number of image frames extracted. Based on the frame segment number of the voice response performance evaluation data table, the processor extracts the actual dwell time (DwellTime) of each frame in the display buffer for the abnormal frame segment (such as segment E001) in the evaluation table. Dwell time refers to the time difference from when the frame is placed on the screen (Scan-out) to when the next frame replaces it. The extracted original duration sequence is subjected to Min-Max normalization to eliminate the difference in basic units between different playback modes (such as 5-second loop and 30fps video). In the formula, the parameters Representative number is The duration (in milliseconds) during which an image frame remains in the display buffer. This represents the duration of the previous frame; This represents the arithmetic mean of the dwell times of all image frames within the sequence. This represents the total number of image frames extracted. The formula The design intent is to capture the degree of "jitter" during frame dwell time. As a smoothing reference term, it is used to measure Expected deviation relative to the average level This directly reflects the abrupt changes between adjacent frames; Example verification: Suppose we extract 3 frames of data from an abnormal segment ( The dwell time sequence is (Unit: milliseconds, indicating noticeable stuttering / fast-forwarding), calculate the average value. ; calculate The item of time ( ): Difference item: ; Reference items: ; Absolute value: ; calculate The item of time ( ): Difference item: ; Reference items: ; Absolute value: ; Total rate of change ; The result A large value indicates that the dwell time within this frame segment exhibits drastic nonlinear fluctuations (erratic speeds), meaning the abnormal rate of change in frame dwell time is extremely high. The processor calculates the dwell time for all target frame segments... The values ​​constitute the set of abnormal change rates of frame dwell time.

[0030] S412: Based on the set of abnormal change rates of frame dwell time, retrieve the standard playback rhythm curve as a reference, compare the Euclidean distance between the current dwell time sequence and the reference curve, define the deviation, identify the frame segments whose deviation exceeds the warning threshold, and obtain the set of rhythm change frame segments. Based on the set of abnormal frame dwell time rates, the processor retrieves a preset standard playback rhythm curve from memory as a reference. It compares the current dwell time sequence with the standard curve. In 30fps video mode, the standard curve is a constant 33.3ms, and in slideshow mode, it is a constant 5000ms. The processor uses an Euclidean distance algorithm to calculate the deviation between the actual captured dwell time sequence vector and the standard curve vector. This deviation reflects the overall deviation of the actual playback rhythm from the ideal rhythm. A warning threshold is set; if the allowable error is ±10%, then for a certain length of sequence, there is a corresponding cumulative allowable deviation threshold. The processor identifies frame segments with deviations exceeding the warning threshold and includes them in the rhythm abrupt change frame segment set. For example, if the calculated deviation far exceeds the threshold, it indicates that the frame segment has experienced severe rhythm collapse, such as a catch-up phenomenon after a long pause. Through this step, the processor accurately locates the target area requiring rhythm intervention.

[0031] S413: Based on the rhythm change frame segment set, bind the deviation value of each frame segment to the frame number in the digital photo frame display logic diagram, arrange them in descending order of deviation, and output the real-time photo frame display abnormal adjustment compensation instruction set. Based on the rhythm mutation frame segment set, the processor binds the deviation value of each frame segment to the frame number in the digital photo frame display logic diagram. This logic diagram details the physical location of each frame's data in the DRAM address space, decoding buffer, and display controller FIFO. The processor sorts all abnormal frame segments in descending order of deviation and generates corresponding adjustment and compensation instructions by looking up a table. For extremely high deviation frame segments with a deviation greater than 100, a frame skipping compensation instruction is generated to force timestamp synchronization; for high deviation frame segments with a deviation between 50 and 100, a dynamic frequency adjustment instruction is generated to increase the GPU frequency to accelerate rendering; for medium deviation frame segments with a deviation between 20 and 50, a thread priority increase instruction is generated. The processor ultimately outputs a real-time photo frame display abnormal adjustment and compensation instruction set containing specific operation codes and target parameters, providing clear execution strategies and parameter guidance for subsequent hardware collaborative management.

[0032] Please see Figure 6 The specific steps for obtaining the logic coordination control sequence of the digital photo frame display are as follows: S511: Call the real-time photo frame display abnormal adjustment and compensation instruction set, extract the number of the risk frame segment in the digital photo frame display logic diagram, map the risk level of the frame segment to the boundary of the display buffer, identify the frame segment information corresponding to the playback protection level, and generate a display task risk distribution map. The processor invokes the real-time frame display anomaly adjustment and compensation instruction set. It extracts the risky frame segment numbers marked as requiring frame skipping or frequency upscaling, and maps the frame numbers to specific display buffer physical address boundaries using page table information from the memory management unit. The processor classifies the playback protection level based on the risk level of deviation from the quantification, where Level_1 corresponds to the extreme risk of crashing or deadlocking, Level 2 corresponds to the risk of significant stuttering, and Level_3 corresponds to the low risk of slight jitter. The processor uses the physical address space as a base map and overlays the risk levels in the form of a heatmap to generate a display task risk distribution map. This map intuitively shows the distribution of risky areas in memory data processing and clarifies the physical address range that the memory controller needs to focus on, thereby achieving a precise mapping from logical layer risks to physical layer resources.

[0033] S512: Based on the display task risk distribution map, extract the cache rescheduling unit number and scheduling priority, match the frame segment risk level with the cache scheduling level, identify the frame segment number with insufficient scheduling coverage, and obtain the display task scheduling risk disconnect list. Based on the display task risk distribution map, the processor scans and extracts the cache rescheduling unit numbers and their scheduling priorities currently allocated to each buffer block. It then performs a matching check between the risk level and the cache scheduling level. The matching rule requires that Level_1 risk must correspond to the highest priority scheduling channel, and Level_2 must correspond to the second highest priority or higher. If a frame segment is found to be in a Level_1 risk state, but its corresponding cache rescheduling unit is only allocated a low-priority DMA channel, the processor determines that the scheduling coverage is insufficient, i.e., a risk disconnect. The processor records the frame segment numbers, current risk levels, and current scheduling levels of all frames that meet the disconnect conditions, generating a display task scheduling risk disconnect list. This list accurately identifies the root cause of the mismatch between resource allocation and task requirements, providing a direct correction target for subsequent dynamic resource reconfiguration.

[0034] S513: Based on the list of risk disconnection in the display task scheduling, and according to the level number in the playback protection priority sequence, extract the key frame segment number that needs to increase the cache allocation weight, output the adjustment control parameters linked with the original cache unit in sequence, and output the digital photo frame display logic collaborative management sequence. Based on the list of display task scheduling risk disconnections, the processor determines the adjustment order according to the playback protection priority sequence. For each critical frame segment to be adjusted, the required cache allocation weight increment is calculated. For frame segments with Level 1 risk but low scheduling priority, the processor generates control parameters that raise the DMA priority to the highest, reserve the minimum bandwidth to meet the resolution and refresh rate requirements, and set the memory area access latency tolerance to the lowest. The processor arranges the adjustment parameters for all disconnected frame segments in chronological order to form a digital photo frame display logic collaborative management sequence, which is sent to the SoC system controller to reconstruct hardware resource allocation in real time. As shown in Table 3, different risk levels correspond to different management strategies. By executing this sequence, the voice command response latency is significantly reduced in high-load scenarios, and screen tearing and stuttering are greatly reduced, ensuring the synchronization and smoothness of voice interaction and screen display. Table 3: Mapping Table of Risk Level and Collaborative Management Parameters By executing this control sequence, experimental data shows that under high-load scenarios, the voice command response latency increases from the average... Reduce to Screen tearing and stuttering were reduced by 92%, significantly improving the smoothness of digital photo frames.

[0035] The voice recognition-based interactive digital photo frame display control system is used to execute the aforementioned voice recognition-based interactive digital photo frame display control method. The system includes: The voice synchronization monitoring module captures the user's voice signal and identifies the start and end time points based on the preset voice command lexicon and the current display status of the digital photo frame. It compares the current image frame number with the playback mode, filters the voice and display synchronization failure intervals, extracts the frame number and timestamp, and generates a voice and display synchronization status recognition dataset. The command localization module is based on the voice and display synchronization status recognition dataset. It identifies the temporal matching relationship between voice command keywords and image switching direction vectors, marks the node numbers of image sequences, matches the digital photo frame display logic diagram, extracts the range of misaligned command and action frames, and establishes the recognition results of abnormal display control areas. The cache linkage module, based on the results of abnormal area identification in display control, retrieves continuous data of cache queue depth sequence and image loading delay features in the frame segment, judges the coordination between cache fill rate and switching frequency, marks the frame segment number that meets the coordination failure threshold, and outputs a voice response performance evaluation data table. The rhythm warning module analyzes the distribution trend of the dwell time of the corresponding image frame and the degree of deviation from the standard playback rhythm curve based on the frame segment number of the voice response performance evaluation data table, extracts the rhythm deviation frame segment number, completes the level identification according to the risk classification standard, and generates a set of real-time frame display abnormal adjustment and compensation instructions. The cache optimization module, based on the real-time photo frame display anomaly control and compensation instruction set, finds the corresponding position number of the risky frame segment in the digital photo frame display logic diagram, retrieves the current cache rescheduling unit configuration list, compares the playback protection priority with the current scheduling level, filters the frame segment numbers that need to update the cache allocation strategy, and outputs the digital photo frame display logic collaborative control sequence.

[0036] The above embodiments illustrate preferred embodiments of the present invention. Any equivalent adjustments to the technical solution based on software engineering methods are within the scope of protection, including but not limited to: implementing algorithm logic using different programming languages, refactoring functional modules into services, adjusting data interaction protocols, and optimizing resource scheduling strategies. Any implementation scheme derived from reasonable modifications to the data processing flow, service call chain, or system architecture layer without departing from the core technology of the present invention should be considered within the protection scope defined by the technical solution of the present invention.

Claims

1. A digital photo frame display control method based on voice recognition interaction, characterized in that, Includes the following steps: S1: Based on the preset voice command lexicon and the current display status of the digital photo frame, capture the voice signal emitted by the user, extract the start and end points of the voice in the time domain, and combine the current image frame number and playback mode identifier to generate a voice and display synchronization status recognition dataset. S2: Based on the voice and display synchronization state recognition dataset, extract the voice command keywords and the current image switching direction vector, analyze the temporal matching degree of the two at the image sequence nodes, filter the image frame intervals where there is misalignment between commands and actions and the voice duration exceeds the preset threshold, and form the display control abnormal area recognition result. S3: Based on the results of the abnormal area identification of the display control, extract the cache queue depth and image loading delay features of the corresponding image frame interval, analyze the coordination between the cache filling rate and the image switching frequency, filter the frame segments where cache overflow or loading lag occurs continuously, and obtain the voice response performance evaluation data table. S4: Based on the voice response performance evaluation data table, analyze the distribution trend of the dwell time of the corresponding image frame in the display buffer, evaluate the degree of deviation from the standard playback rhythm curve, mark the abnormality level of the frame segment according to the deviation magnitude, and output the real-time frame display abnormality adjustment compensation instruction set.

2. The digital photo frame display control method based on voice recognition interaction according to claim 1, characterized in that, The voice and display synchronization status recognition dataset includes the voice start frame number, instruction keyword matching failure marker, current playback mode identifier, and voice end timestamp. The display control abnormal area recognition results include instruction and action misalignment frame segment identifiers, voice continuous over-limit intervals, and image sequence node mismatch blocks. The voice response performance evaluation data table includes continuous frame segments with abnormal cache depth, image loading delay abrupt change areas, cache fill rate mismatch areas, and response failure frame numbers. The real-time frame display abnormal adjustment and compensation instruction set includes risk level labels, frame response deviation values, cache overflow frequency indicators, and playback rhythm offset levels.

3. The digital photo frame display control method based on voice recognition interaction according to claim 1, characterized in that, The specific steps for obtaining the voice and display synchronization status recognition dataset are as follows: S111: Based on the preset voice command dictionary and the current display status of the digital photo frame, capture the user's voice signal and perform endpoint detection, extract the start time and end time of the voice signal, and record the current image frame number and playback mode type to construct a voice and image time-series aligned sequence. S112: Based on the aforementioned speech and image time-series alignment sequence, the following formula is used: ; Calculate the time-series ratio feature value of speech and image, identify the time window when the ratio exceeds the preset threshold, and simultaneously determine whether the instruction keywords are successfully matched in the dictionary, and generate a joint judgment set of instruction validity and duration; in, The feature value representing the time-series ratio of speech to image is... Representing the The duration of a speech segment Representative and the Image switching interval for segmented speech alignment The arithmetic mean of the image switching intervals. This represents the total number of segments where speech and images are aligned. S113: For the joint determination set of the validity and duration of the instruction, extract the voice events that fail to match and exceed the duration limit, bind the corresponding image frame number and playback mode identifier, and generate a voice and display synchronization status recognition dataset.

4. The digital photo frame display control method based on voice recognition interaction according to claim 3, characterized in that, The specific steps for obtaining the identification results of the abnormal display control area are as follows: S211: Based on the voice and display synchronization state recognition dataset, extract the expected image switching direction vector corresponding to the voice command keywords, compare it with the real-time image frame jump direction, identify the symbol consistency between the two at the image sequence nodes, and generate a command and action matching map. S212: Based on the instruction and action matching map, filter the node regions where the symbols are inconsistent and the voice duration is greater than the preset benchmark value. Combined with the continuity constraint of the image sequence, identify the misaligned concentrated frame segments belonging to the same playback cycle to obtain the display control abnormal frame segment division. S213: Invoke the display control abnormal frame segment division, perform integrated analysis on instruction and action matching degree, voice persistence dispersion, cache queue stability and image loading delay, calculate display control abnormality sensitivity, identify the sensitivity layer and mark the response block for frame segment matching, and form display control abnormality area identification result.

5. The digital photo frame display control method based on voice recognition interaction according to claim 4, characterized in that, The specific steps for obtaining the voice response performance evaluation data table are as follows: S311: Based on the display control abnormal area identification results, extract the cache queue depth sequence and image loading delay timestamp of the numbered frame segment in the layer, perform time alignment on the data in the frame segment, identify the cache depth fluctuation amplitude and loading delay mutation amount, and obtain the local response abnormal feature set. S312: Based on the local response anomaly feature set, perform coupled analysis on the buffer fill rate and image switching frequency, construct a buffer and switching coordination index based on the normalized difference between the image switching frequency and the buffer fill rate, filter out frame segments whose coordination index exceeds a preset threshold, and establish a buffer and switching coordination response distribution map. S313: Call the cache and switching collaborative response distribution map, cluster the frames that continuously exceed the preset benchmark value in the collaborative index layer, mark the frame number and time coordinate corresponding to the abnormal area, and obtain the voice response performance evaluation data table.

6. The digital photo frame display control method based on voice recognition interaction according to claim 5, characterized in that, The specific steps for obtaining the real-time photo frame display anomaly adjustment and compensation instruction set are as follows: S411: Based on the frame segment number of the voice response performance evaluation data table, extract the sequence of the dwell time of the image frame under the specified number in the display buffer, perform time normalization processing, calculate the change rate of the dwell time per unit frame, and obtain the set of abnormal change rates of the dwell time of the frame. S412: Based on the set of abnormal change rates of frame dwell time, retrieve the standard playback rhythm curve as a reference, compare the Euclidean distance between the current dwell time sequence and the reference curve, define the deviation, identify the frame segments whose deviation exceeds the warning threshold, and obtain the set of rhythm change frame segments. S413: Based on the rhythm mutation frame segment set, bind the deviation value of each frame segment to the frame number in the digital photo frame display logic diagram, arrange them in descending order of deviation, and output the real-time photo frame display abnormal adjustment compensation instruction set.

7. The digital photo frame display control method based on voice recognition interaction according to claim 1, characterized in that, The method also includes step S5: S5: Call the real-time photo frame display abnormal adjustment and compensation instruction set, identify the corresponding number of the risk frame segment in the digital photo frame display logic diagram, extract the image cache rescheduling unit list, compare the cache scheduling priority with the playback protection sequence, filter the frame segment number that needs to adjust the cache allocation strategy, and output the digital photo frame display logic collaborative control sequence. The digital photo frame display logic collaborative control sequence includes adjusting the target frame segment number, cache depth adjustment parameters, playback protection priority comparison items, and image preloading trigger type.

8. The digital photo frame display control method based on voice recognition interaction according to claim 7, characterized in that, The specific steps for obtaining the digital photo frame display logic collaborative control sequence are as follows: S511: Call the real-time photo frame display abnormality control and compensation instruction set, extract the number of the risk frame segment in the digital photo frame display logic diagram, map the risk level of the frame segment to the boundary of the display buffer, identify the frame segment information corresponding to the playback protection level, and generate a display task risk distribution map. S512: Based on the display task risk distribution map, extract the cache rescheduling unit number and scheduling priority, match the frame segment risk level with the cache scheduling level, identify the frame segment number with insufficient scheduling coverage, and obtain the display task scheduling risk disconnect list. S513: Based on the display task scheduling risk disconnect list, according to the level number in the playback protection priority sequence, extract the key frame segment number that needs to increase the cache allocation weight, output the adjustment control parameters linked with the original cache unit in sequence, and output the digital photo frame display logic collaborative management sequence.

9. A digital photo frame display control system based on voice recognition interaction, characterized in that, The system is used to implement the digital photo frame display control method based on voice recognition interaction as described in any one of claims 1-8, and the system includes: The voice synchronization monitoring module captures the user's voice signal and identifies the start and end time points based on the preset voice command lexicon and the current display status of the digital photo frame. It compares the current image frame number with the playback mode, filters the voice and display synchronization failure intervals, extracts the frame number and timestamp, and generates a voice and display synchronization status recognition dataset. Based on the voice and display synchronization status recognition dataset, the command localization module identifies the temporal matching relationship between voice command keywords and image switching direction vectors, calibrates the image sequence node numbers, matches the digital photo frame display logic diagram, extracts the range of misaligned command and action frames, and establishes the display control abnormal area recognition result. Based on the results of the abnormal area identification of the display control, the cache linkage module retrieves the continuous data of the cache queue depth sequence and image loading delay features in the frame segment, judges the coordination between the cache fill rate and the switching frequency, marks the frame segment number that meets the coordination failure threshold, and outputs the voice response performance evaluation data table. The rhythm warning module analyzes the distribution trend of the dwell time of the corresponding image frame and the degree of deviation of the standard playback rhythm curve based on the frame segment number of the voice response performance evaluation data table, extracts the rhythm deviation frame segment number, completes the level identification according to the risk classification standard, and generates a real-time frame display abnormal adjustment compensation instruction set. The cache optimization module, based on the real-time photo frame display anomaly control and compensation instruction set, finds the corresponding position number of the risky frame segment in the digital photo frame display logic diagram, retrieves the current cache rescheduling unit configuration list, compares the playback protection priority with the current scheduling level, filters the frame segment numbers that need to update the cache allocation strategy, and outputs the digital photo frame display logic collaborative control sequence.