AI voice data interaction system of smart television
By leveraging the AI voice data interaction system of smart TVs, combined with local systems and big data models, the problems of insufficient personalized experience and low interaction efficiency of smart TVs have been solved, enabling accurate content push and personalized management, thereby improving user experience and security.
Patent Information
- Application Number
- CN202511014214.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-12-19
AI Technical Summary
Existing smart TVs suffer from problems such as insufficient personalized experience, low interaction efficiency, rigid interface logic, limited multi-screen interaction, weak content resource integration capabilities, lack of content management, inadequate system maintenance and security, and lack of advanced functions.
The system employs an AI voice data interaction system, comprising an input layer, a processing layer, an output layer, and a support layer. By combining a local system with a big data model, it achieves multimodal result parsing, dynamic resource scheduling, and dynamic rendering. It supports accurate analysis of voice data and personalized recommendations, adapts to parameter adjustments on different display devices, and provides multimodal interaction and cross-platform resource integration.
It improves the response speed and interaction efficiency of smart TVs, enables accurate content delivery and personalized experience, enhances user stickiness, improves visual experience and security, reduces latency and color differences between devices, and supports cross-platform resource access and multi-user management.
Smart Images

Figure CN121173992A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of smart TVs, and particularly to an AI voice data interaction system of a smart TV. BACKGROUND
[0002] A smart TV is a television product based on Internet application technology, with an open operating system and a chip, an open application platform, and a two-way man-machine interaction function, integrating audio, video, entertainment, data and other functions to meet the diversified and personalized needs of users. Its purpose is to provide users with more convenient experiences, and it has become a trend in the television industry.
[0003] 1. Insufficient personalized experience and mechanical content recommendation
[0004] That is, it mainly displays cross-sections and then recommends user behavior (such as viewing history and preferences), and can only provide general recommendations (such as popular lists), and cannot achieve precise push of user preferences.
[0005] Example: A user often watches science fiction movies, but the TV still frequently recommends romance dramas. 2. Low interaction efficiency, missing or weak voice assistant function
[0006] It cannot be controlled by natural language instructions (such as "fast forward to 10 minutes" or "help me find a comedy movie with a rating of 9.0 or above"), and relies on traditional remote control operation, which is not user-friendly for older users or complex search scenarios.
[0007] 3. Interface logic is fixed
[0008] The menu level is complex, lacks intelligent sorting (such as placing frequently used applications at the top), and users need to click multiple times to find the target function.
[0009] 4. Limited multi-screen interaction
[0010] The screen projection function may rely on third-party applications, with high latency or poor compatibility, and cannot seamlessly connect to mobile phone / computer content.
[0011] 5. Weak content resource integration capability
[0012] Cross-platform search is inefficient, and requires manual switching between different video platforms (such as iQIYI and Tencent Video) to find content, and cannot aggregate results through a unified search bar.
[0013] 6. Lack of content management
[0014] It cannot automatically categorize content (such as filtering by type, actor, or year), and users need to manually organize their viewing lists.
[0015] 7. Limited system maintenance and security, and update support
[0016] System vulnerability repair is slow, new features (such as HDR format support) are iteratively lagging behind, and are easily eliminated; privacy protection risks; some low-end televisions may lack data encryption mechanisms, and user viewing habits and other privacy are easily leaked.
[0017] 8. Lack of advanced functions:
[0018] Insufficient image / audio optimization; no AI dynamic tuning (such as optimizing contrast and noise reduction according to the scene); audio-visual experience is inferior to high-end models; no multi-user account, when sharing devices in a family, all user history records are mixed, and personalized settings cannot be distinguished. SUMMARY
[0019] The main purpose of the present application is to provide an AI voice data interaction system for smart TV, aiming to solve the problem of data interaction between existing smart TV and users.
[0020] To achieve the above purpose, the present application provides an AI voice data interaction system for smart TV, comprising:
[0021] An input layer for collecting voice data of users;
[0022] A processing layer for analyzing voice data of users and converting voice data and making intelligent decisions;
[0023] An output layer for generating multi-modal result feedback according to data conversion and intelligent decision-making, and providing it to users;
[0024] A support layer for infrastructure and optimization technology.
[0025] The technical scheme of the present application has the following advantages:
[0026] 1. The local system is combined with the existing data model, and the local data calling and the large model data calling are combined, which improves the rapid response of the intelligent system. Even if a chip with poor operation is used, the operation frequency can be improved through the network large model to reduce the problem of lag, thereby improving the market competitiveness while improving the cost of smart TV;
[0027] 2. The multi-modal result analysis logic is adopted, which solves the problem of messy existing data analysis and user selection difficulty. That is, through historical storage records and user search and use records, the prediction of user demand is improved, the selection tendency of users is improved, and the user experience is improved;
[0028] 3. Through the analysis of voice data, the differentiation of users, emotion judgment and dialect judgment are further realized, thereby improving the stickiness of the intelligent system and customers, and through the predetermined scene, the multi-modal structure analysis is reasonably provided through the big data model,
[0029] For example, after adding + predetermined emotion + predetermined user after multiple search words, the demand of the user can be more accurately judged,
[0030] At the same time, the accuracy of data pushing is improved through multi-dimensional analysis;
[0031] 4. Through dynamic resource scheduling and dynamic rendering, the calling and decoding deepening of video resources are realized, which can be adaptively adjusted through the parameters of different display devices, thereby improving the visual experience of users, and reducing the problem that the existing television or display device cannot adjust the color according to the actual demand, thereby the color difference with the actual color is large, affecting the use of users. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is the framework diagram of the present application;
[0033] Figure 2 is the logic diagram of the present application;
[0034] Figure 3 is the input layer technical annotation flowchart;
[0035] Figure 4 is the processing layer technical annotation flowchart;
[0036] Figure 5 is the output layer technical annotation flowchart. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0038] It should be noted that if the present application embodiments involve directional indications (such as up, down, left, right, front, back, top, bottom, inside, outside, vertical, horizontal, longitudinal, counterclockwise, clockwise, circumferential, radial, axial, …), the directional indications are only used to explain the relative position relationship, movement condition, etc. between the components in a certain posture (as shown in the drawings), if the certain posture changes, the directional indications also change accordingly.
[0039] In addition, if the description of "first" or "second" and the like is involved in the embodiments of the present application, the description of "first" or "second" and the like is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can be explicitly or implicitly included at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of ordinary skilled in the art, when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, also not within the protection scope required by the present application.
[0040] As shown in the figure, an AI voice data interaction system of a smart TV includes: Figures 1 to 5
[0041] An input layer is used for collecting voice data of a user.
[0042] A processing layer is used for analyzing voice data of a user and converting and intelligently deciding voice data.
[0043] An output layer generates a multi-modal result feedback according to data conversion and intelligent decision, and provides it to a user.
[0044] A support layer is used for infrastructure and optimization technology.
[0045] Specifically, the data processing process of the processing layer includes:
[0046] S1: A user transmits voice to the input layer through a Bluetooth remote controller, the input layer identifies voice data and generates text data.
[0047] S2: The text data is sent to a big data processing model, the big data model is used to analyze and list the text data,
[0048] The data structure of the listing includes multi-modal result analysis,
[0049] The big data model analyzes the listed data, filters and selects the user's selection through historical records,
[0050] S3: When the user selects the multi-modal structure, the output layer makes a decision through resource scheduling of the support layer, calls a predetermined program, APP or software, and then executes the predetermined video and audio through a local player and presents it through a display.
[0051] Specifically, the big data processing model is a Deepseek big model, a Quark big data or an Open big data, preferably an open source big model, a Deepseek big model,
[0052] Wherein the model can feedback and analyze the data information sent to it, and list the data, and then tabulate the data, thereby facilitating the user's selection.
[0053] Specifically, the S1 includes user voice, Bluetooth remote control microphone array, noise reduction processing, audio compression encoding and BLE encrypted transmission.
[0054] The microphone array of the Bluetooth remote control is used to analyze the voice data of the predetermined user through multiple microphones, and the voice data is processed by the built-in software.
[0055] Reduce external interference, thereby improving the input accuracy of voice text, and reducing the stability of voice input.
[0056] At the same time, through compression encoding and BLE encrypted transmission, data leakage in the data transmission process is reduced, that is, without going through the predetermined decoding program, it is parsed into garbled voice, thereby improving user privacy.
[0057] Further, the microphone array: beamforming technology suppresses environmental noise.
[0058] Encoding format: Opus low bit rate high fidelity compression.
[0059] Transmission protocol: BLE 5.0 (<20ms delay).
[0060] Specifically, the input layer accesses the iflytek voice input method. Through the open source data processing center, combined with the built-in processing logic, the system cost is saved while improving the user experience.
[0061] Specifically, the voice data is recognized and text data is generated through local or cloud offloading analysis,
[0062] When the instruction is simple, it is analyzed locally.
[0063] When the instruction is complex, it is analyzed through the cloud.
[0064] Specifically, the local can be an intelligent television with only an APP, or through the cloud, such as a big data model.
[0065] Specifically, the instruction includes a shunt strategy.
[0066] Fast local response based on word table matching (such as "volume + 10")
[0067] Intention recognition: NER (Named Entity Recognition) and multi-label classification and instruction analysis, thereby predicting the next instruction of the user.
[0068] Ranking algorithm: CTR estimation (click rate) and user portrait weighting, which is the high-frequency use of the user's CTR estimation and user portrait weighting for the predetermined user, and then as a prediction instruction option.
[0069] Specifically, the output layer (such as Figure 4 ) technical annotation process includes:
[0070] Resource aggregation: The resource aggregation includes copyright verification and CDN options, the resource aggregation plays or selects the user with the lowest data source priority delay, and the copyright verification includes whether the copyright is charged, whether it needs to use a targeted player;
[0071] Dynamic rendering: The dynamic rendering includes JSON Schema driven UI, different templates are matched through different content types, and then the video is decoded through the built-in display module, and the deepening of vision is realized;
[0072] Improve the existing deepening of vision by playing APP adjustment, mainly different TVs, and the color exists deviation, so the adjustment of the app has limitations.
[0073] Decoding optimization: The decoding optimization includes the TV chip hard HEVC / H.266 format.
[0074] Specifically, mixed ASR shunting includes:
[0075] Local model, the local model is a lightweight model optimized for high-frequency instructions, and the lightweight model supports offline recognition when the data bandwidth is <10MB;
[0076] Cloud model, the cloud model is used to support long speech, dialect, and complete ASR service of Chinese and English mixed;
[0077] For example, in the following embodiments:
[0078] DeepSeek model inference.
[0079] Input: cleaned text + user historical behavior features (implicit feedback);
[0080] Output: structured query instruction (such as {type: "video", query: "chicken wings", source: ["TikTok", "Bilibili"], filter: "time <5 minutes"})
[0081] Resource scheduling decision, the resource scheduling decision logic includes:
[0082] Multi-objective optimization: copyright compliance > play smoothness > picture clarity > platform membership status.
[0083] 3.1 Natural Interaction: Human-like Conversational Ability:
[0084] Ambiguous Command Resolution: Support incomplete sentences or colloquial expressions (such as "What did that actor just do?" or "Skip to the exciting part"), without strictly following fixed grammar.
[0085] Context Memory: Can associate multiple rounds of dialogue (such as after the user asks "Recommend science fiction films" and then asks "Don't include American ones", the system automatically filters out non-American science fiction films).
[0086] Emotional Interaction: Through semantic analysis and emotion recognition, respond more human-like (such as when the user says "So boring", recommend comedy films with humorous captions).
[0087] 3.2 Precise Content Service
[0088] ① Cross-platform content penetration:
[0089] Integrate resources from platforms such as iQiyi, Tencent, Bilibili, etc. No need to manually switch APP, directly voice search "play 'Three-Body' " to automatically jump to the platform with copyright.
[0090] ② Dynamic preference learning
[0091] Based on viewing time, pause / fast-forward behavior, ratings, etc. Real-time update user portrait (such as identifying user preference for "suspense + short drama" and preferentially recommending).
[0092] ③ Scene-based recommendation
[0093] Combine external data such as time and weather (such as recommending healing movies on rainy days, automatically adjusting brightness at night and recommending sleep white noise).
[0094] 3.3 Multi-modal interaction upgrade
[0095] ① Voiceprint recognition:
[0096] Distinguish family user identity, automatically switch personalized interface (such as shielding adult content in child mode, enlarging font in old man mode).
[0097] 4.1 Deep technology integration
[0098] Low-latency response
[0099] Local + cloud hybrid computing, even in weak network environment can still execute basic commands (such as adjusting volume).
[0100] Dialect / accent adaptation
[0101] Support Cantonese, Sichuan dialect, and even mixed English and Chinese instructions (such as "play Taylor Swift's MV").
[0102] Professional field optimization
[0103] Train a specialized model for the film and television field to understand complex film titles (such as "play the 1994 Luc Besson-directed 'The Killer'").
[0104] 4.2 Smart home hub
[0105] Cross-device task flow voice command "cast the ball game on the phone to the TV" automatically wakes up the TV and starts casting without manual operation.
[0106] 5.1 Solution selection:
[0107] Android smart solution can choose MTK9633:
[0108] 5.1.1 Android smart solution selects MTK9633 solution, runs android14 system, connects Bluetooth voice remote control, TV pre-installs Xunfei voice system, accesses deepseek big data model, provides voice service, converts text output or voice instruction to operate smart TV applications.
[0109] 5.2 Technical solution:
[0110] Combine Bluetooth voice remote control, voice recognition technology, DeepSeek big data model and smart TV display capability to realize "voice → text → search → multi-modal feedback" closed loop, which involves multi-module cooperation and intelligent optimization design.
[0111] 5.2.1 Technical realization principle
[0112] 1. Voice input and transmission
[0113] Bluetooth Low Energy Protocol (BLE):
[0114] The remote control transmits voice data (compressed as audio stream) through BLE protocol, ensuring low delay (<50ms) and power saving, avoiding the audio delay problem of traditional Bluetooth.
[0115] End-to-end encryption: voice data is encrypted (such as AES-256) during transmission to prevent eavesdropping (for example, user privacy instruction "check my purchase record").
[0116] 2. Hybrid recognition of local + cloud voice recognition (ASR):
[0117] Simple instructions (such as "volume up") are recognized by local lightweight model to reduce delay; complex sentences (such as "find Liu XX's 90s police and gangster film") are processed on the cloud.
[0118] Anti-noise and accent adaptation:
[0119] Beamforming microphone array (built-in remote control) to suppress environmental noise, model optimized for dialects (such as Cantonese, Sichuanese), supports mixed English and Chinese instructions (such as "play Taylor Swift's Love Story").
[0120] 3. Natural Language Understanding (NLU) and DeepSeek Model
[0121] Intent classification and entity extraction:
[0122] Model analyzes user query, such as "today's weather in Shenzhen" → intent = weather query, entity = Shenzhen + today; "how to make Coke chicken wings" → intent = recipe teaching, entity = Coke chicken wings.
[0123] Multi-modal knowledge graph invocation:
[0124] Combining structured data (weather API), unstructured data (video tutorials), and real-time information (hot search rankings), dynamically generating answers.
[0125] For example:
[0126] Weather query: call China Meteorological Administration API to get data, use visualization engine to generate temperature curve graph and precipitation probability animation.
[0127] Recipe teaching: grab high-play teaching videos from Bilibili / TikTok, sort by "rating + user preference" and play.
[0128] 4. Result presentation and interaction:
[0129] Dynamic UI rendering engine:
[0130] Automatic template matching according to result type: weather → card-style text and graphics; recipe → step-by-step video + text summary; movie recommendation → horizontal poster stream + Douban rating.
[0131] Multi-modal feedback:
[0132] Support voice broadcast ("Shenzhen will be sunny today, 28°C") and screen display synchronization, meet different scene needs (such as cooking without looking at the screen).
[0133] 5.2.2 Special features of technical implementation
[0134] 1. End-to-end low-latency optimization:
[0135] Edge computing shunting;
[0136] High-frequency instructions (such as "pause" and "fast forward 30 seconds") are processed by the TV's local NPU, with a response time of less than 200ms, avoiding cloud round-trip delay.
[0137] Preloading and caching:
[0138] Based on the user's historical behavior, resources are preloaded (such as data on cities where the weather is frequently checked). When searching for "Shenzhen weather", the cache is directly called, and only the freshness of the data is verified.
[0139] 2. Multimodal semantic understanding;
[0140] Cross-modal alignment:
[0141] Simultaneously understand speech, text, and images (e.g., when a user says "find movies with a similar style to this poster," upload a poster image taken with a mobile phone for cross-modal search).
[0142] Contextual coherence
[0143] Supports multi-turn dialogue memory:
[0144] For example:
[0145] User: "Zhou XX's song" → System plays "Seven Mile Fragrance" → User: "Change to Li X's version" → System automatically searches for "Seven Mile Fragrance Li X Cover".
[0146] 3. Balancing personalization and privacy:
[0147] Differential privacy training:
[0148] The DeepSeek model updates through federated learning, aggregating user behavior patterns without exposing individual data (e.g., user A likes to watch suspense movies, user B likes to watch documentaries; the model learns commonalities but cannot trace specific users).
[0149] Temporary Anonymous Mode:
[0150] Personalized recommendations can be turned off by using the voice command "Enter Guest Mode" and only general data responses will be used.
[0151] 4. Dynamic resource scheduling:
[0152] QoS priority control: Allocate computing power according to the scenario: prioritize streaming media bandwidth during video playback, and downgrade background search to low power mode.
[0153] Multi-source content aggregation: Automatically compares resources from multiple platforms (e.g., when playing "Oppenheimer"), it prioritizes Tencent Video, which the user has a paid membership for; if the copyright is not available, it prompts "iQiyi can play it, but you need to purchase it separately".
[0154] 5.2.3 Typical Scenarios and Technical Examples
[0155] Scenario 1: Weather Inquiry
[0156] User's voice message: "Today's weather in Shenzhen"
[0157] Technical process:
[0158] ASR converts text + NLU extracts "Shenzhen" and "today";
[0159] DeepSeek calls the weather API to obtain structured data;
[0160] The rendering engine generates a visual chart (temperature / humidity / wind speed animation);
[0161] The results are displayed synchronously (text and picture cards) and voice broadcasted ("Shenzhen today is sunny, 28°C");
[0162] Scenario 2: Recipe teaching;
[0163] User voice: "How to make coke chicken wings";
[0164] Technical process:
[0165] The model identifies the intention as "recipe teaching";
[0166] Aggregate videos from platforms such as Douyin, Bilibili, and Xiaochufang, and sort them by "like rate + user dietary preferences (e.g. low oil)";
[0167] Automatically skip the advertisement segment and start playing from the food preparation step;
[0168] The sidebar synchronously displays the text version of the steps (which can be voice-controlled to "jump to step 3");
[0169] 5.3 Description of the drawings:
[0170] 5.3.1 Bluetooth voice remote control, press and hold the left and right direction keys to activate the Bluetooth module and pair with the smart TV. After successful pairing, press and hold the voice button and speak, and the TV voice module automatically analyzes the language statement and outputs it in text format.
[0171] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation, direct / indirect application in other related technical fields, or direct / indirect application in other related technical fields within the inventive concept of the present application, as described in the present application and the drawings, are included in the patent protection scope of the present application.
Claims
1. An AI voice data interaction system of a smart TV, characterized in that, It comprises: an input layer for collecting voice data of a user; a processing layer for analyzing voice data of a user and converting voice data and making intelligent decisions; an output layer for generating multi-modal result feedback according to data conversion and intelligent decisions, and providing the feedback to the user; a support layer for infrastructure and optimization technology.
2. The AI voice data interaction system of the smart TV of claim 1, wherein: the data processing process of the processing layer comprises: S1: the user transmits voice to the input layer through a Bluetooth remote controller, and the input layer identifies the voice data and generates text data; S2: the text data is sent to a big data processing model, and the big data model is used to analyze and list the text data, the data structure of the listing comprises multi-modal result analysis, the big data model analyzes the listed data, filters and screens the user's choices through historical records, and S3: when the user selects the multi-modal structure, the output layer makes a decision through resource scheduling of the support layer, calls a predetermined program, APP or software, executes the predetermined video and audio through a local player, and presents the video and audio through a display. 3.The AI voice data interaction system of the smart TV of claim 2, wherein: The big data processing model is Deepseek, Quark or Open.
4. The AI voice data interaction system of the smart TV of claim 2, wherein: S1 comprises user voice, a microphone array of the Bluetooth remote controller, noise reduction processing, audio compression encoding and BLE encrypted transmission; the microphone array of the Bluetooth remote controller is used to analyze the voice data of the predetermined user through multiple microphones, and the voice data is subjected to noise reduction processing through the built-in software. 5.The AI voice data interaction system of the smart TV of claim 1, wherein: The input layer accesses the Ifly voice input method. 6.The AI voice data interaction system of the smart TV of claim 1, wherein: The voice data is identified and text data is generated through local or cloud-based shunt analysis, when the instruction is simple, local analysis is used; when the instruction is complex, cloud analysis is used.
7. The AI voice data interaction system of the smart TV of claim 6, wherein: the instruction comprises a shunt strategy; fast local response based on word table matching; intention recognition: NER (named entity recognition) and multi-label classification and instruction analysis to predict the user's next instruction; sorting algorithm: CTR estimation (click rate) and user portrait weighting, the user click rate estimation and user portrait weighting are the high-frequency use of the predetermined user, and are used as the predicted instruction options.
8. The AI voice data interaction system of the smart TV of claim 1, wherein: the output layer technology labeling process comprises: resource aggregation: the resource aggregation comprises copyright verification and CDN options, the resource aggregation uses the data source with the lowest priority delay to play or for the user to select, the copyright verification comprises whether the copyright is charged and whether a targeted player is needed; dynamic rendering: the dynamic rendering comprises JSON Schema driven UI, different templates are matched through different content types to decode the video through the built-in display module, and visual deepening is achieved. Decoding optimization: the decoding optimization includes a television chip hard decoding HEVC / H.266 format. 9.The AI voice data interaction system of the smart television of claim 3, characterized in that: Mixed ASR shunting includes: Local model, the local model is a lightweight model optimized for high-frequency instructions, and the data bandwidth of the lightweight model is less than 10MB, supporting offline identification; Cloud model, the cloud model is used to support complete ASR services for long speech, dialect, and Chinese-English mixed speech; DeepSeek model inference; Input: cleaned text + user historical behavior features; Output: structured query instructions; Resource scheduling decision, the resource scheduling decision logic includes: Multi-objective optimization: copyright compliance > playback smoothness > picture clarity > platform member status. 10.The AI voice data interaction system of the smart television of claim 1, characterized in that: The android smart solution of the interaction system selects the MTK9633 solution, carries the android14 system, externally connects the Bluetooth voice remote controller, preinstalls the ifly voice system in the TV, accesses the deepseek big data model, provides voice services, converts text output or voice instruction operation to the application in the smart television.