Voice control method and device for television, equipment and medium

By performing semantic analysis and pattern recognition on the voice commands of smart TVs, the problems of misoperation and rigidity in existing voice control technologies have been solved, achieving efficient and secure voice interaction and proactive services, thus improving the user experience.

CN122069385APending Publication Date: 2026-05-19SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN COOCAA NETWORK TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing voice control solutions for smart TVs struggle to accurately understand users' diverse and colloquial natural language expressions, posing a high risk of misoperation. Furthermore, they cannot proactively initiate guided services under specific device conditions, making it difficult to balance security and convenience in user experience.

Method used

By semantically parsing user voice commands, identifying intent categories and operation keywords, and determining operation modes based on preset mapping relationships, including immediate execution, page jump, or conditional trigger modes, and combining device status and trigger conditions, differentiated voice interaction and control processing can be achieved.

Benefits of technology

It improves the accuracy and security of voice control, reduces the risk of misoperation, increases response efficiency, and provides proactive services under specific conditions, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069385A_ABST
    Figure CN122069385A_ABST
Patent Text Reader

Abstract

The invention relates to a voice control method and device for a television, equipment and a storage medium. The method comprises the following steps: receiving a voice instruction of a user aiming at a television, performing semantic analysis on the voice instruction to obtain an intention category and an operation keyword, and determining a setting operation of the user on a television system and an operation mode corresponding to the setting operation according to the intention category and the operation keyword, the operation mode comprises an immediate execution mode, a page jump mode or a condition triggering mode, and corresponding voice interaction and control processing is executed based on the setting operation, the operation mode, the current equipment state or the preset triggering condition of the television system. According to the method, the accuracy, the safety and the response efficiency of the user to the television setting operation through voice can be improved, and the problems of interaction rigidness, high misoperation risk and lack of active service capability in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of television control technology, and in particular to a voice control method, apparatus, device, and storage medium for television. Background Technology

[0002] Currently, smart TVs generally support voice control, but existing solutions mostly use fixed command word matching or simple intent recognition mechanisms, making it difficult to accurately understand the diverse and colloquial natural language expressions of users. For example, if a user says "Stop sending notifications" or "Turn off notifications," the TV system may not be able to uniformly map these different statements to the same setting. Furthermore, existing voice interaction typically employs a single processing strategy: either executing all operations directly, posing a high risk of misoperation, or always redirecting to a settings page, forcing users to manually confirm simple operations that could be completed directly through a graphical interface. Moreover, it cannot proactively initiate guided services under specific device conditions, resulting in a rigid confirmation mechanism, poor execution reliability, and a difficulty in balancing security and convenience in user experience. Summary of the Invention

[0003] In view of the above, this application provides a voice control method, apparatus, device and storage medium for television, the purpose of which is to solve the above-mentioned technical problems.

[0004] In a first aspect, this application provides a voice control method for a television, the method comprising: Receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Based on the intent category and operation keywords, determine the user's settings operation on the TV system and the corresponding operation mode of the settings operation, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions, corresponding voice interaction and control processing are performed.

[0005] Secondly, this application provides a voice control device for a television, the voice control device for a television comprising: Parsing module: used to receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Determination module: used to determine the user's setting operation on the TV system and the corresponding operation mode based on the intent category and operation keywords, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Control module: used to perform corresponding voice interaction and control processing based on the settings of the TV system, the operation mode, the current device status or preset trigger conditions.

[0006] Thirdly, this application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the steps of the voice control method for a television as described in any embodiment of the first aspect.

[0007] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the voice control method for a television as described in any embodiment of the first aspect.

[0008] The technical solutions provided in this application have the following advantages compared with the prior art: This application parses voice commands into intent categories and operation keywords, and determines the corresponding TV system settings operations and their operation modes based on preset mapping relationships. This achieves differentiated intelligent processing for different operations: high-risk operations are ensured safety through tiered voice confirmation; low-risk operations are executed directly to improve efficiency; complex settings are precisely navigated to and full-duplex voice interaction is enabled to support continuous control; and conditional triggering modes proactively initiate multi-option voice queries when specific device events or state thresholds are met, achieving environment-driven proactive service. By dynamically integrating operation attributes, risk levels, and device status, the accuracy, security, and response efficiency of users' voice-based TV settings operations are improved, solving the problems of rigid interaction, high risk of misoperation, and lack of proactive service capabilities in existing technologies. Attached Figure Description

[0009] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a preferred embodiment of the voice control method for television according to this application; Figure 2This is a schematic diagram of a preferred embodiment of the voice control device for television according to this application; Figure 3 This is a schematic diagram of a preferred embodiment of the electronic device of this application; The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0013] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.

[0014] Reference Figure 1 The diagram shown is a flowchart illustrating an embodiment of the voice control method for television according to this application. The method is executed by an electronic device (e.g., a television terminal), which can be implemented by a software system and / or a hardware system. The voice control method for television includes: Step S10: Receive the user's voice command for the TV, and perform semantic parsing on the voice command to obtain the intent category and operation keywords; Step S20: Based on the intent category and operation keywords, determine the user's setting operation on the TV system and the corresponding operation mode of the setting operation, wherein the operation mode includes immediate execution mode, page jump mode or conditional trigger mode; Step S30: Based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions, perform corresponding voice interaction and control processing.

[0015] In existing smart TV voice interaction, the voice commands issued by users are often semantically ambiguous or expressed in various ways (e.g., "turn off notifications" or "stop popping up messages"). If these are directly mapped to system functions, it can easily lead to misoperation or failure to respond.

[0016] This embodiment converts voice commands into text statements, then uses a pre-trained intent recognition model to identify the intent category of the statement (such as "notification management" or "storage settings"). Combined with keyword extraction rules corresponding to that intent category, it extracts operation keywords (such as "close" or "clean") from the text statement. Through structured semantic understanding, diverse natural language is uniformly transformed into standardized instruction elements that can be processed by the machine, thereby providing reliable input for subsequent precise control and improving the accuracy and generalization ability of voice command recognition.

[0017] Since different user statements may refer to the same function, and the same function requires different processing strategies in different contexts, it is necessary to map the parsing results to specific TV system settings operations and their execution methods. To this end, this method determines the user's TV system settings operation and the corresponding operation mode based on the intent category and operation keywords. The operation mode includes immediate execution mode, page jump mode, or conditional trigger mode. Specifically, a mapping relationship between intent categories, operation keywords, and TV system settings operations is pre-established in a preset configuration table, and a corresponding operation mode is configured for each TV system settings operation. For example, "Turn off notifications" is mapped to the "Notification switch" setting operation, with an immediate execution mode. "View storage" is mapped to the "Storage management" setting operation, with a page jump mode. "What to do after inserting a USB drive" is associated with the "Peripheral content processing" setting operation, with a conditional trigger mode. Through the pre-configured mapping and mode definition, precise routing of user intents is achieved, allowing the selection of the most suitable interaction path based on different operation characteristics.

[0018] To ensure both user experience and operational security and contextual adaptability, this embodiment performs corresponding voice interaction and control processing based on the TV system's settings, operating mode, current device status, or preset trigger conditions. Specifically, if the operating mode is immediate execution mode, the system determines whether to perform voice confirmation based on the risk level of the TV system settings operation, and executes the operation after assessing its feasibility by referring to the current TV device status (such as storage space and network connection).

[0019] If it is a page jump mode, after completing the necessary confirmation, you will be redirected to the corresponding settings page, where the full-duplex voice interaction capability will be activated, allowing users to control the device continuously by voice.

[0020] In condition-triggered mode, when a preset trigger condition associated with the TV system settings is detected (such as a USB device being inserted or storage usage exceeding 90%), a voice prompt with multiple candidate operation options (e.g., "Do you want to play a video or browse a file?") is proactively played, and the corresponding operation is executed based on the user's selection. By dynamically integrating operation attributes, device status, and environmental events, an intelligent leap from passive response to proactive service is achieved, significantly improving the naturalness of voice interaction and the intelligence level of system response while avoiding accidental operation.

[0021] In one embodiment, the semantic parsing of the voice command to obtain the intent category and operation keywords includes: The voice command is subjected to speech recognition to generate a text statement; The text statement is input into a pre-trained intent recognition model to obtain the intent category that matches the text statement; Based on the intent category, the corresponding keyword extraction rules are invoked to extract operation keywords related to the settings operation of the TV system from the text statement.

[0022] In voice interaction scenarios on smart TVs, user-issued voice commands are often conversational, concise, or varied in expression (e.g., "turn off notifications" or "stop pop-up messages"). Directly using the raw voice signal for function matching makes it difficult to accurately understand the user's true intent. Therefore, structured semantic parsing of voice commands is necessary to extract standardized semantic elements that can be recognized and processed by the system.

[0023] Specifically, the system performs speech recognition on user-input voice commands, converting them into corresponding text statements. These text statements are then input into a pre-trained intent recognition model. This model learns the semantic relationships between different statements and TV setting functions based on a large number of labeled samples, enabling it to output the most matching intent category (e.g., "notification management," "display settings," or "storage cleanup"). Based on the identified intent category, it invokes associated keyword extraction rules (e.g., regular expression templates or named entity recognition strategies) to accurately extract operation keywords related to TV system settings operations (e.g., "on," "off," "cleanup," "switch," etc.) from the original text statement. This transforms unstructured natural language into a structured command combination of intent categories and operation keywords, solving the recognition challenges caused by synonymous and heterogeneous expressions. It also provides a clear and reliable semantic foundation for subsequent operation mapping and pattern judgment, improving the accuracy, robustness, and generalization ability of voice control.

[0024] In one embodiment, determining the user's settings operation on the television system and the corresponding operation mode based on the intent category and operation keywords includes: Based on the intent category, retrieve the set of candidate setting operations associated with the intent category from the preset instruction and operation mapping table; From the candidate setting operation set, filter out the setting operations that match the operation keywords; Based on the setting operation matched by the operation keyword, the operation mode corresponding to the setting operation is read from the attribute field configured in the instruction and operation mapping table.

[0025] In smart TV voice control systems, after accurately extracting the intent category and operation keywords from the user's voice, it is also necessary to precisely map these semantic elements to specific TV system functions (i.e., TV system settings operations) and determine how the function should be executed (e.g., directly closing, jumping to the settings page, or waiting for specific conditions to trigger). Relying solely on simple one-to-one mapping is insufficient to address the following practical problems: for example, the same intent category may correspond to multiple similar settings (e.g., "Notification Management" may include multiple sub-items such as "Application Notification Switch," "System Message Prompts," and "Pop-up Permissions"), while the user often expresses only one of these through operation keywords. Furthermore, different settings operations vary in security, interaction complexity, and contextual dependencies, thus requiring differentiated execution strategies.

[0026] Specifically, based on the identified intent category (e.g., "Storage Management"), a set of all candidate setting operations associated with that intent category is retrieved from a pre-defined instruction and operation mapping table. This mapping table is a structured configuration file (such as JSON or a database table), where multiple candidate TV system setting operations are attached under each intent category field. For example, the "Storage Management" intent might include candidate operations such as "Clear Cache," "Format USB Drive," and "View Storage Details." From this set of candidate setting operations, setting operations that semantically match the current operation keyword (e.g., "Clear") are selected. The matching process can be based on exact keyword matching, synonym expansion (e.g., "Delete" is similar to "Clear"), or fuzzy similarity calculation, ultimately determining a unique target setting operation, such as successfully matching "Clear" with "Clear Cache."

[0027] Based on the successfully matched setting operation, its corresponding attribute field is read from the instruction-operation mapping table. This attribute field explicitly records the operation mode of the setting operation. For example, "Clear Cache" is configured as an immediate execution mode (due to its low risk and strong reversibility), "Format" is configured as an immediate execution mode but associated with a high-risk level (requiring secondary confirmation), and "View Storage Details" is configured as a page redirection mode (requiring access to the settings interface to view details). By pre-configuring the operation mode as an inherent attribute of the setting operation, the system can automatically select the optimal execution path based on the security and interaction requirements of different operations, ensuring the security of high-risk operations while improving the response efficiency of low-risk operations.

[0028] In one embodiment, the execution of corresponding voice interaction and control processing based on the settings operation of the television system, the operating mode, the current device status, or preset trigger conditions includes: If the operation mode is the immediate execution mode, the corresponding voice confirmation process is executed according to the risk level associated with the set operation, and the corresponding execution command is generated in combination with the current TV device status; If the operation mode is a page jump mode, the corresponding voice confirmation process is executed according to the risk level associated with the setting operation, an instruction to jump to the corresponding TV system settings page is generated, and full-duplex voice interaction is enabled on the page; If the operation mode is a condition-triggered mode, when a preset trigger condition associated with the setting operation is detected, a voice query containing multiple candidate operation options is generated, and the corresponding TV system setting operation is executed according to the user's voice selection result of the candidate operation options.

[0029] In smart TV voice control systems, different settings operations vary significantly in terms of security, interaction complexity, and contextual dependence. For example, disabling notifications is a low-risk, reversible operation suitable for direct execution. However, restoring factory settings involves data erasure and requires strict verification. Some operations (such as peripheral content processing) should not even be initiated by the user but should be proactively guided by the system under specific device states (such as USB drive insertion). Therefore, the most suitable voice interaction and control strategy can be dynamically selected based on the inherent attributes and operational context of each TV system's settings operations.

[0030] This embodiment divides the settings operations of the television system into three categories: immediate execution mode, page jump mode, and conditional trigger mode, and defines corresponding processing logic for each mode. When the operation mode is immediate execution mode, a graded voice confirmation process is executed based on the risk level associated with the television system settings operation (e.g., high, medium, low). For example, high-risk operations require secondary voice confirmation, while low-risk operations can be performed without confirmation. After successful confirmation, the feasibility of the operation is determined by referring to the current status of the television device (e.g., storage space, network connection, peripheral access status, etc.), and finally, the corresponding system command is generated and executed.

[0031] When the operation mode is page jump mode, the corresponding voice confirmation is executed according to the risk level associated with the setting operation. After the confirmation is successful, a command is generated to jump to the corresponding TV system settings page, and the full-duplex voice interaction capability is activated on the target page, so that the user can continue to control interface elements (such as "open", "back" and "select the second item") through continuous voice commands on the page.

[0032] When the operation mode is condition-triggered, the system does not respond to the user's voice input. Instead, it continuously monitors the device's operating status. When it detects a preset trigger condition associated with the setting operation (e.g., USB storage device inserted, storage usage exceeding a threshold, video playback stuttering), it actively synthesizes and broadcasts a voice prompt containing multiple candidate operation options (e.g., "USB drive detected, do you want to play video or browse files?"). Based on the user's voice selection of the candidate operation options, it executes the corresponding TV system setting operation.

[0033] Furthermore, the step of executing the corresponding voice confirmation process based on the risk level associated with the setting operation, and generating the corresponding execution command in conjunction with the current television device status, includes: Query the risk level corresponding to the TV system settings operation in the preset configuration table, wherein the risk level includes high risk, medium risk or low risk; Based on the risk level, a voice confirmation strategy bound to the risk level is executed, wherein the voice confirmation strategy includes no confirmation required, single voice confirmation, or double voice confirmation. After meeting the requirements of the voice confirmation strategy, the feasibility of executing the setting operation of the TV system is verified based on the current TV device status, and a corresponding execution instruction is generated.

[0034] If the operation mode is immediate execution, the risk level of the TV system settings operation is queried in the preset configuration table. This preset configuration table is predefined and fixed during the system development phase, and each TV system settings operation is associated with a specific risk level (high risk, medium risk, or low risk). For example, "Restore factory settings," "Clear all viewing history," and "Format internal storage" are marked as high risk. "Delete single application cache" and "Reset network settings" are classified as medium risk. "Turn off system notifications," "Switch audio output mode," and "Adjust screen brightness" are defined as low risk.

[0035] Based on the identified risk level, a corresponding voice confirmation strategy is executed. Specifically, low-risk operations use a "no confirmation required" strategy, and the system directly proceeds to the subsequent verification stage. Medium-risk operations trigger a "single voice confirmation," where the system uses speech synthesis to announce the operation to the user and asks, "Are you sure you want to execute?" The process only continues after the voice recognition module detects that the user has explicitly stated "yes" or similar affirmative statements. High-risk operations employ a "two-stage voice confirmation" strategy. After the initial confirmation, the consequences of the operation are emphasized again (e.g., "This operation will erase all personal data and cannot be recovered"), and the user is required to grant a second explicit authorization to ensure the operation's intent is genuine and irreversible.

[0036] After fulfilling all the requirements of the voice confirmation strategy, the current status of the TV device is collected, including key parameters such as storage space utilization, peripheral connection status, network connectivity, and system resource usage. The feasibility of executing the TV system settings operation is then verified. For example, if the operation is "save the log to a USB drive," but no USB storage device is detected to be connected, the execution is terminated and a prompt is generated. If the operation is "clear the cache" and there is sufficient remaining storage space, the corresponding system execution command is generated, and the underlying system interface is called to complete the actual operation.

[0037] Furthermore, the step of executing the corresponding voice confirmation process based on the risk level associated with the setting operation, and generating an instruction to jump to the corresponding TV system settings page, includes: Query the risk level corresponding to the TV system settings operation in the preset configuration table, wherein the risk level includes high risk, medium risk or low risk; Based on the risk level, a voice confirmation strategy bound to the risk level is executed, wherein the voice confirmation strategy includes no confirmation required, single voice confirmation, or double voice confirmation. After meeting the requirements of the voice confirmation strategy, an instruction to jump to the corresponding TV system settings page is generated based on the page path information recorded in the preset configuration table according to the settings operation of the TV system.

[0038] If the operation mode is page jump mode, after meeting all the requirements of the corresponding voice confirmation strategy, the page path information associated with the TV system's settings operation is read from the preset configuration table. This path information can be the URI or Activity name that uniquely identifies the target settings page within the TV system. Based on the identifier, an instruction to jump to the corresponding TV system settings page is generated and handed over to the TV system for page navigation.

[0039] Furthermore, the step of generating a voice query containing multiple candidate operation options when a preset trigger condition associated with the setting operation is detected includes: Read the preset trigger conditions associated with the setting operation of the TV system from the preset configuration table. The preset trigger conditions include TV device events or TV system status thresholds. Real-time monitoring of the operating data of the television equipment, and matching the operating data with the preset trigger conditions for judgment; When the running data meets the preset triggering conditions, a voice query content containing multiple candidate operation options is generated according to the candidate operation option list configured in the preset configuration table based on the settings operation of the TV system.

[0040] During the use of smart TVs, many user needs are not initiated by voice commands, but rather triggered by changes in device status or external events. For example, when a user inserts a USB drive, they may want to play videos, browse files, or perform a firmware upgrade. When storage space is insufficient, they may need to clear the cache, delete applications, or connect external storage devices. If the TV system passively waits for explicit user commands, its interactive intelligence appears low. Blindly pushing operation suggestions may confuse the user or recommend irrelevant options. Therefore, it is possible to predefine which settings operations can be activated under specific conditions, and when the conditions are met, proactively provide the user with a set of context-related candidate operation options, allowing the user to quickly select via voice.

[0041] The preset trigger conditions associated with the TV system setting operation are read from the preset configuration table. These preset trigger conditions are pre-bound to specific setting operations and are explicitly described as TV device events (such as "USB storage device inserted", "HDMI signal access", "remote control low battery alarm") or TV system status thresholds (such as "built-in storage usage exceeds 90%", "number of consecutive playback stutters ≥ 3", "network latency exceeds 500ms").

[0042] The system continuously monitors the TV's operational data in real time, including hardware event logs, system resource metrics, peripheral device status, and network performance, and matches this data against the aforementioned preset trigger conditions. For example, when a USB device insertion event is detected and there is no playback task currently, the system determines that the "USB drive insertion" trigger condition has been met. Alternatively, when the storage monitoring module reports "available space is less than 1GB," the "insufficient storage space" condition is determined to be met. If the operational data meets a preset trigger condition, a voice prompt containing multiple candidate operation options is generated based on the candidate operation option list configured in the preset configuration table for the TV system settings. This candidate operation option list is a reasonable set of operations designed for the current context. For example, for the "peripheral content processing" setting operation triggered by "USB drive insertion," the candidate options may include "play videos on the USB drive," "browse files on the USB drive," and "upgrade system firmware." For the "storage management" setting operation triggered by "insufficient storage space," the candidate options may include "clear application cache," "uninstall infrequently used applications," and "move content to external storage." These options are translated into natural language queries, such as "USB drive detected. Do you want to play a video, browse files, or upgrade the system?", and the user is then asked to select an option via voice. This gives the TV system environmental awareness and proactive service capabilities, enabling it to provide accurate and timely operation guidance based on the actual device status before the user has even clearly expressed their intention, thus improving the level of interactive intelligence.

[0043] Reference Figure 2 The diagram shown is a functional module schematic of the voice control device 100 for television according to this application.

[0044] The voice control device 100 for television described in this application is installed in an electronic device. Depending on the functions implemented, the voice control device 100 for television includes a parsing module 110, a determining module 120, and a control module 130. These modules can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, and are stored in the memory of the electronic device.

[0045] In this embodiment, the functions of each module / unit are as follows: Parsing module 110: used to receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain intent categories and operation keywords; Determination module 120: is used to determine the user's setting operation on the TV system and the corresponding operation mode of the setting operation based on the intent category and operation keywords, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Control module 130: Used to perform corresponding voice interaction and control processing based on the settings of the television system, the operation mode, the current device status or preset trigger conditions.

[0046] The specific implementation of the voice control device for television described in this application is largely the same as the specific implementation of the voice control method for television described above, and will not be repeated here.

[0047] Reference Figure 3 The diagram shown is a schematic representation of a preferred embodiment of the electronic device of this application.

[0048] The electronic device includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. The memory 113 is used to store computer programs, such as voice control programs for televisions; In some embodiments, the processor 111 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 111 is typically used to control the overall operation of the electronic device, such as performing data interaction or communication-related control and processing. In this embodiment, the processor 111 is used to run program code stored in the memory 113 or process data.

[0049] The communication interface 112 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The communication interface 112 may also be used to establish a communication connection between the electronic device and other electronic devices.

[0050] The memory 113 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 113 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. In other embodiments, the memory 113 may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. of the electronic device. Of course, the memory 113 may include both internal storage units and external storage devices of the electronic device. In this embodiment, the memory 113 is typically used to store the operating system and various computer programs installed on the electronic device, such as program code for a voice control program for a television. In addition, the memory 113 may also be used to temporarily store various types of data that have been output or will be output.

[0051] Figure 3 Only an electronic device with components 111-114 is shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0052] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the voice control method for a television provided in any of the foregoing method embodiments, including: Receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Based on the intent category and operation keywords, determine the user's settings operation on the TV system and the corresponding operation mode of the settings operation, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions, corresponding voice interaction and control processing are performed.

[0053] For a detailed explanation of the above steps, please refer to the above. Figure 1 Description of a flowchart of an embodiment of a voice control method for television.

[0054] Furthermore, this application also proposes a computer-readable storage medium that is both non-volatile and volatile. This computer-readable storage medium is any one or any combination of several of the following: hard disk, multimedia card, SD card, flash memory card, SMC, read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, etc. The computer-readable storage medium includes a data storage area and a program storage area. The program storage area stores a voice control program for a television. When executed by a processor, the voice control program for the television performs the following operations: Receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Based on the intent category and operation keywords, determine the user's settings operation on the TV system and the corresponding operation mode of the settings operation, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions, corresponding voice interaction and control processing are performed.

[0055] The specific implementation of the computer-readable storage medium in this application is largely the same as the specific implementation of the voice control method for television described above, and will not be repeated here.

[0056] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0057] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware simulation platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0058] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A voice control method for a television, characterized in that, The method includes: Receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Based on the intent category and operation keywords, determine the user's settings operation on the TV system and the corresponding operation mode of the settings operation, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions, corresponding voice interaction and control processing are performed.

2. The voice control method for television as described in claim 1, characterized in that, The semantic parsing of the voice command to obtain the intent category and operation keywords includes: The voice command is subjected to speech recognition to generate a text statement; The text statement is input into a pre-trained intent recognition model to obtain the intent category that matches the text statement; Based on the intent category, the corresponding keyword extraction rules are invoked to extract operation keywords related to the settings operation of the TV system from the text statement.

3. The voice control method for television as described in claim 1, characterized in that, The step of determining the user's settings operation on the TV system and the corresponding operation mode based on the intent category and operation keywords includes: Based on the intent category, retrieve the set of candidate setting operations associated with the intent category from the preset instruction and operation mapping table; From the candidate setting operation set, filter out the setting operations that match the operation keywords; Based on the setting operation matched by the operation keyword, the operation mode corresponding to the setting operation is read from the attribute field configured in the instruction and operation mapping table.

4. The voice control method for television as described in claim 1, characterized in that, The process of performing corresponding voice interaction and control based on the settings of the television system, the operating mode, the current device status, or preset trigger conditions includes: If the operation mode is the immediate execution mode, the corresponding voice confirmation process is executed according to the risk level associated with the set operation, and the corresponding execution command is generated in combination with the current TV device status; If the operation mode is a page jump mode, the corresponding voice confirmation process is executed according to the risk level associated with the setting operation, an instruction to jump to the corresponding TV system settings page is generated, and full-duplex voice interaction is enabled on the page; If the operation mode is a condition-triggered mode, when a preset trigger condition associated with the setting operation is detected, a voice query containing multiple candidate operation options is generated, and the corresponding TV system setting operation is executed according to the user's voice selection result of the candidate operation options.

5. The voice control method for television as described in claim 4, characterized in that, The step of executing the corresponding voice confirmation process based on the risk level associated with the setting operation, and generating the corresponding execution command in combination with the current TV device status, includes: Query the risk level corresponding to the TV system settings operation in the preset configuration table, wherein the risk level includes high risk, medium risk or low risk; Based on the risk level, a voice confirmation strategy bound to the risk level is executed, wherein the voice confirmation strategy includes no confirmation required, single voice confirmation, or double voice confirmation. After meeting the requirements of the voice confirmation strategy, the feasibility of executing the setting operation of the TV system is verified based on the current TV device status, and a corresponding execution instruction is generated.

6. The voice control method for television as described in claim 4, characterized in that, The step of executing the corresponding voice confirmation process based on the risk level associated with the setting operation, and generating an instruction to jump to the corresponding TV system settings page, includes: Query the risk level corresponding to the TV system settings operation in the preset configuration table, wherein the risk level includes high risk, medium risk or low risk; Based on the risk level, a voice confirmation strategy bound to the risk level is executed, wherein the voice confirmation strategy includes no confirmation required, single voice confirmation, or double voice confirmation. After meeting the requirements of the voice confirmation strategy, an instruction to jump to the corresponding TV system settings page is generated based on the page path information recorded in the preset configuration table according to the settings operation of the TV system.

7. The voice control method for television as described in claim 4, characterized in that, When a preset trigger condition associated with the setting operation is detected, a voice query containing multiple candidate operation options is generated, including: Read the preset trigger conditions associated with the setting operation of the TV system from the preset configuration table. The preset trigger conditions include TV device events or TV system status thresholds. Real-time monitoring of the operating data of the television equipment, and matching the operating data with the preset trigger conditions for judgment; When the running data meets the preset triggering conditions, a voice query content containing multiple candidate operation options is generated according to the candidate operation option list configured in the preset configuration table based on the settings operation of the TV system.

8. A voice control device for a television, characterized in that, The device includes: Parsing module: used to receive user voice commands for the TV, and perform semantic parsing on the voice commands to obtain the intent category and operation keywords; Determination module: used to determine the user's setting operation on the TV system and the corresponding operation mode based on the intent category and operation keywords, wherein the operation mode includes immediate execution mode, page jump mode or condition trigger mode; Control module: used to perform corresponding voice interaction and control processing based on the settings of the TV system, the operation mode, the current device status or preset trigger conditions.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When executing a program stored in a memory, the processor implements the voice control method for a television as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice control method for a television as described in any one of claims 1 to 7.