A voice interaction control method and device for an 8K set-top box, equipment and medium
By generating voice input prompts on the initial page of the 8K set-top box, receiving and parsing user voice commands, and executing corresponding operations, the problem of complex interactive control in existing 8K set-top boxes is solved, realizing full-page voice interactive control and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SKYWORTH DIGITAL TECH CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-02
AI Technical Summary
The existing 8K set-top boxes have a lengthy interactive control process and complicated user operation, which makes it difficult to match the immersive experience requirements of ultra-high-definition picture quality. In particular, the user operation path is not intuitive in scenarios with multi-level menus and complex content libraries, resulting in frequent interruptions to the viewing experience.
By generating voice input prompts on the initial page of the 8K set-top box, receiving user voice control commands, parsing target interactive units and operation commands, and executing corresponding operations, the system achieves full-page voice interactive control.
It significantly improves the user interaction experience of 8K set-top boxes, reduces operational complexity and cognitive load, improves the accuracy and efficiency of interaction, and is compatible with different types of 8K set-top boxes and applications.
Smart Images

Figure CN122137995A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, and in particular to a voice interaction control method, device, equipment and medium for an 8K set-top box. Background Technology
[0002] Currently, ultra-high-definition video content is experiencing explosive growth. 8K set-top boxes are the core terminals that carry ultra-high-definition audio-visual experiences, and the level of intelligence of their interaction methods directly determines whether users can access massive amounts of ultra-high-definition video content without barriers.
[0003] In the existing interactive control of 8K set-top boxes, after the application is launched, users still need to rely on remote control buttons to complete operations such as channel switching, volume adjustment, and content retrieval to complete subsequent function interactions. The interaction chain is lengthy and has a high learning cost, which is difficult to match the immersive experience requirements brought by 8K picture quality. Especially in scenarios with multi-level menus and complex content libraries, users often interrupt the viewing experience due to the unintuitive operation path. Summary of the Invention
[0004] This invention provides a voice interaction control method, device, equipment, and medium for an 8K set-top box, to solve the problem of fragmented interaction control processes and poor user experience in existing 8K set-top boxes.
[0005] A voice interaction control method for an 8K set-top box includes the following steps: generating voice input prompts matching each initial interaction unit in the initial page based on an initial page displayed on the 8K set-top box; receiving voice control commands from a user based on the voice input prompts; parsing the voice control commands to obtain a target interaction unit and a corresponding target operation command; and controlling the target interaction unit to perform a corresponding target operation according to the target operation command to obtain a target operation result.
[0006] A voice interaction control device for an 8K set-top box includes: a prompt generation module for generating voice input prompts matching each initial interaction unit in the initial page displayed on the 8K set-top box; a voice receiving module for receiving voice control commands from a user based on the voice input prompts; a voice parsing module for parsing the voice control commands to obtain target interaction units and corresponding target operation commands; and an operation execution module for controlling the target interaction units to perform corresponding target operations according to the target operation commands to obtain target operation results.
[0007] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned voice interaction control method for an 8K set-top box.
[0008] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned voice interaction control method for an 8K set-top box.
[0009] In the aforementioned technical solution for the voice interaction control method, device, equipment, and medium of the 8K set-top box, the voice interaction control method includes the following steps: generating voice input prompts matching each initial interaction unit on the initial page displayed on the 8K set-top box; receiving the user's voice control commands based on the voice input prompts; parsing the voice control commands to obtain the target interaction unit and its corresponding target operation command; and controlling the target interaction unit to execute the corresponding target operation according to the target operation command to obtain the target operation result. This method, by analyzing the executable operations of each initial interaction unit on the initial page, can achieve full-page voice interaction control within the 8K set-top box application, significantly improving the user interaction experience of the 8K set-top box. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of a voice interaction control method for an 8K set-top box according to an embodiment of the present invention; Figure 2 This is a detailed flowchart of step S1 in the voice interaction control method of an 8K set-top box in one embodiment of the present invention; Figure 3 This is a detailed flowchart of step S3 in the voice interaction control method of an 8K set-top box in one embodiment of the present invention; Figure 4 This is a detailed flowchart of step S4 in the voice interaction control method of an 8K set-top box in one embodiment of the present invention; Figure 5 This is a detailed flowchart of the feedback of the target operation result after step S4 in the voice interaction control method of an 8K set-top box in one embodiment of the present invention; Figure 6 This is a schematic diagram of a voice interaction control device for an 8K set-top box in one embodiment of the present invention; Figure 7 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] In one embodiment, such as Figure 1 As shown, a voice interaction control method for an 8K set-top box is provided, including the following steps: Step S1: Based on the initial page displayed on the 8K set-top box, generate voice input prompts that match each initial interactive unit on the initial page.
[0014] It's important to note that the 8K set-top box is a terminal device that supports voice recognition and semantic understanding. Through its built-in voice recognition engine, it captures user voice commands in real time and controls the connected display system (such as a TV screen, monitor, or in-vehicle infotainment screen) to display 8K ultra-high-definition content. The initial page is the current interface displayed on the connected display system when the user initiates voice interaction. It shows multiple different types of initial interaction units, each located in a different position on the initial page, possessing independent operational attributes and executable operations that can be activated by voice. Voice input prompts are natural and concise voice guidance messages used to indicate the voice control content that the user can input, including the executable options on the current initial page, such as the functions of each initial interaction unit and the voice trigger keywords.
[0015] In this embodiment, when a user opens the application on the 8K set-top box via voice, the 8K set-top box automatically identifies the initial page currently displayed by the display system, performs structured parsing of the page information of the initial page, obtains the position coordinates, function tags and executable operations of each initial interaction unit, and then combines the semantic model to generate a voice input prompt that accurately matches it, ensuring that each voice input prompt has high recognizability and low ambiguity.
[0016] like Figure 2 As shown, specifically, step S1 includes the following sub-steps: Step S11: Perform information recognition on the initial page to obtain page information.
[0017] It should be noted that the page information includes the page type, layout structure, UI attributes of each initial interaction unit, and executable operations. The UI attributes cover icons, text, button states, and visual hierarchy, while the executable operations correspond to the underlying API calls and business logic of each initial interaction unit.
[0018] In this embodiment, the 8K set-top box is an Android TV smart set-top box with an Android 13 operating system. It has a built-in page information acquisition module and an AI analysis module. The page information acquisition module captures page information in real time, including the UI hierarchy, focus control status, and interactive element attributes of the initial page. The AI analysis module, based on a multimodal fusion algorithm, performs semantic annotation and intent mapping on the page information captured by the page information acquisition module, outputting a structured page description. The initial page is the movie list page of the video playback application. When the user opens the video playback application via voice and enters the movie list page, the page information is identified to obtain the location, type, and executable operations of each initial interactive unit, including the location, type, and executable operations of the "movie cover control," "previous / next page control," and "confirm playback control" within the movie list page.
[0019] Step S12: Parse the page information to obtain the interactive operation instructions that can be executed by each initial interactive unit.
[0020] It should be noted that interactive operation commands are voice-triggered behaviors supported by each initial interactive unit. They can be mapped from natural language to underlying function calls and are closely related to the location, type, and executable operations of each initial interactive unit.
[0021] In this embodiment, the page information acquired by the page information acquisition module is input into the AI analysis module. After receiving the page information, the AI analysis module performs structured parsing and identifies the interactive operation commands supported by each initial interaction unit, such as "the movie cover control supports left and right movement for selection," "the confirm playback control supports confirmation operation," and "the previous / next page control supports page switching operation." By converting the page information into interactive operation commands, the AI analysis module can accurately construct the semantic mapping relationship between voice commands and UI elements, generating voice input prompts with context awareness capabilities.
[0022] Step S13: Generate voice input prompts corresponding to each initial interaction unit based on each interactive operation instruction.
[0023] In this embodiment, the AI analysis module dynamically generates natural language voice input prompts based on the interactive operation instructions output in step S12, such as "Please use voice input to select a movie by saying 'move left' or 'move right,' input 'OK' to play the movie, and input 'previous page' or 'next page' to switch lists." Furthermore, the voice input prompts are broadcast in real-time through the TTS (Text-to-Speech) engine of the Android TV smart set-top box, and the text content of the voice input prompts is dynamically displayed in a semi-transparent overlay in the lower right corner of the screen, ensuring consistent guidance between the audiovisual and visual channels. The overlay supports focus-following highlighting; when the user triggers "move right" with their voice, the focus automatically jumps to the adjacent movie cover and plays a confirmation sound effect, forming a closed-loop feedback. This mechanism significantly improves the operational efficiency and immersion of visually impaired users, making voice interaction accurate, predictable, and error-tolerant.
[0024] In other embodiments, the page information also includes the identifier of the initial interaction unit, its boundary coordinates, interactive state, and preset supported operation types. This identifier and coordinate data are synchronized in real-time to the AI analysis module. Combined with the user's historical interaction preferences and contextual information, the module dynamically selects the three most frequently invoked initial interaction units, generating personalized voice input prompts such as "Say 'Play this' to select the cover, 'Turn the page' to browse more, or 'Confirm' to start watching." The entire process has a response latency of less than 300 milliseconds, ensuring a natural and smooth interaction. This mechanism significantly improves the accuracy of voice interaction and user operation efficiency, especially in multi-level nested interfaces, avoiding accidental triggering caused by ambiguous commands.
[0025] Step S2: Receive the user's voice control commands based on voice input prompts.
[0026] It's important to note that voice control commands are natural language instructions issued by the user based on voice input prompts and their own intentions. After receiving the voice input prompt, the Android TV smart set-top box outputs it to the corresponding user. The user can then select commands based on the voice input prompts and their needs, and output the corresponding voice control commands, which are then received in real time by the Android TV smart set-top box's corresponding voice recognition module.
[0027] In this embodiment, based on the floating prompts in the lower right corner of the screen and the TTS broadcast content, and combined with the actual user needs, the user inputs voice control commands such as "move left," "move right," "confirm," "previous page," or "next page" through the microphone. The voice recognition module of the Android TV smart set-top box receives the corresponding voice control commands in real time to perform subsequent voice interaction control of the Android TV smart set-top box.
[0028] Step S3: Parse the voice control command to obtain the target interaction unit and the corresponding target operation command.
[0029] It should be noted that the target interaction unit refers to the initial interaction unit on the initial page that has the highest semantic match with the voice control command. The target operation command refers to the specific action command executed on the target interaction unit to achieve the actual interactive response.
[0030] In this embodiment, the Android TV smart set-top box also has a built-in voice parsing module. The voice parsing module performs semantic understanding and intent recognition on the voice control commands to accurately locate the target interaction unit and the corresponding target operation command.
[0031] Specifically, such as Figure 3 As shown, step S3 includes the following sub-steps: Step S31: Extract semantic features from voice control commands.
[0032] It should be noted that semantic features refer to key information such as the intention of the action, the object being targeted, and the spatial relationship contained in the voice control commands.
[0033] In this embodiment, the voice parsing module employs a lightweight BERT model (Bidirectional Encoder Representations from Transformers) to extract semantic features in real time, such as directional intent in "move left" and "move right," confirmation action in "confirm," and hierarchical backtracking relationships in "previous page" and "next page." The BERT model has been locally fine-tuned to specifically adapt to the low-computing-power environment of set-top boxes, ensuring recognition accuracy while compressing inference latency to within 300 milliseconds to guarantee real-time interaction. It also supports offline operation, protecting user privacy and stability in weak network environments. It can accurately understand complex semantics, supports multi-dimensional feature extraction, and provides a more natural and precise control experience for voice interaction on 8K set-top boxes.
[0034] In other embodiments, a three-dimensional semantic vector space can be constructed by combining the current focus position with the topology of interface elements. This vector space is synchronously mapped to the coordinate system of the initial UI interaction unit, enabling millisecond-level target interaction unit positioning and operation command generation. For example, when a user issues a "move left" command, the system not only recognizes the directional semantics but also performs dynamic projection and optimal matching in the vector space based on the relative distance between the current focus coordinates and the adjacent initial interaction unit on the left, the interactivity weight, and the Z-axis hierarchy, ensuring zero deviation in command execution.
[0035] Step S32: Identify the target interaction unit from each initial interaction unit based on semantic features.
[0036] In this embodiment, the speech parsing module performs similarity matching between the extracted semantic features and the semantic tags of all initial interactive units on the initial page. A cosine similarity algorithm is used to calculate the matching score, and the initial interactive unit with the highest matching score is selected as the target interactive unit. The semantic tags are automatically generated by the UI framework during the rendering of the initial interactive unit, covering four dimensions: function type, hierarchical position, interaction state, and visual attributes. A dynamic weighting mechanism is introduced in the matching process, assigning higher priority to initial interactive units in the vicinity of the focus area, balancing semantic accuracy and operational intuition.
[0037] Furthermore, low-confidence matching results need to be filtered out using a preset matching threshold. Confirmation of the target interaction unit is only triggered when the highest matching score exceeds the threshold; otherwise, a fuzzy error-tolerance mechanism is triggered, combining contextual history and user operating habits for secondary calibration. For example, when a user issues the "previous page" command twice consecutively, the voice parsing module automatically increases the weight of historical context, including the parent container of the previous target interaction unit in the candidate set to avoid matching failure due to dynamic page refreshes. Simultaneously, a time decay factor is introduced, causing the influence of operation records within three minutes to decrease exponentially, ensuring both stable and timely responses.
[0038] Step S33: Based on semantic features and preset operation mapping rules, identify the target operation command of the target interaction unit.
[0039] It should be noted that the operation mapping rules are a hierarchical structured rule base that covers the precise mapping relationship between semantic features and operation instructions. This includes a basic semantic layer (e.g., "play" → "click"), a scene adaptation layer (e.g., "fast forward" → "seek + 10s" in a video page), and a user preference layer (e.g., automatically binding long-press triggers to high-frequency jump operations). The target operation instruction is the target operation that the target interaction unit needs to execute, generated in real-time by the voice parsing module based on the current interface state, the interaction capabilities of the initial interaction unit, and the user's historical preferences, ensuring a high degree of alignment between the instruction semantics and the functional semantics of the target interaction unit.
[0040] In this embodiment, the speech parsing module inputs semantic features into the rule base, matches and activates corresponding instruction branches layer by layer, and finally generates standardized target operation instructions. Specifically, when the semantic feature is "move right," the "move" action is first matched at the basic semantic layer, and then the scene adaptation layer determines that the current state is a movie selection state, triggering the focus horizontal scrolling logic to generate the target operation instruction "move focus right." When the semantic feature "mute" is recognized, it is first confirmed as a "click" action at the basic semantic layer, and the scene layer determines that the current state is a video playback state, triggering the volume control's on / off logic. Combined with the "automatic recovery after a single mute" setting in the user preference layer, the target operation instruction "volume reset" is preloaded after 30 seconds. The entire mapping process is completed within 80 milliseconds and supports multi-round semantic overlay. For example, "mute and skip the intro" simultaneously activates two operation instruction streams, which are then executed in an orderly manner after being scheduled by the timing arbitrator.
[0041] Furthermore, the target operation command undergoes dual verification by the command verification module: first, it verifies the compatibility between semantic features and the target interaction unit attributes, such as the "mute" command only affecting audio controls; second, it verifies the feasibility of the target operation command in the current UI state, excluding disabled, hidden, or invisible initial interaction units. If verification fails, a semantic conflict warning is immediately provided, and a semantic re-parsing mechanism is activated, reviewing the most recent three rounds of interaction logs and dynamically adjusting the rule base weights—if the user encounters two consecutive verification failures, it automatically degrades to the basic semantic layer for fault-tolerant operations, such as mapping the ambiguous command "turn down the volume" to a general volume slider dragging rather than a precise numerical setting. This verification logic effectively avoids invalid interactions, keeping the false trigger rate below 0.3%.
[0042] Step S4: Control the target interaction unit to execute the corresponding target operation according to the target operation instruction, and obtain the target operation result.
[0043] It should be noted that the target operation result is the interface feedback and state change after the target operation is executed, including explicit feedback such as page jump, initial interaction unit highlighting, animation response, and status prompts, as well as implicit feedback such as background data loading, cache update, and behavior log reporting.
[0044] In this embodiment, the verified target operation command is converted into a low-level control signal executable by the 8K set-top box. This signal is then encapsulated by the device driver layer into a command sequence conforming to the hardware protocol, simultaneously triggering the explicit feedback engine and the implicit response scheduler. Explicit feedback is dynamically adjusted based on the interaction intensity: lightweight operations (such as focus movement) trigger only millisecond-level visual micro-motions, while critical operations (such as mute) are superimposed with sound effects and status icon pulses. Implicit responses are executed in priority chunks—data loading uses a high-speed channel, and log reporting uses asynchronous compression to ensure zero latency in the main interaction. For example, after executing the "mute" command, the interface immediately displays a grayed-out volume icon and a flashing pulse, simultaneously playing a 0.2-second low-frequency prompt tone; the background completes audio stream pause, cache marker update, and encrypted log asynchronous reporting within 5 milliseconds.
[0045] Specifically, such as Figure 4 As shown, step S4 includes the following sub-steps: Step S41: Convert the target operation command into a low-level control signal that can be executed by the 8K set-top box.
[0046] In this embodiment, the underlying control signals are atomic instructions specific to the hardware platform, such as calling `AudioManager.setStreamMute()` on Android, triggering `AVAudioSession.setActive(false)` on iOS, and injecting `MediaElement.muted = true` on the Web. All signals are encapsulated by a unified abstraction layer to ensure consistent behavior across platforms. This abstraction layer also has a built-in protocol adapter that automatically identifies the device model and system version, dynamically selecting the optimal API path—for example, enabling the PrivacySandbox audio control interface on Android 12+ devices, while reverting to the traditional AudioManager solution on older versions. Context signatures are embedded during signal encapsulation, binding user ID, session ID, and operation timing stamps to ensure that instructions are traceable, auditable, and replayable. All underlying signals are transmitted encrypted through a secure channel, with the key dynamically generated by the device's Trusted Execution Environment (TEE) to eliminate the risk of man-in-the-middle tampering.
[0047] Step S42: Control the target interaction unit to complete the target operation according to the underlying control signal and obtain the target operation result.
[0048] In this embodiment, for the underlying control signal corresponding to the "move right" instruction, the driver module receives the underlying control signal and verifies the legality of the operation based on the context signature in the underlying control signal, and verifies the integrity of the key through the TEE. If the verification is successful, the instruction is decoded into a hardware register-level operation sequence, written to the frame buffer offset register of the GPU display controller, and then triggers pixel-level view redraw, controlling the selected item on the movie list page to move one position to the right, completing the corresponding target operation, and obtaining the target operation result. This target operation result includes the page jump state, cursor coordinate offset, animation completion timestamp, and hardware-level response latency data, all of which are signed by the TEE and encapsulated into a structured operation credential, and synchronized in real time to the user's multi-terminal session context engine to ensure millisecond-level consistency of operation states across devices.
[0049] Furthermore, when the "move right" command triggers focus scrolling, the underlying control signals synchronously drive the UI rendering pipeline to update the highlighted area and send the position offset to the focus manager to ensure that the visual displacement is strictly consistent with the logical focus; if scrolling to the end of the list, the preloading logic of the "load more" control is automatically activated, reflecting a deep understanding of the user's intent and proactive response.
[0050] Furthermore, during focus scrolling, the UI framework's focus management API is invoked to perform pixel-level smooth displacement. The displacement distance is dynamically calculated based on 75% of the width of the current initial interaction unit, and the focus highlight style and accessibility description text are updated simultaneously. If there is no valid initial interaction unit on the right, a micro-vibration feedback is played and a voice prompt "End of list reached" is given when the boundary is reached. This mechanism can also embed a real-time conflict detection module. When there is a logical contradiction between the target operation instruction and the current UI state (such as clicking in a disabled state), the semantic re-parsing process is automatically initiated. Combining the visual focus position and the interactivity flag of the initial interaction unit, the semantics of the instruction are dynamically corrected. After all operations are executed, a timestamped operation log is generated for subsequent user behavior modeling and rule base iteration optimization.
[0051] Furthermore, such as Figure 5 As shown, the voice interaction control method for this 8K set-top box also includes the following steps: Step S51: Based on the target operation result, determine whether the initial page should be redirected and obtain the page redirection result.
[0052] In this embodiment, if the initial page does not redirect, the target page is obtained, and the voice interaction control of the 8K set-top box ends. If the initial page redirects, the redirection result is obtained, and the process proceeds to step S52.
[0053] In some embodiments, if the initial page does not redirect, the DOM tree structure and semantic tags of the target page are automatically extracted, and a personalized focus path is generated by combining the user's historical interaction preferences. Simultaneously, a lightweight AR rendering engine is activated, overlaying a semi-transparent guide cursor at the screen edge to synchronously map the intent of the voice command with millisecond-level latency. This guide cursor has dynamic scaling and semantic highlighting capabilities, automatically adjusting the halo radius and transparency based on the hierarchy weight of the currently focused element, and semantically coloring it in real-time by associating keywords in the voice command. When the user's gaze lingers for more than 300 milliseconds, predictive focus pre-activation is triggered, preloading the next interactive resource. Furthermore, the system continuously monitors ambient light intensity and the user's pupil contraction rate, dynamically adjusting the AR cursor brightness and color temperature to match the day-night rhythm. When a multimodal command conflict is detected (e.g., the voice says "back" while the finger hovers over the "search" icon), an intent arbitration protocol is initiated according to the ISO 9241-210 standard for human factors engineering, using time-series confidence weighted fusion of visual, voice, and gesture signals to output a unique optimal action.
[0054] Step S52: Based on the page redirection result, complete the voice interaction control of the 8K set-top box.
[0055] In this embodiment, when the page jump is performed on the initial page, the page jump result is to obtain a new initial page, return to step S1, and repeat the corresponding operations of steps S1-S4 based on the new initial page displayed by the 8K set-top box to perform a new round of voice interaction control on the 8K set-top box until the user actively terminates the interaction or there is no effective response for three consecutive times, which triggers the sleep protocol and completes the voice interaction control of the 8K set-top box.
[0056] Furthermore, after the hibernation protocol is activated, the device enters a low-power listening state, maintaining only the edge wake-up capability of the microphone array and the lightweight inference of the visual focus prediction model; once a wake-up word matching the user's voiceprint characteristics with a confidence level of ≥92% is captured, or the gaze focus is detected to remain in the UI hot zone for more than 1.5 seconds, the full-function interactive state is restored in milliseconds, and the most recent session context snapshot is automatically loaded to ensure the continuity of operation and the consistency of intent.
[0057] It should be noted that, in addition to the video playback application scenarios mentioned above, this invention can also be applied to scenarios such as the settings page of an 8K set-top box and the product details page of a shopping application. For example, in the settings page, the AI analysis module can identify the executable operations of the "brightness adjustment control" and "volume adjustment control," and generate voice prompts such as "increase brightness" and "decrease volume." Users can complete the settings operation by using the corresponding voice control commands, achieving full-page voice interaction control within the 8K set-top box application without the need for a remote control.
[0058] In summary, the voice interaction control method of the present invention includes the following steps: generating voice input prompts matching each initial interaction unit in the initial page based on the initial page displayed on the 8K set-top box; receiving the user's voice control command based on the voice input prompts; parsing the voice control command to obtain the target interaction unit and the corresponding target operation command; controlling the target interaction unit to execute the corresponding target operation according to the target operation command to obtain the target operation result. This method analyzes the executable operations of each initial interaction unit in the initial page, dynamically generating accurate and operable voice input prompts for each initial interaction unit, significantly improving command coverage and semantic matching accuracy; furthermore, through voice input prompts, it guides the user to input the correct voice control command, reducing the user's cognitive load and operation trial-and-error rate; simultaneously, by parsing the voice control command and accurately mapping it to the target interaction unit and the target operation command, it achieves end-to-end semantic closed-loop control, ensuring that the command execution result is fed back to the user in real time, and simultaneously triggering UI status updates and voice confirmation broadcasts, forming a four-step closed loop of "input—parsing—execution—feedback". Therefore, this method can realize full-page voice interaction control within 8K set-top box applications, adapt to different types of 8K set-top boxes and applications, and greatly improve the user interaction experience of 8K set-top boxes.
[0059] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0060] In one embodiment, a voice interaction control device for an 8K set-top box is provided, which corresponds one-to-one with the voice interaction control method for the 8K set-top box in the above embodiments. For example... Figure 6 As shown, the voice interaction control device of this 8K set-top box includes a prompt generation module 101, a voice receiving module 102, a voice parsing module 103, and an operation execution module 104. Detailed descriptions of each functional module are as follows: The prompt generation module 101 is used to generate voice input prompts matching each initial interactive unit in the initial page based on the initial page displayed on the 8K set-top box.
[0061] The voice receiving module 102 is used to receive the user's voice control commands based on the voice input prompts.
[0062] The voice parsing module 103 is used to parse the voice control commands to obtain the target interaction unit and the corresponding target operation command.
[0063] The operation execution module 104 is used to control the target interaction unit to perform the corresponding target operation according to the target operation instruction, and obtain the target operation result.
[0064] Specific limitations regarding the voice interaction control device for 8K set-top boxes can be found in the above description of the voice interaction control method for 8K set-top boxes, and will not be repeated here. Each module in the aforementioned voice interaction control device for 8K set-top boxes can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0065] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice interaction control method for an 8K set-top box.
[0066] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the voice interaction control method for the 8K set-top box described in the above embodiment, for example... Figure 1 S1-S4, as shown, will not be described again here to avoid repetition. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the voice interaction control device for the 8K set-top box, for example... Figure 6 The functions of the prompt generation module 101, voice receiving module 102, voice parsing module 103, and operation execution module 104 shown are not described again here to avoid repetition.
[0067] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the voice interaction control method of the 8K set-top box described in the above embodiment, for example... Figure 1 S1-S4, as shown, will not be described again here to avoid repetition. Alternatively, when this computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the voice interaction control device for the 8K set-top box, for example... Figure 6The functions of the prompt generation module 101, voice receiving module 102, voice parsing module 103, and operation execution module 104 shown are not described again here to avoid repetition. The computer-readable storage medium can be non-volatile or volatile.
[0068] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0069] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0070] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A voice interaction control method for an 8K set-top box, characterized in that, Including the following steps: Based on the initial page displayed on the 8K set-top box, voice input prompts matching each initial interactive unit in the initial page are generated; Based on the voice input prompts, receive the user's voice control commands; The voice control command is parsed to obtain the target interaction unit and the corresponding target operation command; The target interaction unit is controlled to perform the corresponding target operation according to the target operation instruction, and the target operation result is obtained.
2. The voice interaction control method as described in claim 1, characterized in that, The initial page displayed on the 8K set-top box generates voice input prompts matching each initial interactive unit on the initial page, including: The initial page is subjected to information recognition to obtain page information; The page information is parsed to obtain the interactive operation instructions that can be executed by each of the initial interaction units; Based on each of the interactive operation instructions, a voice input prompt corresponding to each of the initial interactive units is generated.
3. The voice interaction control method as described in claim 1, characterized in that, The step of parsing the voice control command to obtain the target interaction unit and the corresponding target operation command includes: Extract semantic features from the voice control commands; Based on the semantic features, the target interaction unit is identified from each of the initial interaction units; Based on the semantic features and the preset operation mapping rules, the target operation command of the target interaction unit is identified.
4. The voice interaction control method as described in claim 1, characterized in that, The step of controlling the target interaction unit to execute the corresponding target operation according to the target operation instruction and obtaining the target operation result includes: The target operation command is converted into a low-level control signal that can be executed by the 8K set-top box; The target interaction unit is controlled to complete the target operation according to the underlying control signal, and the target operation result is obtained.
5. The voice interaction control method as described in claim 1, characterized in that, The voice interaction control method further includes: Based on the target operation result, determine whether the initial page should be redirected, and obtain the page redirection result; Based on the page redirection result, the voice interaction control of the 8K set-top box is completed.
6. The voice interaction control method as described in claim 5, characterized in that, The step of completing the voice interaction control of the 8K set-top box based on the page redirection result includes: When a page jump occurs on the initial page, a new initial page is obtained, and the initial page displayed on the 8K set-top box is returned to, generating voice input prompts that match the initial interactive units on the initial page.
7. The voice interaction control method as described in claim 1, characterized in that, The step of receiving the user's voice control commands based on the voice input prompt includes: Based on the voice input prompts and user needs, guide the user to input the voice control commands and receive the voice control commands.
8. A voice interaction control device for an 8K set-top box, characterized in that, include: The prompt generation module is used to generate voice input prompts that match each initial interactive unit in the initial page based on the initial page displayed on the 8K set-top box. The voice receiving module is used to receive the user's voice control commands based on the voice input prompts; The voice parsing module is used to parse the voice control commands to obtain the target interaction unit and the corresponding target operation command; The operation execution module is used to control the target interaction unit to perform the corresponding target operation according to the target operation instruction, and obtain the target operation result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the voice interaction control method for the 8K set-top box as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice interaction control method for the 8K set-top box as described in any one of claims 1 to 7.