Voice interaction method, system and device, electronic equipment and medium

By acquiring the text information and component attributes of voice input operations, the target interactive component is identified and the interactive command is executed, solving the problem of the inflexible interaction of existing voice assistants, realizing accurate interaction between users and device interface components, and improving the user experience.

CN121963727APending Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-01-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing voice assistants cannot flexibly interact with various visible components in the device interface via voice, which affects the user experience.

Method used

By acquiring the voice text information and component attribute information corresponding to the voice input operation, the component description is extracted, the target interactive component is determined, and the corresponding component interaction instructions are executed.

Benefits of technology

It enables flexible and precise interaction between users and device interface components, improving the convenience of voice interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963727A_ABST
    Figure CN121963727A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method and device, electronic equipment and a medium, and the method comprises the steps: obtaining voice text information corresponding to a voice input operation and component attribute information of at least one component in a current interface in response to the voice input operation for the current interface; performing component description extraction on the voice text information to obtain component feature description information and interaction type information; performing component positioning based on the component feature description information and the component attribute information of the at least one component, and determining a target interaction component in the at least one component; and for the target interaction component, executing a component interaction instruction corresponding to the interaction type information. By utilizing the scheme provided by the invention, the target component which the user wants to interact in the current interface can be accurately positioned in combination with the component feature description information extracted from the voice text information input by the user through voice, so that a simulation interaction effect between the user and the target component is generated, and the voice interaction experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and in particular to a voice interaction method, system, device, electronic device and medium. Background Technology

[0002] A voice assistant is a device that uses technologies such as speech recognition and natural language processing to enable users to interact with the device using natural language.

[0003] However, in the relevant existing technologies, most voice assistants can only perform some predefined fixed functions. When the user triggers a voice input operation, the system matches the voice text with functional keywords to realize the call of fixed functions, such as opening settings or increasing the volume. Users cannot flexibly interact with various visible components in the device interface by voice, which affects the user's voice interaction experience. Summary of the Invention

[0004] This application provides a voice interaction method, device, equipment, storage medium, and computer program product, which allows users to interact flexibly and accurately with various components on the current interface simply by voice input. This breaks the limitation of traditional voice assistants that can only perform fixed operations, and improves the convenience of component interaction and the user's voice interaction experience.

[0005] According to one aspect of the embodiments of this application, a voice interaction method is provided, the method comprising: In response to a voice input operation on the current interface, obtain the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. Component descriptions are extracted from the voice and text information to obtain component feature description information and interaction type information; Based on the component feature description information and the component attribute information of each of the at least one component, the component is located to determine the target interactive component among the at least one component; For the target interactive component, execute the component interaction instruction corresponding to the interaction type information.

[0006] According to one aspect of the embodiments of this application, a voice interaction device is provided, the device comprising: The first information acquisition module is used to respond to a voice input operation on the current interface and acquire the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. The component description extraction module is used to extract component descriptions from the voice text information to obtain component feature description information and interaction type information. A component localization module is used to locate components based on the component feature description information and the component attribute information of each of the at least one component, and to determine the target interactive component among the at least one component. The interaction instruction execution module is used to execute component interaction instructions corresponding to the interaction type information for the target interaction component.

[0007] According to one aspect of the embodiments of this application, an electronic device is provided, including: a processor; Memory used to store computer programs executable by the processor; The processor is configured to execute the computer program to implement any of the above-described voice interaction methods.

[0008] According to one aspect of the present application, a computer-readable storage medium is provided, which, when a computer program in the storage medium is executed by a processor of an electronic device, enables the electronic device to perform any of the above-described voice interaction methods.

[0009] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, and a processor reading from the computer-readable storage medium and executing the computer program to implement any of the above-described voice interaction methods.

[0010] The voice interaction method, apparatus, device, storage medium, and computer program product provided in this application have the following technical effects: In a voice interaction scenario between a user and a terminal interface, this application, in response to a voice input operation on the current interface, acquires the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. It then extracts component descriptions from the voice text information to obtain component feature description information and interaction type information. Based on the component feature description information and the component attribute information of at least one component, it performs component localization to determine the target interactive component among at least one component. By extracting the component feature description information from the user's voice input voice text information and based on the component feature description information and the component attribute information of the current interface, it accurately locates the target component that the user wishes to interact with, thereby generating a simulated interaction effect between the user and the target component. This breaks the limitation of traditional voice assistants that can only perform fixed operations. Users can interact flexibly and accurately with various components in the current interface simply by voice input, improving the convenience of component interaction and the user's voice interaction experience. Attached Figure Description

[0011] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the application environment of a voice interaction method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a voice interaction method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a process of inputting component feature description information and component attribute information of at least one component into a large language model for component localization to determine the target interactive component in at least one component, according to an embodiment of this application. Figure 4 This is a flowchart illustrating another voice interaction method provided in an embodiment of this application; Figure 5 This is a flowchart illustrating another voice interaction method provided in an embodiment of this application; Figure 6 This is a flowchart illustrating another voice interaction method provided in an embodiment of this application; Figure 7 This is a flowchart illustrating a voice interaction method for a television terminal provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application; Figure 9 This is a block diagram of an electronic device for voice interaction provided in an embodiment of this application; Figure 10 This is a block diagram of another electronic device for voice interaction provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0016] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0017] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0018] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text translation, semantic understanding, machine translation, question answering, and knowledge graphs.

[0019] Large Language Model (LLM): Also known as a large model, natural language model, or large-scale language model, it refers to a natural language processing model with a large number of parameters and training data. The training process of a large language model typically employs unsupervised learning, that is, training the model using a large-scale text corpus to learn the probability distribution and rules of language. During training, the large language model usually uses a language model as the objective function, optimizing the model parameters by maximizing the predicted probability of the next word.

[0020] User interface (UI): refers to the interface through which users interact with the system, including visible elements such as buttons, menus, and pop-ups.

[0021] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment for a voice interaction method provided in an embodiment of this application. The application environment may include at least a terminal device 101 and a server 102.

[0022] Terminal device 101 and server 102 are connected via a wireless or wired network. Terminal device 101 responds to a user-triggered voice input operation on the current interface, acquires the corresponding voice text information and component attribute information of at least one component on the current interface, and sends the voice text information and component attribute information of at least one component to server 102. Server 102 extracts component descriptions from the voice text information to obtain component feature description information and interaction type information. Based on the component feature description information and component attribute information of at least one component, server 102 locates the component, identifies the target interactive component among at least one component, and sends the target interactive component and interaction type information to terminal device 101, so that terminal device 101 executes the component interaction command corresponding to the interaction type information for the target interactive component. Optionally, the terminal device 101 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, or in-vehicle terminal, but is not limited to these.

[0023] Optionally, the server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Optionally, it is a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0024] In some embodiments, a voice interaction application provided by a server 102 is installed on the terminal device 101. The terminal device 101 can perform functions such as voice interaction requests through this voice interaction application. Optionally, the voice interaction application is an application in the operating system of the terminal device 101, or an application provided by a third party. The user can trigger a voice input operation on the application interface of the voice interaction application through the terminal device 101. The terminal device 101 obtains the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. The application sends the voice text information and the component attribute information of at least one component to the server 102. The server 102 extracts component descriptions from the voice text information to obtain component feature description information and interaction type information. Based on the component feature description information and the component attribute information of at least one component, the server locates the component, determines the target interactive component among the at least one component, and sends the target interactive component and interaction type information to the terminal device 101 so that the terminal device 101 executes the component interaction instruction corresponding to the interaction type information for the target interactive component.

[0025] Those skilled in the art should understand that the terminal device 101 and server 102 described above are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0026] The following describes the application scenarios of the voice interaction method provided in the embodiments of this application. The embodiments of this application provide a voice interaction method, device, electronic device, computer-readable storage medium, and computer program product, which can be applied to various scenarios. Examples are given below.

[0027] In some embodiments, the above-described voice interaction method can be applied to voice interaction scenarios in television terminals. In response to a user's voice input operation on the television interface, the method obtains the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the television interface. The method extracts component descriptions from the voice text information to obtain component feature description information and interaction type information. Based on the component feature description information and the component attribute information of at least one component, the method locates the component and determines the target interactive component in at least one component. Then, for the target interactive component, the method executes the component interaction command corresponding to the interaction type information, thereby generating a simulated interaction effect between the user and the television components and improving the user's television interaction experience.

[0028] In some embodiments, the above-described voice interaction method can be applied to voice interaction scenarios of in-vehicle terminals. In response to a user's voice input operation on the vehicle interface, the method obtains the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the vehicle interface. The method extracts component descriptions from the voice text information to obtain component feature description information and interaction type information. Based on the component feature description information and the component attribute information of at least one component, the method locates the component and determines the target interactive component in at least one component. Then, for the target interactive component, the method executes the component interaction command corresponding to the interaction type information, thereby generating a simulated interaction effect between the user and the vehicle components and improving the user's vehicle interaction experience.

[0029] It should be noted that, in addition to the above-mentioned application scenarios, the voice interaction method provided in this application can also be widely applied to other scenarios to meet the diverse needs of different user groups, and this application does not impose any restrictions.

[0030] Figure 2 This is a flowchart illustrating a voice interaction method provided in an embodiment of this application. The method is executed by a computer device, which can be a terminal device or a server. Figure 2 As shown, the method may include: S201, in response to a voice input operation on the current interface, obtain the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface.

[0031] In one specific embodiment, the current interface can be the interface displayed on the screen of the terminal device. Specifically, the current interface can be the current application interface of the target application running on the terminal device. For example, the target application may include, but is not limited to: video application, social application, and reading application. The terminal device may include, but is not limited to: smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, and vehicle terminal, etc.

[0032] In one specific embodiment, voice input can be triggered by the user operating the terminal device. In some application scenarios, the user can activate the voice application on the terminal device before triggering the voice input. For example, the user can activate the voice application on the terminal device via a voice interaction button on a remote control wirelessly connected to the terminal device.

[0033] In one specific embodiment, the speech-text information is used to express the component interaction intent of the user (speaker) triggering the speech input operation. The speech-text information can be the text information obtained after speech recognition of the speech data collected for the speech input operation. Specifically, the speech data can be preprocessed to improve the robustness of subsequent speech recognition. This preprocessing can include, but is not limited to: sampling and frame segmentation, denoising and echo cancellation, endpoint detection, and multi-microphone beamforming. Illustratively, the sampling rate can be 16kHz, the speech format can be 16-bit PCM, the speech frame length can be 25 ms, and the frame shift can be 10 ms; denoising and echo cancellation can employ time-domain / frequency-domain filtering (e.g., Wiener filtering, spectral subtraction) or deep learning models (e.g., waveform-level deep neural networks or convolutional recurrent networks); endpoint detection can employ traditional algorithms based on energy and spectral features or endpoint detection based on deep neural networks, which can reduce latency and invalid transmissions.

[0034] In a specific embodiment, features can be extracted from the voice acquisition data to obtain voice feature data, and then voice recognition can be performed based on the voice feature data to obtain voice-text information. For example, the voice acquisition data can be represented as waveform data, and the voice feature data can use acoustic feature data such as Mel-frequency cepstral coefficients. Optionally, feature enhancement can be performed on the voice feature data to improve the robustness and performance of voice recognition in complex scenarios. For example, feature enhancement can employ mean-variance normalization, acoustic context splicing, and adaptive recognition based on speaker identity feature vectors.

[0035] In one specific embodiment, an ASR (Automatic Speech Recognition) model can be used to perform speech recognition on speech feature data to obtain speech text information. Optionally, the ASR model can be a traditional hybrid model or an end-to-end model. Illustratively, a traditional hybrid model can consist of an acoustic model (e.g., deep neural network, convolutional neural network, LSTM), a hidden Markov model, a decoder, and an external N-gram language model. An end-to-end model can be a CTC (Connectionist Temporal Classification) model, an RNN-T (RNN Transduce) model, an encoder-decoder model based on attention mechanisms or Transformers. In an optional embodiment, for the traditional hybrid model, the external language model can be shallowly fused. For example, during the decoding stage (e.g., beam search), the score of the external language model and the score of the ASR model can be linearly weighted to improve language coherence. Alternatively, the external language model can be deeply fused. For example, the internal representations (e.g., hidden states) of the external language model can be deeply fused in the early or middle stages of ASR model decoding.

[0036] In an optional embodiment, the speech text information can also be post-processed to further improve the accuracy of speech recognition. Illustratively, the post-processing can include: punctuation restoration and case prediction, number / unit normalization (e.g., normalizing "one hundred and twenty" to "120"), colloquialization correction (removal of filler words), and calibration of the confidence of the speech text information (mapping model probabilities to confidence scores).

[0037] In an optional embodiment, to ensure low latency in speech recognition, a streaming model can be used as the ASR model or an edge-cloud collaborative recognition method can be adopted. Speech activity detection and preliminary ASR in key scenarios can be performed on the device side, while complex sentences or large models can be processed in the cloud. In addition, the ASR model can be compressed and deployed to the edge device of the mobile terminal (e.g., TV box, remote control). For illustration, the model compression here can adopt methods such as quantization, distillation, and pruning.

[0038] In a specific embodiment, each component in the current interface can be a visual interface component (UI component) with a specific function. The component attribute information of each component is used to define the component characteristics of each component. Optionally, the component attribute information of each component may include, but is not limited to: component identifier (id), component display text (text, label), component type information (type: button, list, checkbox, input, slider, etc.), component position information (boundary box: horizontal coordinate x, vertical coordinate y, width w, height h), component visibility information (visibility), component interactivity information (enabled), component hierarchy information (z-index), component hierarchy relationship information (parent_id, children_ids), component auxiliary description information (aria), component recognition text (OCR), component interaction state information (state: focus, selected, checked, hover), and other sub-attribute information.

[0039] In one specific embodiment, at least one component in the current interface can be a visible and interactive interface component among at least one interface component of the current interface. Specifically, the component attribute information of each of the at least one interface component can be obtained in the following way: 1) In response to a voice input operation on the current screen, obtain the root view tree of the current screen.

[0040] 2) Traverse at least one component node in the root view tree to obtain the node attribute information of at least one component node.

[0041] Specifically, the current root view tree is obtained through the getRootView method of the current Activity instance; the layout manager is obtained through the rootView.getLayoutManager method; all component nodes of the current root view tree are obtained through the layoutManager.getItemCount method; and then all component nodes are traversed in turn through a for loop to obtain the node attribute information of each component node.

[0042] 3) Use at least one component node of the root view tree as at least one interface component in the current interface; use the node attribute information of each of the at least one component node as the component attribute information of each of the at least one interface component.

[0043] In an optional embodiment, the above method may further include: Extract the component attribute information of at least one UI component in the current interface; At least one UI component is defined as a UI component whose component attribute information indicates that it is in a visible and interactive state.

[0044] Specifically, component attribute information may include component visibility information and component interactivity information. From at least one UI component, UI components that are visible (visibility=true) and interactive (enabled=true) are selected as at least one component.

[0045] As can be seen from the above embodiments, in the scenario of component interaction based on user voice, the component that the user expects to interact with should be a component that the user can currently see and interact with. By filtering out "visible and interactive" components from at least one interface component, the subsequent component location range is narrowed, the efficiency of component location is improved, and thus the efficiency of voice interaction is improved.

[0046] S202, extract component descriptions from the voice and text information to obtain component feature description information and interaction type information.

[0047] Specifically, the component feature description information describes the component features that the user expects to interact with. This is illustrative, and the component features may include, but are not limited to, spatial location, component content, component function, component identifier, and other feature dimensions. The interaction type information indicates the type of interaction operation the user expects to perform. This is illustrative, and the interaction operation types may include, but are not limited to, click, select, scroll, swipe, input, return, open, and other operation types.

[0048] In an optional embodiment, before extracting component descriptions from the speech text information, speech recognition error correction can be performed on the speech text information to improve the accuracy of the expression of the text information. Specifically, speech recognition error correction can adopt an error correction method based on edit distance or an error correction method based on a sequence-to-sequence model to correct word errors introduced during the automatic speech recognition process.

[0049] In a specific embodiment, the above-described component description extraction of speech text information to obtain component feature description information and interaction type information may include: S2021 (not shown in the figure) inputs voice and text information into the interaction intent classification model to classify the interaction intent and obtain interaction type information.

[0050] Specifically, the interaction intent classification model can be any existing artificial intelligence model with interaction type classification capabilities; this application does not impose any particular limitations on it. Illustratively, the interaction intent classification model can employ a large language model, or a joint model of text feature extraction and a classifier.

[0051] In one specific embodiment, the interaction intent classification model includes a large language model. The process of inputting voice-text information into the interaction intent classification model for interaction intent classification to obtain interaction type information may include: inputting voice-text information and interaction classification prompts into the large language model, classifying the voice-text information for interaction intent, and obtaining interaction type information. The interaction classification prompts are used to indicate the interaction type corresponding to the voice-text information from a variety of preset interaction types. Illustratively, the various preset interaction types can be pre-set based on the component interaction requirements in actual applications. For example, the various preset interaction types can be set as: click, select, scroll, swipe, input, return, and open.

[0052] In one specific embodiment, the interaction intent classification model includes a text feature extraction model and a classifier. The process of inputting speech-text information into the interaction intent classification model to classify the interaction intent and obtain interaction type information can include: inputting the speech-text information into the text feature extraction model to extract text features, obtaining text feature information; and then inputting the text feature information into the classifier to classify the interaction intent and obtain interaction type information. Illustratively, the text feature extraction model can use the BERT model, and the classifier can use a Softmax classifier or a multi-label classifier.

[0053] S2022 (not shown in the figure) inputs the voice and text information into the slot extraction model to extract component descriptions and obtain component feature description information.

[0054] Specifically, the slot extraction model can be any existing artificial intelligence model capable of extracting component description slots; this application does not impose any particular limitations on it. Illustratively, the slot extraction model can employ a large language model, a sequence labeling model (e.g., a BIO labeling model, a conditional random field model), or a pointer network.

[0055] In a specific embodiment, the component feature description information can be structured component feature representation information. This information may include feature description information for at least one component feature dimension. Each component feature dimension's feature description information can consist of a set of slot-slot pairs, where the slot represents the component feature dimension and the slot value represents the feature description information under that dimension. Illustratively, the at least one component feature dimension may include, but is not limited to, spatial location dimensions (e.g., row dimension, column dimension, hierarchy dimension), type dimensions (e.g., function type dimension, content type dimension), and identifier dimensions.

[0056] In an optional embodiment, the interaction intent classification model and the slot extraction model can be jointly trained to improve the overall text semantic understanding effect.

[0057] In an optional embodiment, the above-mentioned extraction of component descriptions from speech text information to obtain component feature description information and interaction type information may further include: inputting speech text information into a large language model to extract component descriptions to obtain interaction type information and component feature description information.

[0058] Specifically, the large language model can classify the interaction type of speech and text information to obtain interaction type information. Then, combined with the interaction type information, it can perform part-of-speech tagging and dependency parsing on multiple text segments in the speech and text information to obtain grammatical relationship information between multiple text segments. Based on the grammatical relationship information, it can extract component descriptions from multiple text segments to obtain component feature description information.

[0059] In an optional embodiment, if pronouns exist in the multiple text segments, the large language model can also combine the speech context information prior to triggering the current speech input operation to perform coreference resolution on the pronouns.

[0060] For example, taking the voice text information "click the third one in the second row on the left" as an example, the large language model can classify the interaction type of the voice text information and obtain the interaction type information as "click". Then, it can extract component descriptions from multiple text segments "click", "left", "second row" and "third" in the voice text information to obtain the text segment "left" corresponding to the region dimension, the text segment "second row" corresponding to the row dimension and the text segment "third" corresponding to the column dimension. The text segment "left" corresponding to the region dimension is structured to obtain the feature description information of the region dimension (region = left), the text segment "second row" corresponding to the row dimension is structured to obtain the feature description information of the row dimension (row = 2), and the text segment "third" corresponding to the column dimension is structured to obtain the feature description information of the column dimension (column = 3).

[0061] S203, based on component feature description information and component attribute information of at least one component, perform component localization and determine the target interactive component in at least one component.

[0062] In a specific embodiment, the above-described component localization based on component feature description information and component attribute information of at least one component, determining the target interactive component among at least one component, may include: The component feature description information and the component attribute information of at least one component are input into the large language model to locate the component and determine the target interactive component.

[0063] Specifically, the large language model can locate components based on component feature description information and the component attribute information of at least one component, and determine the target interactive component in at least one component.

[0064] In one specific embodiment, the component feature description information includes: feature description information for at least one component feature dimension, such as... Figure 3 As shown, the above-mentioned component feature description information and component attribute information of at least one component are input into the large language model for component localization, and the target interactive component in at least one component is determined to include: Input the feature description information of at least one component feature dimension and the component attribute information of at least one component into the large language model, and the large language model performs the following operations: S301, based on the positioning priority of at least one component feature dimension, determine the current component feature dimension among at least one component feature dimensions; and select at least one component as the current matching component.

[0065] Specifically, the positioning priority of the at least one component feature dimension can be used to determine the positioning order of the at least one component feature dimension in the component positioning process. Optionally, the positioning priority of the at least one component feature dimension can be determined by combining the order of the text segments corresponding to each of the at least one component feature dimension in the speech text information, or by combining the grammatical relationship information between the text segments corresponding to each of the at least one component feature dimension.

[0066] S302, determine candidate components from the currently matched components whose component attribute information matches the feature description information of the current component feature dimension.

[0067] In a specific embodiment, the candidate components whose component attribute information matches the feature description information of the current component feature dimension are determined from the current matching components, including: S3021 (not shown in the figure), determine the sub-attribute information in the component attribute information of the currently matched component that corresponds to the feature dimension of the current component.

[0068] Specifically, component attribute information may include, but is not limited to: component identifier (id), component display text (text, label), component type information (type: button, list, checkbox, input, slider, etc.), component position information (boundary box: horizontal coordinate x, vertical coordinate y, width w, height h), component visibility information (visibility), component interactivity information (enabled), component hierarchy information (z-index), component hierarchy relationship information (parent_id, children_ids), component auxiliary description information (aria), component recognition text (OCR), component interaction state information (state: focus, selected, checked, hover), etc. If a certain sub-attribute information corresponds to the current component feature dimension, it can be said that the sub-attribute information can be used as the basis for the component to take a value under the current component feature dimension.

[0069] To illustrate, taking the current component feature dimension as a row dimension, the sub-attribute information corresponding to the current component feature dimension can be the vertical coordinate y, which is the ordinate of the top-left corner of the component's bounding box; taking the current component feature dimension as a column dimension, the sub-attribute information corresponding to the current component feature dimension can be the horizontal coordinate x, which is the abscissa of the top-left corner of the component's bounding box; taking the current component feature dimension as a content type dimension, the sub-attribute information corresponding to the current component feature dimension can be at least one of the following: component display text, component identification text, or component auxiliary description information.

[0070] S3022 (not shown in the figure): Determine the sub-attribute filtering conditions based on the feature description information of the current component feature dimension.

[0071] Specifically, the sub-attribute filtering conditions can be the filtering conditions corresponding to the sub-attribute information. For example, taking the current component feature dimension as the row dimension, the sub-attribute filtering condition can be the filtering range of the vertical coordinate y; taking the current component feature dimension as the column dimension, the sub-attribute filtering condition can be the filtering range of the horizontal coordinate x; taking the current component feature dimension as the content type dimension, the sub-attribute filtering condition can be that there is a semantic relationship between the sub-attribute information and the feature description information of the content type dimension. Here, semantic relationship can refer to the semantic similarity between the sub-attribute information and the feature description information of the content type dimension being greater than a preset similarity threshold.

[0072] S3023 (not shown in the figure): Determine candidate components from the currently matched components whose sub-attribute information meets the sub-attribute filtering conditions.

[0073] S303, based on the positioning priority of at least one component feature dimension, update the current component feature dimension; take the candidate component as the current matching component, and repeatedly execute the component positioning process from the current matching component to determine the candidate component whose component attribute information matches the feature description information of the current component feature dimension, until the preset positioning end condition is reached.

[0074] Specifically, the preset positioning end condition can be preset according to the component positioning accuracy requirements in actual applications. For example, the preset positioning end condition can be that the number of current candidate components is 1, or that at least one component feature dimension has been visited.

[0075] S304. The candidate component that reaches the preset positioning end condition is used as the target interactive component.

[0076] For example, the component feature description information includes: feature description information of the region dimension (region = left side), feature description information of the row dimension (row = 2), and feature description information of the column dimension (column = 3). First, the region dimension is used as the current component feature dimension, and the sub-attribute information corresponding to the region dimension is determined as the horizontal coordinate x. Based on the feature description information of the region dimension (region = left side), a first filtering range of the horizontal coordinate x is determined, and at least one component of the current interface whose horizontal coordinate x satisfies the first filtering range is selected as candidate components (i.e., components in the left region). Next, the row dimension is used as the current component feature dimension, and the sub-attribute information corresponding to the row dimension is determined as the vertical coordinate y. The components in the left region are sorted according to the vertical coordinate y, the row boundary of each row of components in the left region is identified, and the feature description information of the row dimension is used as the basis for the selection. Information (row = 2): Determine the second filtering range for vertical coordinate y. Select components in the left area whose vertical coordinate y satisfies the second filtering range as candidate components (i.e., components in the second row). Then, use the column dimension as the current component feature dimension and determine the sub-attribute information corresponding to the column dimension as the horizontal coordinate x. Sort the components in the second row according to the horizontal coordinate x to obtain the column sequence of components. Based on the feature description information of the column dimension (column = 3), determine the third filtering range for horizontal coordinate x. Select components in the second row whose horizontal coordinate x satisfies the third filtering range as candidate components (i.e., components in the third column). At this point, all three component feature dimensions contained in the component feature description information have been visited, reaching the preset positioning end condition. Select the current candidate component as the target interactive component and output the component attribute information of the target interactive component.

[0077] As can be seen from the above embodiments, by inputting the feature description information of at least one component feature dimension and the component attribute information of at least one component into the large language model, the positioning order of at least one component feature dimension is determined based on the positioning priority of at least one component feature dimension, and the components are screened sequentially according to the positioning order, combined with the feature description information of each component feature dimension, thereby narrowing the range of candidate components and improving the accuracy of component positioning.

[0078] S204, For the target interactive component, execute the component interaction instruction corresponding to the interaction type information.

[0079] Specifically, a predefined instruction description language (ACTION {type:.......,target:component_id, coords:[ x, y, w, h ], meta:{row,col,index}}) can be used. This instruction description language serves as an intermediate representation between voice / text information and component interaction instructions. For example, the instruction description language can use JSON. The large language model can generate the instruction description language based on the interaction type information and the component attribute information of the target interactive component. The backend executes the component interaction instructions corresponding to the instruction description language, thus achieving the same response effect as manually operating the interface components through voice interaction. For instance, taking the voice / text information "Click the third item in the second row on the left" as an example, the intermediate representation could be ACTION {type:CLICK, target:component_id, coords:[x, y,w, h], meta:{row:2,col:3}}.

[0080] In an optional embodiment, the voice text information, component attribute information of at least one component in the current interface, and an example of the instruction description language can also be input into the large language model. The large model directly outputs the instruction description language, reducing the processing latency of voice interaction through end-to-end component location processing. For example, the input data for the large language model can be: System:You are an assistant that maps user utterances to UIactions.Only use conponent IDs provided. Context: - Active app: JiguangTV -UI components: 1) id=List-serialstype=gridtext="TV scrolling list"bounds=[120,220,800,600] 2) id=item-2-3 type=cardtext="XXX TV series" bounds=[320,428,160,188] User: Please open the third one in the second row of the TV series list for me. Instruction: Return JSON:{action:CLICK, target_id:"...", meta:{...}}.If no match return {action:"NO_MATCH"}." As can be seen from the above embodiments, in response to a voice input operation on the current interface, the system obtains the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. Component descriptions are extracted from the voice text information to obtain component feature description information and interaction type information. Then, based on the component feature description information and the component attribute information of at least one component, component localization is performed to determine the target interactive component among at least one component. By extracting the component feature description information from the voice text information of the user's voice input and accurately locating the target component that the user wishes to interact with, a simulated interaction effect between the user and the target component is generated. This breaks the limitation of traditional voice assistants that can only perform fixed operations. Users can interact flexibly and accurately with various components in the current interface simply by voice input, improving the convenience of component interaction and the user's voice interaction experience.

[0081] In a specific embodiment, such as Figure 4 As shown, before determining the target interactive component in at least one component by locating the component based on component feature description information and component attribute information of at least one component, the above method may further include: S401, Determine the interactive hotspots in the current interface.

[0082] Specifically, the interactive hotspot can be the interactive focal area in the current interface, that is, the area where the user is likely to interact. In a specific embodiment, determining the interactive hotspot in the current interface may include: S4011 (not shown in the figure) obtains the interface layout information of the current interface, the focus component of at least one component, and the voice context information before the voice input operation is triggered.

[0083] Specifically, the interface layout information is used to characterize the layout of each interface area in the current interface. The focus component in at least one component can be the component that is in the focus state in at least one component. Here, the focus state can refer to the selected state or the active state. The voice context information can refer to the text information corresponding to at least one round of voice acquisition data entered by the user (speaker) before triggering the voice input operation.

[0084] S4012 (not shown in the figure) predicts the interaction hotspots of the current interface based on interface layout information, voice context information, and focus components, and determines the interaction hotspots.

[0085] Specifically, interface layout information, speech context information, and focused components can be input into a large language model. The large language model then combines the speech context information and focused components to predict interactive hotspots within the interface areas covered by the interface layout information, thus determining the interactive hotspots within the current interface area. For illustration, the interactive hotspots can be the interface areas where focused components are located in various interface areas of the current interface.

[0086] As can be seen from the above embodiments, combining voice context information and focus components to predict the interaction hotspots in the interface areas involved in the interface layout information and determine the interaction hotspots in the current interface area can improve the accuracy of interaction hotspot prediction.

[0087] S402, determine the hot zone component corresponding to the interactive hot zone from at least one component.

[0088] In one specific embodiment, the component attribute information of each component includes: the component position information of each component, and the above-mentioned determination of the hot zone component corresponding to the interactive hot zone from at least one component may include: At least one component that has a binding relationship with the interactive hotspot is designated as the hotspot component; And / or, designate at least one component whose component location information is located in the interactive hot zone as a hot zone component.

[0089] Specifically, the interactive hotspot is a UI area within the current interface. The binding relationship between the UI area and components can be pre-set based on the component layout requirements of the actual application. For example, if the content of a component matches the content of the UI area, a binding relationship can be set between the component and the UI area. For instance, if component 'a' is the entry point for TV series 1, and UI area A is the list of TV series displayed, then the content of component 'a' is considered to match the content of UI area A. Specifically, components located in the interactive hotspot can include components within the interactive hotspot and components located above the interactive hotspot.

[0090] As can be seen from the above embodiments, at least one component that has a binding relationship with the interactive hot zone is designated as a hot zone component, and / or at least one component whose component position information is located in the interactive hot zone is designated as a hot zone component, thereby specifically narrowing the selection range of components.

[0091] The above-mentioned component localization based on component feature description information and component attribute information of at least one component, determining the target interactive component in at least one component, may include: S403, based on the component feature description information and the component attribute information of the hot zone component, perform component localization and determine the target interactive component in the hot zone component.

[0092] Specifically, the details of step S403 are similar to those of the aforementioned step S203, and will not be repeated here.

[0093] As can be seen from the above embodiments, the interactive hotspots in the current interface are determined, and the hotspot components corresponding to the interactive hotspots are determined from at least one component. Then, the component feature description information and the component attribute information of the hotspot components are input into the large language model for component localization, and the target interactive components in the hotspot components are determined. By utilizing the hotspot information, the hit robustness of component localization is further improved.

[0094] In an optional embodiment, such as Figure 5 As shown, before executing the component interaction instructions corresponding to the interaction type information for the target interactive component, the above method may further include: S501, obtain the voice context information and historical component interaction information before triggering the voice input operation.

[0095] Specifically, voice context information can refer to the text information corresponding to at least one round of voice acquisition data input by the user (speaker) before triggering the voice input operation; historical component interaction information can characterize the user's interaction with at least one component in the current interface before triggering the voice input operation. Historical component interaction information can include: at least one historical interaction record, each historical interaction record containing data such as interaction time, interaction component identifier, and interaction result.

[0096] S502 performs contextual correlation analysis on the target interactive component based on voice context information to determine contextual correlation indicators.

[0097] Specifically, contextual correlation metrics are used to characterize the degree of correlation between the target interactive component and the speech context information. In a specific embodiment, the above-mentioned contextual correlation analysis of the target interactive component based on speech context information to determine the contextual correlation metrics may include: S5021 (not shown in the figure) determines the component type description information from the component feature description information extracted from the speech context information.

[0098] Specifically, the detailed content of "extracting component feature description information from speech context information" is similar to the detailed content of "extracting component description from speech text information to obtain component feature description information and interaction type information" in the aforementioned step S202, and will not be repeated here.

[0099] In a specific embodiment, the component type description information is the description information belonging to the component type dimension in the component feature description information. Here, the component type dimension may include, but is not limited to, the function type dimension and the content type dimension.

[0100] S5022 (not shown in the figure), determine the component type information and component display text of the target interactive component from the component attribute information of the target interactive component.

[0101] Specifically, the component type information and component display text are extracted from the component attribute information of the target interactive component.

[0102] S5023 (not shown in the figure) performs type association analysis on component type information and component type description information to obtain type association index.

[0103] Specifically, the type association index is used to characterize the degree of type association between the component type information of the target interactive component and the component type description information in the voice context information. In a specific embodiment, the larger the value of the type association index, the higher the degree of type association between the component type information and the component type description information is considered.

[0104] For example, the voice context information is "turn up the brightness", the current interface is the settings interface, the voice text information corresponding to the current voice input operation is "click the first one at the top", and the component type description information extracted from the voice context information is "type=adjust brightness". If the component type information of the target interactive component is "slider", it can be considered that the component type "slider" is highly related to the component type of the setting component "adjust brightness". If the component type information of the target interactive component is "text label", it can be considered that the component type "text label" is less related to the component type of the setting component "adjust brightness". Therefore, the type correlation index corresponding to "slider" is greater than the type correlation index corresponding to "text label".

[0105] S5024 (not shown in the figure) performs semantic association analysis on the component display text and component type description information to obtain semantic association indicators.

[0106] Specifically, the semantic association index is used to characterize the degree of semantic association between the component display text of the target interactive component and the component type description information in the voice context information. In a specific embodiment, the higher the value of the semantic association index, the higher the degree of semantic association between the component display text and the component type description information is considered.

[0107] For example, the voice context information is "open the TV series list", the voice text information corresponding to the current voice input operation is "click the third one in the second row on the left", and the component type description information extracted from the voice context information is "category=TV series". If the component display text of the target interactive component is "TV series 《xxxx》", it can be considered that the semantic association between "TV series 《xxxx》" and "category=TV series" is high. If the component display text of the target interactive component is "Variety show 《xxxx》", it can be considered that the semantic association between "Variety show 《xxxx》" and "category=TV series" is low. Therefore, the semantic association index corresponding to "TV series 《xxxx》" is greater than the semantic association index corresponding to "Variety show 《xxxx》".

[0108] For example, the voice context information is "find a science fiction movie", and the voice text information corresponding to the current voice input operation is "click the third one in the second row on the left". The component type description information extracted from the voice context information is "genre=science fiction". If the component display text of the target interactive component is "Interstellar", it can be determined from the local knowledge base that the semantic association between "Interstellar" and "genre=science fiction" is high. If the component display text of the target interactive component is "Dream of the Red Chamber", it can be determined from the local knowledge base that the semantic association between "Dream of the Red Chamber" and "genre=science fiction" is low. Therefore, the semantic association index corresponding to "Interstellar" is greater than the semantic association index corresponding to "Dream of the Red Chamber".

[0109] S5025 (not shown in the figure) obtains the context association index based on at least one of the type association index or semantic association index.

[0110] In an optional embodiment, the type association index can be used as the context association index, or the semantic association index can be used as the context association index. Alternatively, the type association index and the semantic association index can be weighted and fused to obtain the context association index.

[0111] As an illustration, the contextual relevance metric can be expressed as the following formula: f_ij = w_type·type_ij + w_text·text_ij Where i represents speech context information, j represents the target interactive component, f_ij represents the context association index, w_match represents the weight of the type association index, match_ij represents the type association index, match_ij∈ {0,1}, w_text represents the weight of the semantic association index, text_ij represents the semantic association index, text_ij∈ {0,1}.

[0112] As can be seen from the above embodiments, by determining the component type description information from the component feature description information extracted from the voice context information, determining the component type information and the component display text of the target interactive component from the component attribute information of the target interactive component, performing type association analysis on the component type information and the component type description information to obtain a type association index, and performing semantic association analysis on the component display text and the component type description information to obtain a semantic association index, and obtaining a context association index based on at least one of the type association index or the semantic association index, the degree of association between the target interactive component and the voice context information can be reasonably evaluated.

[0113] S503, based on historical component interaction information, performs historical interaction correlation analysis on the target interactive component to determine historical interaction correlation indicators.

[0114] Specifically, historical interaction correlation metrics are used to characterize the degree of user's historical interaction preferences for the target interactive component. In a specific embodiment, the above-mentioned historical interaction correlation analysis of the target interactive component based on historical component interaction information, to determine the historical interaction correlation metrics, includes: S5031 (not shown in the figure): Based on historical component interaction information, determine the number of historical interactions and the time of the last interaction of the target interactive component.

[0115] Specifically, historical component interaction information may include: at least one historical interaction record, and statistical analysis of at least one historical interaction record to obtain the historical interaction count and last interaction time of the target interactive component.

[0116] S5032 (not shown in the figure) is based on the number of historical interactions, and frequency normalization is performed to obtain the historical interaction frequency index.

[0117] Specifically, the historical interaction frequency metric is used to characterize the frequency of historical interactions between a user and a target interactive component. Illustratively, the historical interaction frequency metric can be expressed as the following formula:

[0118] Where j represents the target interactive component, s_freq_j represents the historical interaction frequency index, f_j represents the number of historical interactions, and N represents the total number of records with at least one historical interaction record.

[0119] S5033 (not shown in the figure) is based on the time difference between the current time and the last interaction time, and time normalization is performed to obtain the recent interaction correlation index.

[0120] Specifically, the recent interaction correlation metric is used to characterize the degree of recent user interaction with the target interactive component. Illustratively, the recent interaction correlation metric can be expressed by the following formula:

[0121] Where j represents the target interactive component, s_last_j represents the recent interaction correlation index, current_time represents the current time, t_last_j represents the time of the last interaction, and τ represents the time decay factor.

[0122] S5034 (not shown in the figure), determine the historical interaction correlation index based on at least one of the historical interaction frequency index or the recent interaction correlation index.

[0123] In an optional embodiment, the historical interaction frequency index can be used as the historical interaction correlation index, or the recent interaction correlation index can be used as the historical interaction correlation index. Alternatively, the historical interaction frequency index and the recent interaction correlation index can be weighted and fused to obtain the historical interaction correlation index.

[0124] As an illustration, historical interaction correlation indicators can be expressed by the following formula: s_hist_j = w_freq·s_freq_j + w_last·s_last_j, where s_hist_j represents the historical interaction correlation index, w_freq represents the weight of the historical interaction frequency index, s_freq_j represents the historical interaction frequency index, w_last represents the weight of the recent interaction correlation index, and s_last_j represents the recent interaction correlation index.

[0125] As can be seen from the above embodiments, based on historical component interaction information, the historical interaction count and last interaction time of the target interactive component are determined. Based on the historical interaction count, frequency normalization processing is performed to obtain the historical interaction frequency index. Based on the time difference between the current time and the last interaction time, time normalization processing is performed to obtain the recent interaction correlation index. Based on at least one of the historical interaction frequency index or the recent interaction correlation index, the historical interaction correlation index is determined. This can reasonably assess the user's historical interaction preference for the target interactive component.

[0126] S504. Based on at least one of the contextual correlation indicators or historical interaction correlation indicators, perform interaction confidence analysis on the target interaction component to determine the interaction confidence indicators of the target interaction component.

[0127] Specifically, the interaction confidence index of the target interaction component is used to evaluate the positioning reliability of the target interaction component. In an optional embodiment, the context association index can be used as the interaction confidence index, or the historical interaction association index can be used as the interaction confidence index. Alternatively, the context association index and the historical interaction association index can be weighted and fused to obtain the interaction confidence index.

[0128] In an optional embodiment, in addition to contextual correlation metrics and historical interaction correlation metrics, the following metrics can also be combined to perform interaction confidence analysis on the target interaction component: (1) Text similarity index s_text: used to measure the text similarity between the component display text / component label of the target interactive component and the voice text information.

[0129] (2) Spatial conformity index s_spatial: used to measure the spatial conformity of the target interactive component in the current interface, that is, the consistency between the component position of the target interactive component and the actual visible position.

[0130] (3) Visibility index s_vis: used to measure the visibility of the target interactive component, that is, to assess whether the target interactive component is visible, whether it is in the viewport, and whether it is enabled.

[0131] (4) Interaction type matching index s_type: used to measure the matching degree between the function type of the target interactive component and the interaction type information corresponding to the voice and text information.

[0132] (5) Character recognition reliability index s_ocr: used to measure the consistency between the component recognition text extracted by OCR and the component display text.

[0133] As an illustration, the character recognition reliability index can be expressed by the following formula:

[0134] Where j represents the target interactive component, s_ocr_j represents the character recognition reliability index, OCR_j represents the component-recognized text, T_j represents the component-displayed text, ED(OCR_j, T_j) represents the edit distance between OCR_j and T_j, ED(OCR_j, T_j) measures the minimum number of single-character editing operations (including insertion, deletion, and replacement) required to convert the string OCR_j into T_j, len(OCR_j) represents the length of the component-recognized text, and len(T_j) represents the length of the component-displayed text. It is understood that if the system does not extract the component-recognized text through OCR, then s_ocr_j = 0.

[0135] Accordingly, at least two of the following indicators can be weighted and fused to obtain the interaction confidence index: text similarity index, spatial conformity index, visibility index, interaction type matching index, character recognition reliability index, context association index, and historical interaction association index.

[0136] The above-mentioned execution of component interaction instructions corresponding to the interaction type information for the target interactive component may include: S2041, based on the interaction confidence index, execute the component interaction instruction corresponding to the interaction type information for the target interaction component.

[0137] Specifically, if the interaction confidence index is greater than the first preset confidence threshold, the component interaction instruction corresponding to the interaction type information is executed for the target interaction component; if the interaction confidence index is less than or equal to the first preset confidence threshold, the component interaction instruction corresponding to the interaction type information is not executed temporarily.

[0138] In an optional embodiment, if the interaction confidence index is greater than a second preset confidence threshold and less than or equal to a first preset confidence threshold, a secondary confirmation message is displayed to the user, such as "Please confirm whether to click the third one in the second row on the left?". If the interaction confidence index is less than or equal to the second preset confidence threshold, an interaction command execution failure message is displayed to the user. The second preset confidence threshold is less than the first preset confidence threshold. Illustratively, the first and second preset confidence thresholds can be preset based on the accuracy requirements of voice interaction in the application.

[0139] As can be seen from the above embodiments, based on voice context information, context association analysis is performed on the target interactive component to determine context association indicators. Based on historical component interaction information, historical interaction association analysis is performed on the target interactive component to determine historical interaction association indicators. Based on at least one of the context association indicators or historical interaction association indicators, interaction confidence analysis is performed on the target interactive component to determine the interaction confidence indicators of the target interactive component. Finally, based on the interaction confidence indicators, a decision is made on whether to execute the component interaction command corresponding to the interaction type information for the target interactive component. By improving the quality of interaction confidence analysis, the positioning reliability of the target interactive component is guaranteed, thereby improving the accuracy of subsequent component interaction command execution.

[0140] In an optional embodiment, such as Figure 6 As shown, after extracting component descriptions from the speech-text information to obtain component feature description information and interaction type information, the above method may further include: S601, if the target interactive component in at least one component is not identified, obtain a screenshot of the current interface.

[0141] S602, perform component recognition on the screenshot of the interface, and determine the component recognition text of at least one recognized component and the recognition position information of at least one recognized component in the screenshot of the interface.

[0142] Specifically, component recognition here can employ object detection and OCR (Optical Character Recognition) technologies. To illustrate, component recognition can first be performed on the screenshot of the interface based on the object detection model to obtain the recognition location information of at least one component. Then, optical character recognition can be performed on the bounding box corresponding to the recognition location information of each component to obtain the component recognition text of each component.

[0143] S603, based on component recognition text and recognition location information, performs visual verification on the component attribute information of at least one component.

[0144] Specifically, visual verification here can include consistency verification of component display text and consistency verification of component position, to confirm whether the component position and component display text actually visible to the user are consistent with the component position and component display text recorded by the system, thereby correcting component attribute information that is inconsistent in the verification.

[0145] As can be seen from the above embodiments, in the event of component location failure, visual verification of component attribute information is performed by combining interface screenshots, thereby correcting inconsistent component attribute information and improving the accuracy of subsequent voice interaction.

[0146] In an optional embodiment, if the voice text information expresses complex interactive requirements, the voice text information can be decomposed to obtain sub-text information corresponding to each interactive requirement in the complex interactive requirements, and the component interaction instructions corresponding to each sub-text information can be executed sequentially. If the expected interactive component corresponding to any sub-text information fails to be located or the component interaction instruction corresponding to any sub-text information fails to be executed, all component interaction instructions already executed based on the voice text information are rolled back to ensure the transactional nature of the interaction instructions. For example, if the voice text information is "Open the TV series list and select the third one in the second row", the voice text information can be decomposed into sub-text information 1 "Open the TV series list" and sub-text information 2 "Select the third one in the second row". If the component interaction instruction corresponding to sub-text information 1 "Open the TV series list" is executed successfully, while the expected interactive component corresponding to sub-text information 2 "Select the third one in the second row" fails to be located or the component interaction instruction fails to be executed, the component interaction instruction corresponding to sub-text information 1 "Open the TV series list" is rolled back, that is, the interface state before the "TV series list" is opened is returned.

[0147] In an optional embodiment, security and permission checks can also be performed on the voice and text information. If the interaction type involved in the voice and text information is a sensitive interaction type (e.g., payment, account management), a secondary confirmation message can be forcibly displayed to the user to ensure the security of the voice interaction. At the same time, a command blacklist can be set to prevent accidental triggering of commands in the blacklist.

[0148] In an optional embodiment, the large language model can be deployed on a server. The terminal device can package the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface to obtain an interaction instruction data packet, and upload the interaction instruction data packet to the server. The server can parse and process the interaction instruction data packets received from different terminal devices in parallel to reduce network overhead.

[0149] As can be seen from the technical solutions provided in the embodiments of this specification above, in the voice interaction scenario between the user and the terminal interface, in response to a voice input operation on the current interface, the system obtains the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. Component descriptions are extracted from the voice text information to obtain component feature description information and interaction type information. The component feature description information and the component attribute information of at least one component are then input into a large language model for component localization to determine the target interactive component among at least one component. By extracting the component feature description information from the voice text information of the user's voice input, and using the large language model based on the component feature description information and the component attribute information of the current interface, the system accurately locates the target component that the user wishes to interact with, thereby generating a simulated interaction effect between the user and the target component. This breaks the limitation of traditional voice assistants that can only perform fixed operations. With the limitations of voice input, users can interact flexibly and accurately with various components on the current interface simply by using voice input, improving the convenience of component interaction and the user's voice interaction experience. In addition, contextual correlation analysis can be performed on the target interactive component based on voice context information to determine contextual correlation indicators, and historical interaction correlation analysis can be performed on the target interactive component based on historical component interaction information to determine historical interaction correlation indicators. Based on at least one of the contextual correlation indicators or historical interaction correlation indicators, interaction confidence analysis can be performed on the target interactive component to determine the interaction confidence indicators of the target interactive component. Finally, based on the interaction confidence indicators, a decision is made on whether to execute the component interaction command corresponding to the interaction type information for the target interactive component. By improving the quality of interaction confidence analysis, the positioning reliability of the target interactive component is guaranteed, thereby improving the accuracy of subsequent component interaction command execution.

[0150] In one specific embodiment, the current interface can be the interface displayed on the screen of the terminal device. Taking a television terminal as an example, the terminal device may be a television terminal. Figure 7This is a flowchart illustrating a voice interaction method for a television terminal provided in an embodiment of this application. Specifically, as shown... Figure 7 As shown, the voice interaction method may include the following steps: S701: Users can trigger voice input operations by using the voice interaction button on the remote control device connected to the TV terminal; S702: The television terminal responds to the voice input operation by collecting user voice data and component attribute information of at least one component in the current television interface. S703: The TV terminal performs voice enhancement processing on the user's voice data to obtain enhanced voice data; S704: The television terminal performs text recognition on the enhanced voice data to obtain voice-text data; S705: The TV terminal extracts component descriptions from the voice and text information to obtain component feature description information and interaction type information; S706: The television terminal inputs the component feature description information and the component attribute information of at least one component into the large language model to locate the component and determine the target interactive component in at least one component. S707: The television terminal converts the component attribute information and interaction type information of the target interactive component into a structured instruction description language; S708: The TV terminal determines the component interface of the target interactive component from the component interface list in the backend; S709: The TV terminal uses the component interface of the target interactive component to execute the component interaction instructions corresponding to the instruction description language.

[0151] As can be seen, this solution accurately locates the UI components that the user expects to interact with by acquiring real-time information about the UI components of the current TV interface, leveraging the powerful semantic understanding capabilities of the large language model, and combining the text data of the user's voice input with the UI component information. This breaks the constraints of fixed operations in traditional voice assistants, allowing users to simulate interactions with all visible components on the TV interface simply by using voice. This greatly enhances the user's TV experience and provides a better solution for barrier-free interaction, enabling more users to use electronic devices smoothly.

[0152] This application also provides a voice interaction device. Figure 8 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application, as shown below. Figure 8 As shown, the above-mentioned device includes: The first information acquisition module 810 is used to respond to a voice input operation on the current interface and acquire the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. The component description extraction module 820 is used to extract component descriptions from voice and text information to obtain component feature description information and interaction type information. The component localization module 830 is used to locate components based on component feature description information and component attribute information of at least one component, and to determine the target interactive component in at least one component. The interaction instruction execution module 840 is used to execute component interaction instructions corresponding to the interaction type information for the target interactive component.

[0153] In one specific embodiment, the component localization module 830 is further configured to input component feature description information and component attribute information of at least one component into a large language model for component localization, and determine the target interactive component in at least one component.

[0154] In one specific embodiment, the component feature description information includes: feature description information for at least one component feature dimension, and the component description extraction module 820 is further configured to: Input the feature description information of at least one component feature dimension and the component attribute information of at least one component into the large language model, and the large language model performs the following operations: Based on the positioning priority of at least one component feature dimension, determine the current component feature dimension among at least one component feature dimension; Select at least one component as the current matching component; From the currently matched components, identify candidate components whose component attribute information matches the feature description information of the current component's feature dimension; Based on the positioning priority of at least one component feature dimension, update the current component feature dimension; take the candidate component as the current matching component, and repeatedly execute the component positioning process from determining the candidate component whose component attribute information matches the feature description information of the current component feature dimension from the current matching component to taking the candidate component as the current matching component, until the preset positioning end condition is reached. The candidate component that meets the preset positioning termination condition will be used as the target interactive component.

[0155] In a specific embodiment, the component description extraction module 820 is further configured to: determine the sub-attribute information corresponding to the feature dimension of the current component in the component attribute information of the currently matched component; determine the sub-attribute filtering conditions based on the feature description information of the feature dimension of the current component; and determine the candidate components whose sub-attribute information satisfies the sub-attribute filtering conditions from the currently matched components.

[0156] In one specific embodiment, the above-described apparatus further includes: The interactive hotspot determination module is used to determine the interactive hotspots in the current interface; A hot zone component determination module is used to determine the hot zone component corresponding to the interactive hot zone from at least one component; The aforementioned component positioning module 830 is also used to locate components based on component feature description information and component attribute information of hot zone components, and to determine the target interactive component in the hot zone components.

[0157] In one specific embodiment, the above-mentioned interactive hotspot determination module is further used for: Obtain the current interface layout information, the focus component in at least one component, and the voice context information prior to triggering the voice input operation; Based on interface layout information, voice context information, and focus components, the interaction hotspots of the current interface are predicted and determined.

[0158] In one specific embodiment, the component attribute information of each component includes: component location information of each component; the aforementioned hot zone component determination module is further used for: At least one component that has a binding relationship with the interactive hotspot is designated as the hotspot component; And / or, designate at least one component whose component location information is located in the interactive hot zone as a hot zone component.

[0159] In one specific embodiment, the above-described apparatus further includes: The second information acquisition module is used to acquire voice context information and historical component interaction information before the voice input operation is triggered. The context association analysis module is used to perform context association analysis on target interactive components based on voice context information and determine context association indicators. The historical interaction correlation analysis module is used to perform historical interaction correlation analysis on target interactive components based on historical component interaction information, and to determine historical interaction correlation indicators. The interaction confidence analysis module is used to perform interaction confidence analysis on the target interaction component based on at least one of the contextual correlation indicators or historical interaction correlation indicators, and to determine the interaction confidence indicators of the target interaction component. The interaction instruction execution module 840 is also used to: based on the interaction confidence index, execute the component interaction instruction corresponding to the interaction type information for the target interaction component.

[0160] In one specific embodiment, the above-mentioned context association analysis module is further used for: Component type description information is determined from the component feature description information extracted from the speech context information; From the component attribute information of the target interactive component, determine the component type information and the component display text of the target interactive component; Type association analysis is performed between component type information and component type description information to obtain type association indicators; Semantic association analysis is performed on the component display text and component type description information to obtain semantic association indicators; The context association index is obtained based on at least one of the type association index or the semantic association index.

[0161] In one specific embodiment, the aforementioned historical interaction association analysis module is further used for: Based on historical component interaction information, determine the number of historical interactions and the time of the last interaction of the target interactive component; Based on the number of historical interactions, frequency normalization is performed to obtain the historical interaction frequency index. Based on the time difference between the current time and the last interaction time, time normalization is performed to obtain recent interaction correlation indicators. Historical interaction correlation indicators are determined based on at least one of the following indicators: historical interaction frequency indicators or recent interaction correlation indicators.

[0162] In an optional embodiment, the above-described apparatus further includes: The screenshot acquisition module is used to acquire a screenshot of the current interface when the target interactive component in at least one component has not been identified. The component recognition module is used to recognize components in the screenshot and determine the component recognition text and the recognition location information of at least one component in the screenshot. The visual verification module is used to perform visual verification on the component attribute information of at least one component based on component recognition text and recognition location information.

[0163] In an optional embodiment, the above-described apparatus further includes: The component attribute extraction module is used to extract the component attribute information of at least one interface component in the current interface. The component filtering module is used to select at least one UI component whose component attribute information indicates that it is in a visible and interactive state as at least one component.

[0164] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0165] Figure 9 This is a block diagram of an electronic device for voice interaction provided in an embodiment of this application. The electronic device can be a terminal device, and its internal structure diagram can be as follows: Figure 9As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice interaction method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse. Figure 10 This is a block diagram of another electronic device for voice interaction provided in an embodiment of this application. The electronic device can be a server, and its internal structure diagram can be as follows: Figure 10 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice interaction method. Those skilled in the art will understand that Figure 9 or Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing a computer program executable by the processor; wherein the processor is configured to execute the computer program to implement the voice interaction method provided in the various optional implementations described above.

[0166] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the voice interaction methods provided in the various optional implementations described above. In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the voice interaction methods provided in the various optional implementations described above.

[0167] It is understood that in the specific implementation of this application, user-related data is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0168] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media in the processing steps of the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0169] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0170] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A voice interaction method, characterized in that, The method includes: In response to a voice input operation on the current interface, obtain the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. Component descriptions are extracted from the voice and text information to obtain component feature description information and interaction type information; Based on the component feature description information and the component attribute information of each of the at least one component, the component is located to determine the target interactive component among the at least one component; For the target interactive component, execute the component interaction instruction corresponding to the interaction type information.

2. The method according to claim 1, characterized in that, The step of locating components based on the component feature description information and the component attribute information of each of the at least one component, and determining the target interactive component among the at least one component, includes: The component feature description information and the component attribute information of each of the at least one component are input into the large language model for component localization to determine the target interactive component.

3. The method according to claim 2, characterized in that, The component feature description information includes: feature description information for at least one component feature dimension; the step of inputting the component feature description information and the component attribute information of the at least one component into a large language model for component localization, and determining the target interactive component, includes: The feature description information of each of the at least one component feature dimension and the component attribute information of each of the at least one component are input into the large language model, and the large language model performs the following operations: Based on the positioning priority of the at least one component feature dimension, determine the current component feature dimension among the at least one component feature dimensions; and use the at least one component as the current matching component; From the currently matched components, determine candidate components whose component attribute information matches the feature description information of the current component's feature dimension; Based on the positioning priority of the at least one component feature dimension, update the current component feature dimension; take the candidate component as the current matching component, and repeatedly execute the process of determining the candidate component whose component attribute information matches the feature description information of the current component feature dimension from the current matching component to taking the candidate component as the current matching component, until the preset positioning end condition is reached. The candidate component that reaches the preset positioning termination condition will be used as the target interactive component.

4. The method according to claim 3, characterized in that, The candidate components determined from the currently matched components whose component attribute information matches the feature description information of the current component's feature dimension include: Determine the sub-attribute information in the component attribute information of the currently matched component that corresponds to the feature dimension of the current component; Based on the feature description information of the current component feature dimension, determine the sub-attribute filtering conditions; From the currently matched components, candidate components whose sub-attribute information satisfies the sub-attribute filtering conditions are determined.

5. The method according to claim 1, characterized in that, Before determining the target interactive component among the at least one components by locating the component based on the component feature description information and the component attribute information of each of the at least one component, the method further includes: Determine the interactive hotspots in the current interface; Determine the thermal zone component corresponding to the interactive thermal zone from the at least one component; The step of locating components based on the component feature description information and the component attribute information of each of the at least one component, and determining the target interactive component among the at least one component, includes: Based on the component feature description information and the component attribute information of the hot zone component, the component is located to determine the target interactive component in the hot zone component.

6. The method according to claim 5, characterized in that, Determining the interactive hotspots in the current interface includes: Obtain the interface layout information of the current interface, the focus component among the at least one component, and the voice context information prior to triggering the voice input operation; Based on the interface layout information, the voice context information, and the focus component, the interaction hotspot is predicted for the current interface to determine the interaction hotspot.

7. The method according to claim 5, characterized in that, The component attribute information for each component includes: the component position information for each component, and determining the hotspot component corresponding to the interactive hotspot from the at least one component includes: The component that has a binding relationship with the interactive hot zone among the at least one components is designated as the hot zone component; And / or, the component whose component location information is located in the interactive hot zone is designated as the hot zone component.

8. The method according to claim 1, characterized in that, Before executing the component interaction instruction corresponding to the interaction type information for the target interactive component, the method further includes: Obtain the voice context information and historical component interaction information prior to triggering the voice input operation; Based on the voice context information, a context association analysis is performed on the target interactive component to determine the context association index; Based on the historical component interaction information, a historical interaction correlation analysis is performed on the target interactive component to determine the historical interaction correlation index. Based on at least one of the context association indicators or the historical interaction association indicators, an interaction confidence analysis is performed on the target interaction component to determine the interaction confidence indicators of the target interaction component. The step of executing the component interaction instruction corresponding to the interaction type information for the target interactive component includes: Based on the interaction confidence index, for the target interaction component, execute the component interaction instruction corresponding to the interaction type information.

9. The method according to claim 8, characterized in that, The step of performing context association analysis on the target interactive component based on the voice context information to determine the context association indicators includes: Component type description information is determined from the component feature description information extracted from the speech context information; From the component attribute information of the target interactive component, determine the component type information and the component display text of the target interactive component; The component type information and the component type description information are subjected to type association analysis to obtain type association indicators; Semantic association analysis is performed on the component display text and the component type description information to obtain semantic association indicators; The context association index is obtained based on at least one of the type association index or the semantic association index.

10. The method according to claim 8, characterized in that, The step of performing historical interaction correlation analysis on the target interactive component based on the historical component interaction information to determine historical interaction correlation indicators includes: Based on the historical component interaction information, determine the number of historical interactions and the last interaction time of the target interactive component; Based on the historical interaction count, frequency normalization is performed to obtain the historical interaction frequency index. Based on the time difference between the current time and the last interaction time, time normalization is performed to obtain the recent interaction correlation index. The historical interaction correlation index is determined based on at least one of the historical interaction frequency index or the recent interaction correlation index.

11. The method according to any one of claims 1 to 10, characterized in that, After extracting component descriptions from the speech-text information to obtain component feature description information and interaction type information, the method further includes: If the target interactive component among the at least one component is not identified, obtain a screenshot of the current interface; Component recognition is performed on the screenshot of the interface to determine the component recognition text of at least one recognized component in the screenshot of the interface and the recognition location information of the at least one recognized component; Based on the component-identified text and the identification location information, visual verification is performed on the component attribute information of each of the at least one component.

12. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Extract the component attribute information of at least one interface component in the current interface; The interface component whose component attribute information indicates that it is in a visible and interactive state is designated as the at least one interface component.

13. A voice interaction device, characterized in that, The device includes: The first information acquisition module is used to respond to a voice input operation on the current interface and acquire the voice text information corresponding to the voice input operation and the component attribute information of at least one component in the current interface. The component description extraction module is used to extract component descriptions from the voice text information to obtain component feature description information and interaction type information. A component localization module is used to locate components based on the component feature description information and the component attribute information of each of the at least one component, and to determine the target interactive component among the at least one component. The interaction instruction execution module is used to execute component interaction instructions corresponding to the interaction type information for the target interaction component.

14. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the voice interaction method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the voice interaction method as described in any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes a computer program that is loaded and executed by a processor to implement the voice interaction method as described in any one of claims 1 to 12.