Display device and user intention recognition method
By combining user operations and voice signals, the mobile cursor of the display device is controlled to move on the user interface, solving the problem that the display device is difficult to recognize user intentions under specific conditions, and achieving a wider user intention recognition and a more efficient operation experience.
Patent Information
- Application Number
- CN202510293938.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for existing display devices to accurately identify user intentions under specific conditions, resulting in limitations in user intention identification.
By combining user operations and voice signals, user intention recognition is performed using agent technology, controlling the mobile cursor to move to any position on the user interface, and user intention recognition is performed by combining the displayed content and voice recognition results.
It improves the versatility and accuracy of user intention recognition, and users do not need to press keys one by one to select intention targets, which improves operational efficiency and accuracy.
Smart Images

Figure CN120340480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a display device and a user intention recognition method. Background Art
[0002] With the rapid development of display devices, the functions that display devices can provide for users are becoming increasingly rich. Taking a TV as an example, there are more and more TV scenarios. It is not only used as a device for watching TV programs at home, but also can be used for gaming, playing electronic photo albums, information display, etc.
[0003] In this context, most display devices have the function of recognizing user interaction intentions. However, currently, display devices usually need to meet specific conditions (such as specific scenarios, specific instructions, etc.) to achieve user intention recognition. In some scenarios where the user has an interaction intention but deviates from these specific conditions, it is difficult for the display device to recognize the user intention, resulting in certain limitations in current user intention recognition. Summary of the Invention
[0004] This application provides a display device and a user intention recognition method to solve the problem of limitations in current user intention recognition.
[0005] In a first aspect, some embodiments provide a display device, including: a display configured to display a user interface; a communication module configured to communicate with a control device; the control device is used to respond to a user operation and control a moving cursor on the user interface to move to any position on the user interface; a microphone configured to receive a voice signal emitted by the user; a controller configured to: when the emission time of the voice signal matches the operation time of the user operation, determine the movement trajectory of the moving cursor on the user interface; based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the voice signal, perform user intention recognition through an agent to obtain a user operation intention.
[0006] For the above display device, on the one hand, by combining the user operation triggered by the user through the control device with the user voice signal and using agent technology, user intention recognition is performed, thereby obtaining a user operation intention. Compared with the current method of user intention recognition that can only be performed under specific conditions, this solution can be applied to scenarios other than specific conditions, improving the generality of user intention recognition. On the other hand, the above control device can respond to user operations and control the moving cursor on the user interface to move to any position on the user interface. In this way, the user can use the control device to quickly and accurately select an intention target in the user interface. Compared with a traditional remote control, the user no longer needs to select the intention target by pressing one button at a time, effectively improving the accuracy and efficiency of user operations, and thus improving the accuracy and efficiency of user intention recognition.
[0007] In a second aspect, some embodiments further provide a user intention recognition method, which is applied to the display device provided in the first aspect. The display device includes: a display, a communication module, a microphone, and a controller. The method includes: when the emission time of the voice signal emitted by the user matches the operation time of the user operation triggered by the user through the control device, determining the movement trajectory of the moving cursor on the user interface; the control device is used to respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface; based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the voice signal, the intelligent agent is used to perform user intention recognition to obtain the user operation intention.
[0008] In the above user intention recognition method, on the one hand, by combining the user operation triggered by the user through the control device with the user voice signal and using the intelligent agent technology, user intention recognition is performed to obtain the user operation intention. Compared with the current method of user intention recognition that can only be performed under specific conditions, this solution can be applied to scenarios other than specific conditions, improving the versatility of user intention recognition. On the other hand, the above control device can respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface. In this way, the user can use the control device to quickly and accurately select the intention target in the user interface. Compared with the traditional remote control, the user no longer needs to select the intention target by pressing each key one by one, effectively improving the accuracy and efficiency of the user operation, and thus improving the accuracy and efficiency of user intention recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided in some embodiments of the present application;
[0011] Figure 2 It is a schematic diagram of the hardware configuration of the display device provided in some embodiments of the present application;
[0012] Figure 3 It is a schematic diagram of the hardware configuration of the control device provided in some embodiments of the present application;
[0013] Figure 4 It is a schematic diagram of the software configuration of the display device provided in some embodiments of the present application;
[0014] Figure 5 Schematic diagram of control directed to a remote controller provided in some embodiments of the present application;
[0015] Figure 6 Relationship diagram among an intelligent agent, voice signals, and a control device provided in some embodiments of the present application;
[0016] Figure 7 Interaction diagram among modules during user intention recognition provided in some embodiments of the present application;
[0017] Figure 8 Scene schematic diagram of user intention recognition provided in some embodiments of the present application;
[0018] Figure 9 Flow schematic diagram of user intention recognition provided in some embodiments of the present application;
[0019] Figure 10 Flow schematic diagram of a user intention recognition method provided in some embodiments of the present application. Detailed implementation manners
[0020] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0021] It should be noted that the brief description of terms in the present application is only for facilitating the understanding of the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0022] The terms "first", "second", "third", etc. in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such used terms can be interchanged under appropriate circumstances.
[0023] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclusively include. For example, a product or device comprising a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.
[0024] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to the element.
[0025] In the embodiments of the present application, the display device 200 generally refers to a device with the capabilities of screen display and data processing. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0026] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided in some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300, and the control device 100. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.
[0027] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 to perform data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and the connection communication can be realized through network communication protocols to achieve the purpose of one-to-one control operation and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0028] As Figure 1 also shown, the display device 200 also performs data communication with the server 400 through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0029] The display device 200 can provide a broadcast receiving TV function, and can also additionally provide a smart network TV function with computer support functions, including but not limited to, network TV, smart TV, Internet Protocol TV (IPTV), etc.
[0030] Figure 2 For some embodiments of the present application Figure 1 is the hardware configuration block diagram of the display device 200 in
[0031] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0032] In some embodiments, the detector 230 is used to collect signals of the external environment or for external interactions. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; alternatively, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0033] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving image display. The display 260 is used to receive the image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of the menu control interface, and the user control UI interface, etc.
[0034] In some embodiments, the communication device 220 is a component for communicating with external devices or the server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0035] The communication device 220 can enable the display device 200 to communicate with external devices or the server 400 in a wireless or wired connection manner. Among them, the wired connection can connect the display device 200 with external devices through components such as data lines and interfaces. The wireless connection can connect the display device 200 with external devices through wireless signals or wireless networks. The display device 200 can directly establish a connection relationship with external devices or indirectly establish a connection relationship through gateways, routers, connection devices, etc.
[0036] In some embodiments, the controller 250 can include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and the first interface to the nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0037] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, that is, the tuner demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0038] In some embodiments, the user may input a user command through the graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI).
[0039] In some embodiments, the audio output device 270 may be the built-in speaker of the display device 200 or an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0040] In some embodiments, the user input interface 280 can be used to receive instructions from the user input.
[0041] Figure 3 For some embodiments provided by this application Figure 1 The hardware configuration block diagram of the control device in. As Figure 3 shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0042] The control device 100 is configured to control the display device 200, and can receive the input operation instructions of the user, and convert the operation instructions into instructions recognizable and responsive by the display device 200, playing an intermediary role in the interaction between the user and the display device 200.
[0043] In some embodiments, the control device 100 may be an intelligent device. For example: the control device 100 can install various applications for controlling the display device 200 according to user needs.
[0044] In some embodiments, as Figure 1 shown, after the mobile terminal 300 or other intelligent electronic devices install the application for controlling the display device 200, they can play a similar function to the control device 100.
[0045] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and internal.
[0046] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, and other near-field communication modules.
[0047] User input / output interface 140, where the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a button 144, and other input interfaces.
[0048] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The communication interface 130 is configured in the control device 100, such as modules for WiFi, Bluetooth, NFC, etc., and can encode user input instructions through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol and send them to the display device 200.
[0049] A memory 190, which is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of a controller. The memory 190 can store various control signal instructions input by a user.
[0050] A power supply 180, which is used to provide operating power support for each component of the control device 100 under the control of a controller.
[0051] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can (control the display device) provide a user interface, allow a user to interact with the display device 200, and support running various application programs.
[0052] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system developed specifically for the display device.
[0053] The operating system can be divided into different modules or levels according to the functions implemented.
[0054] For example, as Figure 4 shown, in some embodiments, the system is divided into four layers, from top to bottom, which are the application layer (abbreviated as the "application layer"), the application framework layer (abbreviated as the "framework layer"), the system library layer, and the kernel layer.
[0055] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run the applications and interact with users based on the applications. At least one application can run in the application layer. These applications can be window programs, system setting programs, or clock programs that come with the operating system; they can also be applications developed by third-party developers. In specific implementation, the application packages in the application layer are not limited to the above examples.
[0056] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center, which decides to make the applications in the application layer take actions. Through the API interface, applications can access the resources in the system and obtain the services of the system during execution.
[0057] As Figure 4 shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc. Among them, the view system can design and implement the interfaces and interactions of applications. The view system includes lists, grids, text boxes, buttons, etc. The managers include at least one of the following modules: The activity manager is used to interact with all the activities running in the system; the location manager is used to provide access to the system location service for system services or applications; the package manager is used to retrieve various information related to the application packages currently installed on the device; the notification manager is used to control the display and clearing of notification messages; the window manager is used to manage the icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0058] In some embodiments, the activity manager is used to manage the life cycles of various applications and the general navigation back function, such as controlling the exit, opening, and backward of applications. The window manager is used to manage all window programs, such as obtaining the size of the display screen, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the change of the display window. For example, shrinking the display window, jittering the display, or distorting the display.
[0059] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction libraries included in the system runtime layer, such as C / C++ instruction libraries, to implement the functions that the framework layer is intended to achieve.
[0060] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, as Figure 4 shown, hardware drivers can be configured in the kernel layer, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.
[0061] It should be noted that the above examples are only simple divisions of the functions of the operating system, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of levels and the specific level types included in the operating system can be in other forms.
[0062] With the rapid development of display devices, the functions that display devices can provide for users are becoming more and more abundant. Taking a TV as an example, there are more and more TV scenarios. It is not only used as a device for watching TV programs at home, but also can be used for gaming, playing electronic photo albums, information display, etc. In this context, most display devices have the function of recognizing user interaction intentions.
[0063] Currently, user's multimodal information such as voice, lip movement, face, voiceprint, etc. is usually used to wake up or wake up the display device without waking up, so as to realize user intention recognition. However, this often needs to be achieved under specific conditions such as specific scenarios and specific instructions. In some scenarios where the user has an interaction intention but deviates from these specific conditions, for example, when the user is watching TV and suddenly finds a relatively interesting target, and the user directly says "introduce this to me", in this case, the display device actually doesn't know what the pronoun "this" refers to, so it is difficult to accurately recognize the user's intention. To sum up, there are certain limitations in the current display device for user intention recognition.
[0064] To solve this technical problem, in some embodiments, a display device is provided, and the display device includes:
[0065] A display, configured to display a user interface.
[0066] Among them, the user interface is a media interface for interaction and information exchange between an application or an operating system and a user. The common manifestation form of the user interface is the graphical user interface, which refers to the user interface related to computer operations displayed in a graphical manner. It can be an interface element such as an icon, a window, a control, etc. displayed on the display screen of a display device, where the control can include visible interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, etc.
[0067] In some embodiments, the display is further configured to display the interface display content corresponding to the user operation intention.
[0068] In some embodiments, the display can also be configured to display the movement trajectory of the moving cursor. Specifically, the display can completely display the movement trajectory of the moving cursor in the user interface, such as highlighting the movement trajectory with light, so that the user can intuitively see the operation effect.
[0069] Optionally, the display can also be configured to always only display the light point of the moving cursor, that is, the current position of the cursor on the user interface, so as to avoid blocking the interface content. Specifically, it can be configured according to actual needs, and this embodiment does not limit this.
[0070] The communication module is configured to communicate with the control device; the control device is used to respond to the user operation and control the moving cursor on the user interface to move to any position of the user interface.
[0071] Among them, the control device can be an intelligent device for receiving the operation initiated by the user on the user interface, such as a remote control device (such as a remote control handle, a pointing remote control), or a terminal device (such as a smart phone, a tablet) installed with control software or a program for responding to the user operation and controlling the movement of the moving cursor.
[0072] In a preferred example, the control device is a pointing remote control. As the name implies, a pointing remote control means pointing and remotely controlling anywhere. The user can directly point to any position in the display device through the pointing remote control. Figure 5 Shows the control schematic diagram of the pointing remote control. Figure 5 The circle in the shown user interface is the moving cursor. From the perspective of interaction, the pointing remote control can solve the problems of clumsy operation and complex buttons of traditional remote controls. In some scenarios, it can also solve the clumsiness of directive voice. The user does not need to clearly state the specific position of the intended target in the user interface, such as "play the first audio in the third row". The user only needs to point to the first audio in the third row with the pointing remote control and say "play this" or "play" to achieve the intended operation of playing the audio.
[0073] Optionally, after pointing to the audio, the user can also press the "Confirm" button on the pointing remote control to play the audio.
[0074] The user operation can refer to an operation on the user interface triggered by the user through a control device, which can be but is not limited to operations such as circle selection, point selection, and sliding. Moving the cursor refers to a cursor that can move flexibly in the user interface and can be flexibly moved to any position in the user interface through the control of the control device.
[0075] Optionally, the user operation can be an operation triggered by the user physically moving the control device. For example, if the user directly holds the control device and draws a circle in the air, it can be mapped to a circle selection operation on the user interface.
[0076] Optionally, the user operation can also be an operation triggered by the user on the touch screen of the control device. For example, if the user performs any sliding, circle selection, etc. operations on the touch screen of the control device, it can also be mapped to an operation on the user interface.
[0077] The microphone is configured to receive the voice signal emitted by the user.
[0078] Among them, the voice signal refers to the voice signal initiated by the user to the display device. In the voice control of the display device, the voice signal can be collected by the microphone and converted into an electrical signal, and then the text information in the voice signal can be recognized based on the electrical signal, so as to identify the user's intention according to the text information.
[0079] Optionally, the microphone can be a microphone array composed of at least two microphones.
[0080] The controller is configured to: determine the movement trajectory of the moving cursor on the user interface when the emission time of the voice signal matches the operation time of the user operation.
[0081] Among them, the emission time refers to the trigger time of the voice signal, and the operation time refers to the trigger time of the user operation.
[0082] Optionally, if the trigger time of the voice signal is the same as the trigger time of the user operation, it can be considered that the two match.
[0083] Optionally, if the time difference between the trigger time of the voice signal and the trigger time of the user operation is within a preset time range (such as within 1 second), it can be considered that the two match.
[0084] The movement trajectory can be understood as the trajectory when the moving cursor moves on the user interface, which can be a straight line or a curve, or a circle or an irregular shape. The specific form of the movement trajectory is determined based on the user operation. Whatever shape the user operation is, the movement trajectory can be the same shape.
[0085] In some embodiments, the movement trajectory of the moving cursor can be determined based on the interface coordinates in the user interface mapped by the operation trajectory of the user operation.
[0086] Exemplarily, the controller can receive the user operation triggered by the user through the control device and the voice signal sent by the user through the microphone. And perform matching detection on the user operation and the voice signal. Specifically, it can respectively obtain the triggering time of the user operation and the generation time of the voice signal. If the two times are the same or the time difference between the two times is within the preset time range, it is considered that the user operation and the voice signal match. Then, further determine the operation trajectory mapped by the user operation on the user interface, that is, the movement trajectory of the moving cursor. If the two times are inconsistent or the time difference between the two times is not within the preset time range, it is considered that the user operation and the voice signal do not match, and the current user intention recognition may not be responded to.
[0087] Based on the display content in the display area corresponding to the movement trajectory and the speech recognition result of the voice signal, the user intention is recognized through the agent to obtain the user operation intention.
[0088] Among them, the display area may refer to the interface area in the user interface corresponding to the movement trajectory. For example, if the movement trajectory is a circle, then the display area is the area enclosed by this circle. If the movement trajectory is a line, then the display area may be the area where the interface element currently pointed to by the moving cursor is located. For example, the user controls the moving cursor to move from element A to element B through the control device. Then the display area at this time may be the area where element B is located, that is, the area containing element B.
[0089] The display content may refer to the content in the display area. For example, the object type (person, animal, text, video, audio, etc.) of the target object in the display area. Of course, in addition to the object type, the display content may also include but is not limited to the size of the display area, the position of the display area in the user interface, etc.
[0090] Optionally, in practical applications, the display content can be stored as a content file as the input of the agent. In one example, the content file may include the object type of the target object in the display area.
[0091] The speech recognition result may refer to the result obtained by recognizing the voice signal, including but not limited to the text information of the voice signal, the speech recognition duration, the speech recognition accuracy rate, etc. User intention recognition refers to the process of recognizing the user operation intention. The user operation intention refers to the requirement or goal expressed by the user operation, including but not limited to the information intention, that is, the intention of the user to obtain certain information, the operation intention, that is, the intention of the user to execute a certain operation, etc.
[0092] Agents are generally based on certain specific architectures, such as react, reflection, RAFA, etc. It can be software, hardware, or a system, with autonomy, adaptability, and interaction capabilities. The agent perceives changes in the environment, makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal. In this embodiment, the agent can understand the user's intention based on the user's operation and voice signal, and take corresponding actions or provide relevant information according to the user's intention.
[0093] In some embodiments, the agent can be composed of a large model, a functional module containing tools, a memory module, etc. Among them, the memory is used to store sufficient context to provide memory for the large model. The large model is responsible for analyzing the input information and planning an execution plan that meets the user's intention. Then, it calls the tool according to the execution plan. The tool has the ability to search or process information. After obtaining the output result of the tool, it is sent to the large model, and the large model determines whether the output result meets the user's needs. If it meets, it is fed back to the display device. If not, the above steps are repeated until a result that meets the user's needs is obtained or the maximum number of loops is reached.
[0094] In some embodiments, the large model can be a large language model. This requires the user to express their needs clearly every time. If the user does not express clearly and lacks some important information, the feedback result of the agent will be inaccurate. Therefore, in this embodiment, by combining the display content of the display area corresponding to the user's operation and the language signal emitted by the user. For example, when the user is watching TV and outputs "Introduce this star", and at the same time selects the specific position of the star or makes a circle selection, then the agent can clearly know the user's intention target, and the feedback result will be more accurate.
[0095] Exemplarily, Figure 6 shows the relationship diagram among the agent, the voice signal, and the control device. Refer to Figure 6, after the controller obtains the user operation triggered by the user through the control device and the voice signal issued by the user, it can send the user operation and the voice signal to the intelligent agent, and the intelligent agent can perform user intention recognition based on the user operation and the voice signal. Specifically, the controller can input the content file corresponding to the display content of the display area hit by the user operation and the text information of the voice signal into the intelligent agent. The large model in the intelligent agent can perform user intention recognition based on the content file and the text information to obtain the user operation intention, and at the same time plan an execution plan that meets the user operation intention. Then, based on the execution plan, an appropriate function module is called, and by using the tools in the function module, information search or processing is realized to obtain an output result. The large model can further determine whether the output result meets the user operation intention. If it meets, the output result is fed back to the display device. If it does not meet, the plan is re-planned until a result that meets the user operation intention is obtained or the maximum number of loops is reached.
[0096] In this embodiment, on the one hand, by combining the user operation triggered by the user through the control device and the user voice signal, and using the intelligent agent technology, user intention recognition is performed, so as to obtain the user operation intention. Compared with the current method of user intention recognition that can only be carried out under specific conditions, this solution can be applied to scenarios other than specific conditions, improving the generality of user intention recognition. On the other hand, the above control device can respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface. In this way, the user can quickly and accurately select the intention target in the user interface by using the control device. Compared with the traditional remote control, the user no longer needs to select the intention target by pressing each key one by one, effectively improving the accuracy and efficiency of the user operation, and thus improving the accuracy and efficiency of user intention recognition.
[0097] In some embodiments, the controller is further configured to: obtain the operation trajectory data of the user operation from the interaction data sent by the control device.
[0098] Among them, the operation trajectory data may refer to the trajectory data when the user initiates a user operation through the control device, and may include but is not limited to direction, speed, acceleration, etc.
[0099] Exemplarily, when the control device receives the operation initiated by the user, it can collect the trajectory data generated during the user operation and send the trajectory data to the display device. For use when there is a matching voice signal for the user operation, the display device maps the trajectory data into the user interface to form the cursor trajectory of the moving cursor in the user interface.
[0100] Take the pointing remote control as an example. In some embodiments, when the user physically moves the pointing remote control to perform operations such as circle selection on the user interface, the pointing remote control will collect the trajectory data generated during the user operation, specifically, it can be the movement trajectory of the pointing remote control itself. Exemplarily, the gyroscope in the pointing remote control will collect the movement trajectory data of the pointing remote control and send it to the display device.
[0101] In some embodiments, if the user triggers an operation through the touch screen on the pointing remote control, the controller or control chip in the pointing remote control can directly send the coordinate information of the operation in the touch screen to the display device, and the display device can convert it into interface coordinates based on the mapping relationship between the touch screen coordinates and the user interface coordinates. The trajectory formed by the interface coordinates is the cursor trajectory of the moving cursor.
[0102] In some embodiments, when determining the movement trajectory of the moving cursor on the user interface when the emission time of the voice signal matches the operation time of the user operation, the controller is further configured to: when the emission time of the voice signal matches the operation time of the user operation, convert the operation trajectory data into a set of coordinates in the user interface; determine the target trajectory formed by the set of coordinates in the user interface, and use the target trajectory as the movement trajectory of the moving cursor on the user interface.
[0103] Among them, the set of coordinates may include at least one interface coordinate point. The target trajectory refers to the trajectory formed by each coordinate point in the set of coordinates in the user interface.
[0104] Exemplarily, when the emission time of the voice signal matches the operation time of the user operation, the controller can perform coordinate mapping on the operation trajectory data. Specifically, it maps each trajectory point included in the operation trajectory data to the user interface, so as to obtain the interface coordinate points of each trajectory point in the user interface, and the set composed of these interface coordinate points is the set of coordinates. After the controller obtains the set of coordinates, it can further determine the trajectory formed by each interface coordinate point in the set of coordinates in the user interface, and use this trajectory as the movement trajectory of the moving cursor on the user interface.
[0105] Optionally, when there is no matching voice signal for the user operation, the display device may not process the operation trajectory data.
[0106] In this embodiment, by mapping the operation trajectory data of the user operation to the user interface, the absolute coordinates of the user operation on the user interface can be obtained. In this way, the display area selected by the user in the user interface can be accurately positioned, the positioning accuracy of the display area can be improved, and thus the accuracy of the display content can be improved, laying a foundation for the intelligent agent to accurately recognize the user's intention based on the display content.
[0107] In some embodiments, during the process of determining the display area corresponding to the movement trajectory, the controller is further configured to: determine the initial area corresponding to the movement trajectory; expand the initial area under the condition that the initial area meets the area expansion condition to obtain an expanded target area; and use the target area as the display area hit by the movement trajectory.
[0108] The initial area herein refers to the original area hit by the movement trajectory. It should be noted that in practical applications, there may be situations where the user's selected range is too small or the user fails to select a complete object, which will affect the integrity and accuracy of the content displayed within the display area, thereby affecting the accuracy of the intelligent agent's intention recognition. Based on this, this embodiment introduces an area expansion operation to adaptively expand the area that meets the area expansion condition to ensure the integrity and accuracy of the area content.
[0109] The area expansion condition refers to the pre-set prerequisite for area expansion. The target area refers to the expanded area. It can be understood that area expansion can also be regarded as a process of complete target segmentation. Therefore, when the area has been expanded to include the complete target or content, the expansion can be stopped, which is applicable to scenarios with complex interface elements and non-fixed sizes. Or it can be stopped when the expansion reaches a specified size, which is applicable to scenarios with simple interface elements and fixed sizes.
[0110] Optionally, the area expansion condition can be that the area of the initial area is smaller than a preset area threshold, which means the initial area is too small and needs to be expanded. The preset area threshold is the pre-set area threshold, and when it is lower than this threshold, the area is considered too small, such as 1% of the entire user interface area. Of course, in practical applications, an appropriate area threshold can also be defined according to the actual situation, and this embodiment does not limit this.
[0111] Optionally, the area expansion condition can also be whether the display content within the initial area is complete. If it is not complete, area expansion is required. Edge detection can be used to determine whether the display content is complete, that is, perform edge detection on the display content within the initial area. If the detection result shows that the content edges are discontinuous, it means the content is incomplete; if the detection result shows that the content edges are continuous, it means the content is complete. It should be noted that the area expansion condition is not limited to the above two points, and in practical applications, appropriate area expansion conditions can be set according to the actual situation.
[0112] Exemplarily, the controller can first determine the initial area hit by the movement trajectory. Then, it determines whether the initial area meets the area expansion condition, such as whether the area of the initial area is less than a preset area threshold, or whether the display content within the initial area is complete. If the area of the initial area is less than the preset area threshold, or the display content within the initial area is incomplete, the initial area is expanded. The expansion stops until the specified area size is reached or the complete content is included. The area at this time is the display area hit by the final movement trajectory.
[0113] In this embodiment, by expanding the area corresponding to the user operation in the display area, the integrity and accuracy of the display content in the display area are ensured, thereby guaranteeing the accuracy of the user intention recognition by the intelligent agent based on the display content.
[0114] In some embodiments, the display content includes the object type of the target object in the display area; during the process of the controller performing speech recognition on the display content and voice signals corresponding to the movement trajectory and performing user intention recognition through the intelligent agent, the controller is further configured to: based on the object type of the target object in the display area corresponding to the movement trajectory, screen out the target function module that matches the object type from each candidate function module for identifying user intentions; based on the object type and the speech recognition result of the voice signal, call the target function module through the intelligent agent to perform user intention recognition.
[0115] Among them, the function module can refer to a module that contains the tools required for the intelligent agent to perform user intention recognition, and can be used to implement the functions of information search or processing. Each function module can be configured with one tool. The target function module refers to the function module finally called by the intelligent agent, and the target function module can be one or more.
[0116] For example, if the target object is a person, then it is more inclined to configure a function module that includes tools such as person introduction, person movie and TV query, and person award query. If the target object is a flower, then it is more inclined to configure a function module that includes tools such as flower introduction and flower cultivation.
[0117] It should be noted that the tools of the agent can be understood as the capabilities of the agent. Different tool configurations can enable the agent to have different capabilities. For example, if a search tool is configured for the agent, then the agent can autonomously call the search tool to obtain relevant information from a third-party platform. Considering that if a large number of tools are configured for the agent, there will be many combinations between tools when it comes to user intent recognition, that is, there are multiple task planning paths. This also means that the agent may execute multiple rounds of tool calls to achieve the user intent. To a certain extent, it will cause problems such as a long user intent recognition process and high recognition costs. Therefore, in this embodiment, it is selected to match the function modules, that is, tools, according to the object type, so as to improve the adaptability of the tools and overcome the problems of too long recognition time and too high recognition costs caused by the agent calling too many tools.
[0118] Exemplarily, after the controller identifies the object type of the target object, it can screen out the target function modules that match the object type from each candidate function module. Then add these target function modules to the agent configuration and update the agent. By using the updated agent to call the tools in the target function module, the user intent recognition can be achieved.
[0119] In some embodiments, for user intent recognition under specific conditions, in order to ensure the executability of the task, personalized tool configuration may not be performed, but all tools are configured into the agent.
[0120] In some embodiments, in addition to the method of screening function modules based on the object type, in practical applications, screening can also be performed according to business requirements.
[0121] In some embodiments, in the process of screening out the target function modules that match the object type from each candidate function module for identifying user intent, the controller is further configured to: respectively obtain the function attribute information of each candidate function module for identifying user intent; respectively match each function attribute information with the object type to obtain the matching results between each function attribute information and the object type; and use the candidate target function modules with matching results as the target function modules.
[0122] Among them, the function attribute information may refer to the relevant attribute information of the candidate function module, which may include but is not limited to the tool name, function description, etc. included. The matching result can be used to represent the matching degree between the function attribute information and the object type, and the matching result can include matching and non-matching.
[0123] Optionally, the corresponding matching level (high matching degree, medium matching degree, low matching degree) may be further determined according to the matching degree between each functional attribute information and the object type, and the functional module with a high matching degree may be used as the target functional module.
[0124] Exemplarily, for each candidate functional module, the controller may respectively obtain the functional attribute information of each candidate functional module, such as the tool name and function description it contains, so as to analyze the matching degree between the candidate functional module and the object type based on this functional attribute information.
[0125] For example, assume that an object type is a person, and the function description of a candidate functional module is person introduction, then the two match. If an object type is an animal, and the function description of a candidate functional module is person introduction, then the two do not match.
[0126] In this embodiment, by screening out the functional modules that match the object type, the intelligent agent can be provided to call the tools in the functional module to perform user intention recognition. The adaptability of the functional module or the tool to the object type is improved, thereby improving the accuracy of user intention recognition.
[0127] In some embodiments, during the process of detecting the object type of the target object in the display area, the controller is further configured to: perform target detection on the display area to obtain a target detection frame in the display area; perform type recognition on the target object in the target detection frame to obtain the object type of the target object.
[0128] Among them, target detection refers to the process of detecting the target of interest in the display area and marking the bounding box of the target of interest, that is, the target detection frame. Type recognition refers to the process of recognizing the type of the target object in the target detection frame.
[0129] In some embodiments, after obtaining the display area, the controller may extract the area image of the display area. The controller can then perform target detection on the area image to detect the target of interest in the area image and mark the corresponding detection frame. Then, type recognition is performed on the target object in the detection frame to obtain the object type of the target object.
[0130] Optionally, object detection can be implemented using a deep learning model for object detection, such as the YOLO model, the EfficientDet model, etc. Object type recognition can be implemented using an image recognition model, such as a deep convolutional neural network model, a long short-term memory network model, etc. It should be noted that object detection and object type recognition can be implemented using one model, that is, the object detection task and the object type recognition task can be embedded in one model. Of course, two models can also be used to implement the two tasks separately. It can be specifically determined according to actual needs.
[0131] In some embodiments, for the content file corresponding to the display content, the content file may include not only the object type of the target object, but also the image of the target object.
[0132] Optionally, if there are multiple object detection frames in the display area, that is, the user may circle multiple intended objects. The object type of each target object in each object detection frame can be recognized, and the recognized object type and the image of the target object are stored in the content file. The agent performs user intention recognition based on each content file to obtain the operation intention of the user for each target object.
[0133] Optionally, in the case where the user circles multiple intended objects, the controller can also display the target objects in each object detection frame to the user and ask the user to confirm which one or more targets the intention recognition is specifically for. After receiving the confirmation instruction from the user, object type recognition is performed. This can avoid performing intention recognition on the targets that the user has selected multiple times or incorrectly selected, thereby improving the accuracy and efficiency of user intention recognition.
[0134] In this embodiment, object detection is performed on the display area to obtain the object detection frame within the display area. Then, the object type of the target object within the object detection frame is recognized to improve the accuracy of object type recognition, thereby improving the accuracy of subsequent function module matching for identifying user intentions based on the object type.
[0135] In some embodiments, when the controller executes the speech recognition result of the display content and the speech signal corresponding to the movement trajectory and performs user intention recognition through the agent, it is further configured to: perform interaction intention detection based on the display content and the speech recognition result of the speech signal corresponding to the movement trajectory to obtain an interaction intention detection result; in the case where the interaction intention detection result indicates that the user has an interaction intention, perform user intention recognition through the agent.
[0136] Among them, interactive intent detection can be understood as an operation to detect whether the user has an interactive intent. The interactive intent detection result is used to represent whether the user has an interactive intent. When the user has an interactive intent, the agent then performs user intent recognition to obtain the user's operation intent.
[0137] Exemplarily, the controller can first determine whether to trigger the wake-free interaction of the display device based on the display content corresponding to the movement trajectory and the speech recognition result of the voice signal. If triggered, it is considered that the user has an interactive intent; if not triggered, it is considered that the user does not have an interactive intent. Specifically, the controller can input the display content and the voice signal into the wake-free engine, and the wake-free engine allows the user to trigger the interaction without using a specific wake word.
[0138] In some embodiments, a pre-trained interactive intent detection model is deployed in the wake-free engine. The interactive intent detection model extracts the features of the display area (such as coordinate features, object features of the target object, etc.) and the text features of the speech recognition result, which is the text information, to detect the user's interactive intent. If the output result is 1 or other identifiers indicating that the user has an interactive intent, it can be considered that the user has an interactive intent. If the output result is 0 or other identifiers indicating that the user does not have an interactive intent, it can be considered that the user does not have an interactive intent. Among them, the interactive intent detection model can be a binary classification model.
[0139] Optionally, the output result can also be wake-up and non-wake-up. When the output result is wake-up, it is considered that the user has an interactive intent. When the output result is non-wake-up, it is considered that the user does not have an interactive intent.
[0140] In this embodiment, by determining whether the user has an interactive intent and using the agent to perform user intent recognition on the premise that the user has an interactive intent, the situation of invalid recognition by the agent can be avoided, thereby improving the efficiency of the agent's user intent recognition.
[0141] In some embodiments, the interactive intent detection result further includes the complexity of intent recognition; the process of controlling the user intent recognition by the agent is further configured to: when the complexity of intent recognition is not lower than the complexity threshold, perform user intent recognition by the agent deployed in the cloud; when the complexity of intent recognition is lower than the complexity threshold, perform user intent recognition by the agent deployed locally.
[0142] Among them, the complexity refers to the difficulty level of user intention recognition, which can be represented as the complexity score of user intention recognition or the complexity level. For example, for some user intentions that only involve single-step operations, such as "open the application", it can be considered that the complexity is relatively low. For some user intentions that involve multi-step operations, such as "search for all TV dramas and movies of this star in the past year", this involves search operations for "star", "in the past year", "TV drama", and "movie", and the complexity is relatively high. Of course, this is only an example. In actual applications, appropriate complexity judgment rules can be defined according to the actual situation, and this embodiment does not limit this.
[0143] The complexity threshold is a threshold set in advance for the complexity of user intention recognition. If it is not lower than this threshold, it can be considered that the complexity is high, and then the user intention recognition is performed by an agent deployed in the cloud. Since the agent deployed in the cloud relies on cloud computing resources, it has powerful computing capabilities and supports the recognition of complex intentions. When it is lower than this threshold, it can be considered that the complexity is low, and then the user intention recognition is performed by an agent deployed locally. The agent deployed locally, that is, on the display device side, relies on limited local computing resources and is suitable for lightweight intention recognition tasks.
[0144] In this embodiment, by reasonably allocating the intention recognition tasks of the cloud agent and the local agent according to the complexity of intention recognition, the efficiency of user intention recognition can be effectively improved.
[0145] In some embodiments, during the process of determining the speech recognition result of the speech signal, the controller is further configured to: extract the signal features of the speech signal; based on the signal features, perform validity detection on the speech signal to obtain the validity detection result of the speech signal; in the case where the validity detection result indicates that the speech signal belongs to a valid speech signal, perform speech text recognition on the speech signal to obtain text information matching the speech signal.
[0146] Among them, the signal features refer to the features of the speech signal, which can include but are not limited to features such as the energy, spectrum, and fundamental frequency of the signal. The signal features can be used to determine whether the speech signal is a valid speech signal or an invalid speech signal. A valid speech signal refers to a human voice signal, that is, the speech issued by the user. An invalid speech signal refers to a noise signal such as background sound or environmental noise. The validity detection result can be used to indicate whether the speech signal belongs to a valid speech signal, and can include that the speech signal belongs to a valid speech signal and that the speech signal does not belong to a valid speech signal.
[0147] Exemplarily, after receiving the voice signal sent by the microphone, the controller may first perform a validity detection on the voice signal, that is, detect whether the voice signal belongs to a valid voice signal. Specifically, the signal features in the voice signal can be extracted through VAD (Voice Activity Detection) technology, and then based on the signal features, it can be determined whether the voice signal is valid. In the case that the voice signal belongs to a valid voice signal, the voice signal is further subjected to text recognition to obtain text information. Herein, the text information refers to the text converted from the voice signal.
[0148] In this embodiment, by performing a validity detection on the voice signal, the valid voice signal and the invalid voice signal can be quickly distinguished, thereby improving the quality of the voice signal, and further improving the accuracy of voice recognition and the accuracy of user intention recognition.
[0149] In a specific embodiment, Figure 7 shows an interaction diagram among various modules during the user intention recognition process. Specifically, the user initiates a user operation through a control device, and at the same time, the microphone receives the user's voice signal. The controller can obtain the operation trajectory data of the user operation from the interaction data sent by the control device, and in the case that the user operation matches the voice signal, convert the operation trajectory data into a set of coordinates in the user interface. Then, the target trajectory formed by each coordinate point in the coordinate set in the user interface is determined, and this target trajectory is used as the movement trajectory of the moving cursor in the user interface. Based on the display content of the display area corresponding to the movement trajectory and the voice recognition result of the voice signal, the user intention is recognized through an agent to obtain the user operation intention.
[0150] Figure 8 shows a schematic diagram of the user intention recognition scenario. Figure 8 As shown, the user specifically uses a pointing remote control to circle a display area in the user interface of the display device through the control device, and at the same time, the user emits a voice signal. The content file corresponding to the display area (including the object type of the target object in the display area) and the voice signal enter the wake-free engine of the display device for judging the user interaction intention. In the case that it is determined that the user has an interaction intention, the wake-free engine sends the display content and the text information of the voice signal to the agent, and the user intention is recognized through the agent. The agent can feedback the recognized user operation intention to the display device, and the display device can display it by means of speaker playback, text display, action execution, etc.
[0151] Figure 9 shows a schematic flowchart of the user intention recognition. Figure 9As shown, the user triggers a user operation by pointing at the remote control while initiating speech. The display device performs VAD detection on the speech signal to determine whether the speech signal is a valid speech signal. In the case where the speech signal is a valid speech signal, the speech signal and the user operation are processed. Specifically, it includes identifying the text information of the speech signal and obtaining the content file corresponding to the user operation, where the content file includes the object type and image of the target object within the display area hit by the user operation.
[0152] In some embodiments, the operation trajectory data of the user operation is converted into a set of coordinates in the user interface, and the set of coordinates contains at least one absolute coordinate. The area formed by these coordinates in the user interface is the display area. In the case where the display area meets the area expansion condition, it can be expanded. The object type is obtained by performing target detection on the display area to obtain the target detection box within the display area, and then identifying the type of the target object within the target detection box.
[0153] In some embodiments, the text information and the content file of the user operation enter the wake-free engine of the display device. An interaction intention detection model is deployed in the wake-free engine to detect whether the user has an interaction intention. In the case where the user has an interaction intention, user intention recognition is performed through an agent.
[0154] In some embodiments, the agent tool configuration can be based on the object type to configure tools that match the object type in the agent framework. In this way, the agent can call these tools for user intention recognition. The output result of the agent can be stored in the context module for the interaction intention detection model to call.
[0155] It should be noted that this embodiment is compatible with traditional user intention recognition solutions. That is, the user's speech signal can still be sent to the keyword wake-up engine. If the speech signal contains a keyword, it wakes up, indicating that the user has an interaction intention. At this time, the agent tool configuration can also be performed. The agent tools in the traditional solution do not need to be personalized configured, and all tools can be configured into the agent.
[0156] In some embodiments, if not woken up but there is a speech signal, the text information of the speech signal can still be recognized and stored in the context module. If there is no speech signal but there is a user operation, the display device may not perform user intention recognition.
[0157] In this embodiment, on the one hand, by combining the user operation triggered by the user through the control device with the user voice signal and using the agent technology, user intention recognition is performed to obtain the user operation intention. Compared with the current method of user intention recognition that can only be performed under specific conditions, this solution can be applied to scenarios other than specific conditions, improving the generality of user intention recognition and thus enhancing the user experience. On the other hand, the above control device can respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface. In this way, the user can quickly and accurately select the intention target in the user interface using this control device. Compared with the traditional remote control, the user no longer needs to select the intention target by pressing each key one by one, effectively improving the accuracy and efficiency of the user operation and thus enhancing the accuracy and efficiency of user intention recognition.
[0158] In some embodiments, the present application further provides a user intention recognition method, which is applied to the above display device. In this embodiment, as Figure 10 shown, the user intention recognition method includes the following steps:
[0159] Step S1002, when the emission time of the voice signal emitted by the user matches the operation time of the user operation triggered by the user through the control device, determine the movement trajectory of the moving cursor on the user interface; the control device is used to respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface;
[0160] Step S1004, based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the voice signal, perform user intention recognition through the agent to obtain the user operation intention.
[0161] In some embodiments, the method further includes: obtaining the operation trajectory data of the user operation; when the emission time of the voice signal matches the operation time of the user operation, determining the movement trajectory of the moving cursor on the user interface further includes: when the emission time of the voice signal matches the operation time of the user operation, converting the operation trajectory data into a coordinate set in the user interface; determining the target trajectory formed by the coordinate set in the user interface, and taking the target trajectory as the movement trajectory of the moving cursor on the user interface.
[0162] In some embodiments, determining the display area corresponding to the movement trajectory further includes: determining the initial area corresponding to the movement trajectory; when the initial area meets the area expansion condition, expanding the initial area to obtain the expanded target area; taking the target area as the display area hit by the movement trajectory.
[0163] In some embodiments, the display content includes the object type of the target object in the display area; based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the speech signal, the user intention recognition by the intelligent agent further includes: based on the object type of the target object in the display area corresponding to the movement trajectory, screening out the target function module that matches the object type from each candidate function module for identifying the user intention; based on the object type and the speech recognition result of the speech signal, calling the target function module by the intelligent agent to perform user intention recognition.
[0164] In some embodiments, detecting the object type of the target object in the display area further includes: performing target detection on the display area to obtain a target detection frame in the display area; performing type recognition on the target object in the target detection frame to obtain the object type of the target object.
[0165] In some embodiments, screening out the target function module that matches the object type from each candidate function module for identifying the user intention includes: respectively obtaining the function attribute information of each candidate function module for identifying the user intention; respectively matching each function attribute information with the object type to obtain the matching result between each function attribute information and the object type; using the candidate target function module with a matching result indicating a match as the target function module.
[0166] In some embodiments, the user intention recognition by the intelligent agent based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the speech signal includes: performing interaction intention detection based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the speech signal to obtain an interaction intention detection result; in the case where the interaction intention detection result indicates that the user has an interaction intention, performing user intention recognition by the intelligent agent.
[0167] In some embodiments, the interaction intention detection result further includes the complexity of intention recognition; the user intention recognition by the intelligent agent further includes: in the case where the complexity of intention recognition is not lower than the complexity threshold, performing user intention recognition by the intelligent agent deployed in the cloud; in the case where the complexity of intention recognition is lower than the complexity threshold, performing user intention recognition by the intelligent agent deployed locally.
[0168] In some embodiments, the speech recognition result includes text information matching the speech signal; the process of determining the speech recognition result of the speech signal further includes: extracting the signal features of the speech signal; based on the signal features, performing validity detection on the speech signal to obtain the validity detection result of the speech signal; in the case where the validity detection result indicates that the speech signal is a valid speech signal, performing speech text recognition on the speech signal to obtain text information matching the speech signal.
[0169] In the user intention recognition method of this embodiment, on the one hand, by combining the user operations triggered by the user through the control device and the user voice signal, and using the agent technology, user intention recognition is performed to obtain the user operation intention. Compared with the current method of user intention recognition that can only be carried out under specific conditions, this solution can be applied to scenarios other than specific conditions, improving the generality of user intention recognition. On the other hand, the above control device can respond to user operations and control the moving cursor on the user interface to move to any position on the user interface. In this way, the user can quickly and accurately select the intention target in the user interface using this control device. Compared with the traditional remote control, the user no longer needs to select the intention target by pressing each key one by one, effectively improving the accuracy and efficiency of user operations, and thus improving the accuracy and efficiency of user intention recognition.
[0170] In some embodiments, the present application further provides a user intention recognition device, which is applied to the above display device. In this embodiment, the user intention recognition device includes:
[0171] A cursor trajectory determination module, configured to determine the movement trajectory of the moving cursor on the user interface when the emission time of the voice signal emitted by the user matches the operation time of the user operation triggered by the control device; the control device is configured to respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface;
[0172] A user intention recognition module, configured to perform user intention recognition through an agent based on the display content of the display area corresponding to the movement trajectory and the speech recognition result of the voice signal to obtain the user operation intention.
[0173] In some embodiments, the device is further configured to: obtain the operation trajectory data of the user operation from the interaction data sent by the control device; the cursor trajectory determination module is further configured to: when the emission time of the voice signal matches the operation time of the user operation, convert the operation trajectory data into a coordinate set in the user interface; determine the target trajectory formed by the coordinate set in the user interface, and use the target trajectory as the movement trajectory of the moving cursor on the user interface.
[0174] In some embodiments, the device is further configured to: determine the initial area corresponding to the movement trajectory; under the condition that the initial area meets the area expansion condition, expand the initial area to obtain the expanded target area; use the target area as the display area hit by the movement trajectory.
[0175] In some embodiments, the displayed content includes the object type of the target object within the display area; the user intention recognition module is further configured to: based on the object type of the target object within the display area corresponding to the movement trajectory, screen out a target function module that matches the object type from each candidate function module for identifying the user intention; and based on the object type and the speech recognition result of the speech signal, call the target function module through an agent to perform user intention recognition.
[0176] In some embodiments, the device is further configured to: perform target detection on the display area to obtain a target detection frame within the display area; and perform type recognition on the target object within the target detection frame to obtain the object type of the target object.
[0177] In some embodiments, the device is further configured to: respectively obtain the function attribute information of each candidate function module for identifying the user intention; respectively match each function attribute information with the object type to obtain the matching result between each function attribute information and the object type; and use the candidate target function module with a matching result indicating a match as the target function module.
[0178] In some embodiments, the user intention recognition module is further configured to: perform interaction intention detection based on the displayed content of the display area corresponding to the movement trajectory and the speech recognition result of the speech signal to obtain an interaction intention detection result; and in the case where the interaction intention detection result indicates that the user has an interaction intention, perform user intention recognition through an agent.
[0179] In some embodiments, the interaction intention detection result further includes the complexity of intention recognition; the user intention recognition module is further configured to: in the case where the complexity of intention recognition is not lower than a complexity threshold, perform user intention recognition through an agent deployed in the cloud; and in the case where the complexity of intention recognition is lower than the complexity threshold, perform user intention recognition through an agent deployed locally.
[0180] In some embodiments, the speech recognition result includes text information matching the speech signal; the device is further configured to: extract the signal features of the speech signal; based on the signal features, perform validity detection on the speech signal to obtain a validity detection result of the speech signal; and in the case where the validity detection result indicates that the speech signal is a valid speech signal, perform speech text recognition on the speech signal to obtain text information matching the speech signal.
[0181] Each module in the above user intention recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0182] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method steps are implemented.
[0183] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method steps are implemented.
[0184] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random-access memories (ReRAMs), magnetoresistive random-access memories (MRAMs), ferroelectric random-access memories (FRAMs), phase-change memories (PCMs), graphene memories, etc. Volatile memories can include random-access memories (RAMs) or external caches, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random-access memory (SRAM) or dynamic random-access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0185] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0186] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A display device, characterized in that, Including: A display configured to display a user interface; A communication module configured to communicate with a control device; The control device is used to respond to a user operation and control a moving cursor on the user interface to move to any position on the user interface; A microphone configured to receive a voice signal emitted by a user; A controller configured to: Determine a moving trajectory of the moving cursor on the user interface when an emission time of the voice signal matches an operation time of the user operation; Based on display content of a display area corresponding to the moving trajectory and a speech recognition result of the voice signal, perform user intention recognition through an agent to obtain a user operation intention.
2. The display device according to claim 1, wherein The controller is further configured to: Obtain operation trajectory data of the user operation from interaction data sent by the control device; When determining the moving trajectory of the moving cursor on the user interface when the emission time of the voice signal matches the operation time of the user operation, the controller is further configured to: Convert the operation trajectory data into a coordinate set in the user interface when the emission time of the voice signal matches the operation time of the user operation; Determine a target trajectory formed by the coordinate set in the user interface, and use the target trajectory as the moving trajectory of the moving cursor on the user interface.
3. The display device according to claim 1, wherein When determining the display area corresponding to the moving trajectory, the controller is further configured to: Determine an initial area corresponding to the moving trajectory; Expand the initial area under the condition that the initial area meets an area expansion condition to obtain an expanded target area; Use the target area as the display area hit by the moving trajectory.
4. The display device according to claim 1, characterized in that, The display content includes an object type of a target object in the display area; When the controller performs user intention recognition through an agent based on the display content of the display area corresponding to the moving trajectory and the speech recognition result of the voice signal, the controller is further configured to: Based on the object type of the target object in the display area corresponding to the moving trajectory, screen out a target function module that matches the object type from each candidate function module for identifying user intentions; Based on the object type and the speech recognition result of the voice signal, call the target function module through the agent to perform user intention recognition.
5. The display device according to claim 4, wherein, When detecting the object type of the target object in the display area, the controller is further configured to: Perform target detection on the display area to obtain a target detection frame in the display area; Perform type recognition on the target object in the target detection frame to obtain the object type of the target object.
6. The display device according to claim 4, wherein When screening out a target function module that matches the object type from each candidate function module for identifying user intentions, the controller is further configured to: Respectively obtain function attribute information of each candidate function module for identifying user intentions; Match each of the function attribute information with the object type respectively to obtain the matching results between each of the function attribute information and the object type; Represent the matching results as the candidate target function modules that match, and use them as the target function modules.
7. The display device according to claim 1, wherein When the controller executes the display content corresponding to the display area of the movement trajectory and the speech recognition result of the speech signal, and performs user intention recognition through the agent, it is further configured as follows: Detect the interaction intention based on the display content corresponding to the display area of the movement trajectory and the speech recognition result of the speech signal, to obtain the interaction intention detection result; When the interaction intention detection result indicates that the user has an interaction intention, perform user intention recognition through the agent.
8. The display device according to claim 7, wherein The interaction intention detection result further includes the complexity of intention recognition; when the controller performs user intention recognition through the agent, it is further configured as follows: When the complexity of intention recognition is not lower than the complexity threshold, perform user intention recognition through the agent deployed in the cloud; When the complexity of intention recognition is lower than the complexity threshold, perform user intention recognition through the agent deployed locally.
9. The display device according to claim 1, wherein The speech recognition result includes the text information that matches the speech signal; when determining the speech recognition result of the speech signal, the controller is further configured as follows: Extract the signal features of the speech signal; Based on the signal features, perform validity detection on the speech signal to obtain the validity detection result of the speech signal; When the validity detection result indicates that the speech signal is a valid speech signal, perform speech text recognition on the speech signal to obtain the text information that matches the speech signal.
10. A method for user intention recognition, characterized in that, The method includes: When the emission time of the speech signal emitted by the user matches the operation time of the user operation triggered by the user through the control device, determine the movement trajectory of the moving cursor on the user interface; the control device is used to respond to the user operation and control the moving cursor on the user interface to move to any position on the user interface; Based on the display content corresponding to the display area of the movement trajectory and the speech recognition result of the speech signal, perform user intention recognition through the agent to obtain the user operation intention.
Citation Information
Cited By
Display device and circle selection screenshot method
CN121099129A