Display device and control method of display device
By using a semantic understanding model that combines intent recognition, slot filling, and default intent models in display devices, the problem of display devices being unable to accurately determine user intent is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202210917564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-08-01
AI Technical Summary
Existing display devices often provide default fallback responses after responding to user voice commands, failing to accurately determine the user's intent and thus reducing the user experience.
A semantic understanding model consisting of an intent recognition model, a slot filling model, and a default intent model is adopted to improve the accuracy of intent determination by recognizing and analyzing user voice commands.
It improves the accuracy of display devices in responding to user voice commands, reduces default fallback responses, and enhances the user experience.
Smart Images

Figure CN115273848B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and more particularly to a display device and a method for controlling the display device. Background Technology
[0002] With the rapid development of display devices, the functions they can provide to users are becoming increasingly diverse. Currently, display devices include televisions, set-top boxes, and other products with display screens. Taking televisions as an example, televisions are expanding their applications beyond simply watching television programs in the home; they can also be used for gaming, displaying electronic photo albums, and showcasing information.
[0003] Currently, many display devices have the ability to interact with users via voice assistants. However, it is often found that after responding to a user's voice command, the display device only provides a default, fallback response. For example, displaying "This problem is too difficult, Xiao X is still learning" on the display device. Such a response fails to meet the user's needs and reduces the user experience.
[0004] Therefore, how to more accurately determine the user's intent after the user outputs a voice command has become a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a display device and a control method for the display device. The method uses a semantic understanding model composed of an intent recognition model, a slot filling model, and a default intent model, which can more accurately determine the user's intent and improve the user experience.
[0006] In a first aspect, a display device is provided, comprising:
[0007] A monitor is used to display the user interface.
[0008] User interface, used to receive input signals;
[0009] The controllers, which are connected to the display and the user interface respectively, are configured as follows:
[0010] When a user inputs a voice command, the system recognizes the text content within the voice command.
[0011] The content text is input into a semantic understanding model, and the result to be used is output. The semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model that utilize a common encoding module. The result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model.
[0012] If the first confidence level is not greater than the first preset confidence level, and / or the second confidence level is not greater than the second preset confidence level, then the encapsulation result is determined based on the third confidence level and the third result, and corresponding operations are performed based on the encapsulation result.
[0013] In some embodiments, the controller, which performs the step of determining the encapsulation result based on the third confidence level and the third result, to perform the operation corresponding to the encapsulation result, is further configured to:
[0014] The third result includes search intent or non-search intent;
[0015] When the third result is a search intent, it is detected whether the third confidence level is greater than the third preset confidence level;
[0016] If the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent;
[0017] If the first result is a media asset search intent, then the encapsulated result is determined to include a search intent and a second result, so as to perform a search operation that uses the second result as the search object.
[0018] In some embodiments, the controller is further configured to: determine the encapsulation result as a default intent when the third result is not a search intent, or when the third confidence level is less than or equal to a third preset confidence level, so as to perform the operation of displaying default information.
[0019] In some embodiments, the controller is further configured to: if the first result is not a media asset search intent, determine that the encapsulated result includes a search intent and content text, and perform a search operation with the content text as the search object.
[0020] In some embodiments, the controller is further configured to:
[0021] If the first confidence level is greater than the first preset confidence level, and the second confidence level is greater than the second preset confidence level, then based on the first result and the second result, a packaging result is determined, and corresponding operations are performed based on the packaging result.
[0022] In some embodiments, the controller is further configured to save the content text when the third confidence level is less than or equal to a third preset confidence level.
[0023] In some embodiments, the controller is further configured to train the semantic understanding model using sample data.
[0024] In some embodiments, the controller, which performs training of the semantic understanding model using sample data, is further configured to:
[0025] The sample data includes first sample data and second sample data; the intent recognition model further includes an intent recognition classifier; the slot filling model further includes a slot filling classifier; the default intent model further includes a default intent classifier.
[0026] The intent recognition model and slot filling model are trained using the first sample data to determine and lock the parameters in the common encoding module, intent recognition classifier, and slot filling classifier; the default intent model of the common encoding module, including the locked parameters, is trained using the second sample data to determine the parameters of the default intent classifier.
[0027] Secondly, a method for controlling a display device is provided, comprising:
[0028] When a user inputs a voice command, the system recognizes the text content within the voice command.
[0029] The content text is input into a semantic understanding model, and the result to be used is output. The semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model that utilize a common encoding module. The result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model.
[0030] If the first confidence level is not greater than the first preset confidence level, and / or the second confidence level is not greater than the second preset confidence level, then the encapsulation result is determined based on the third confidence level and the third result, and corresponding operations are performed based on the encapsulation result.
[0031] In some embodiments, the step of determining the encapsulation result based on the third confidence level and the third result, and then performing the operation corresponding to the encapsulation result, includes:
[0032] The third result includes search intent or non-search intent;
[0033] When the third result is a search intent, it is detected whether the third confidence level is greater than the third preset confidence level;
[0034] If the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent;
[0035] If the first result is a media asset search intent, then the encapsulated result is determined to include a search intent and a second result, so as to perform a search operation that uses the second result as the search object.
[0036] The display device and control method provided in the above embodiments utilize a semantic understanding model composed of an intent recognition model, a slot filling model, and a default intent model to more accurately determine the user's intent and improve the user experience. The method includes: when a user-inputted voice command is received, recognizing the text content in the voice command; inputting the text content into the semantic understanding model and outputting a result to be used, wherein the semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model using a common encoding module; the result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model; if the first confidence level is not greater than a first preset confidence level, and / or the second confidence level is not greater than a second preset confidence level, then based on the third confidence level and the third result, a packaging result is determined to perform a corresponding operation based on the packaging result. Attached Figure Description
[0037] Figure 1 An operational scenario between a display device and a control device according to some embodiments is illustrated;
[0038] Figure 2 A hardware configuration block diagram of a control device 100 according to some embodiments is shown;
[0039] Figure 3 A hardware configuration block diagram of a display device 200 according to some embodiments is shown;
[0040] Figure 4 A software configuration diagram of a display device 200 according to some embodiments is shown;
[0041] Figure 5 An exemplary schematic diagram of a user interface provided according to some embodiments is shown;
[0042] Figure 6 An exemplary flowchart is shown for a control method of a display device according to some embodiments;
[0043] Figure 7 A schematic diagram of yet another user interface provided according to some embodiments is shown as an example;
[0044] Figure 8 An exemplary schematic diagram of the structure of a semantic understanding model provided according to some embodiments is shown;
[0045] Figure 9 An exemplary schematic diagram of yet another user interface provided according to some embodiments is shown;
[0046] Figure 10 A flowchart of another control method for a display device provided according to some embodiments is shown as an example. Detailed Implementation
[0047] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0048] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0049] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0050] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0051] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.
[0052] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.
[0053] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.
[0054] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.
[0055] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.
[0056] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.
[0057] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.
[0058] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0059] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0060] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.
[0061] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0062] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.
[0063] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.
[0064] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).
[0065] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0066] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.
[0067] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.
[0068] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0069] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can execute operations related to the object selected by the user command.
[0070] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.
[0071] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0072] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0073] See Figure 4 In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Layer"), and the kernel layer.
[0074] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs, system settings programs, or clock programs that come with the operating system; they may also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0075] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0076] like Figure 4 As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: ActivityManager, which interacts with all activities running in the system; LocationManager, which provides access to system location services for system services or applications; PackageManager, which retrieves various information related to application packages currently installed on the device; NotificationManager, which controls the display and clearing of notification messages; and WindowManager, which manages icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0077] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling display window changes (e.g., shrinking the display window, shaking the display, distorting the display, etc.).
[0078] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0079] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0080] In some embodiments, the display device has the function of interacting with the user through a voice assistant. However, it often fails to provide a default, fallback response after the display device responds to the user's voice command. For example, displaying "This problem is too difficult, Xiaox is still learning" on the display device fails to meet the user's needs and reduces the user experience. Figure 5 A schematic diagram of a user interface provided according to some embodiments is shown as an example. Figure 5 The user interface displays a default fallback response: "This question is too difficult; Xiao X is still learning."
[0081] To address the aforementioned technical problems, this application provides a control method for a display device. This method uses a semantic understanding model composed of an intent recognition model, a slot filling model, and a default intent model, which can more accurately determine the user's intent and improve the user experience.
[0082] Figure 6 A flowchart of a control method for a display device according to some embodiments is illustrated. The method includes:
[0083] S100. When a voice command input by the user is received, the content text in the voice command is recognized. In this embodiment of the application, the voice command is recognized as content text. The specific recognition process is not limited, and any method that can convert the voice command into content text is acceptable.
[0084] The voice commands described in this embodiment can be based on user voice input. For example, a user can say a wake-up word and the operation they wish the display device to perform to input a voice command. For instance, a user can say, "Xiao X, check today's weather," where "Xiao X" is the wake-up word and "check today's weather" is the operation the user wishes the display device to perform. In this case, the voice command is input. In another example, the user can press a voice input button on the control device and say the operation they wish the display device to perform to input a voice command.
[0085] In this application embodiment, there are various scenarios in which the display device can receive voice commands input by the user. For example, Figure 7 An exemplary schematic diagram of yet another user interface provided according to some embodiments is shown. Figure 7 The display device is playing a video, and while the video is playing, the display device can receive user commands.
[0086] S200. Input the content text into the semantic understanding model and output the result to be used.
[0087] Figure 8 An exemplary schematic diagram of the structure of a semantic understanding model provided according to some embodiments is shown. The semantic understanding model includes an intent recognition model, a slot-filling model, and a default intent model utilizing a common encoding module.
[0088] The results to be used include a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model.
[0089] The intent recognition model is used for intent recognition to obtain a first result and a first confidence level corresponding to the first result. The slot filling model is used for parameter extraction to obtain a second result and a second confidence level corresponding to the second result. The default intent model is used for default intent recognition to obtain a third result and a third confidence level corresponding to the third result. The third result includes search intent and no search intent.
[0090] Taking the voice command "what's the weather tomorrow in Berlin" as an example, the first result corresponding to the intent recognition model can be "weather.query" (weather, search intent), the second result corresponding to the slot filling model can be "{datetime:tomorrow,city:Berlin}" (date: tomorrow, city: Berlin), and the third result corresponding to the default intent model can be "non-search intent".
[0091] In this embodiment, the intent recognition model, slot filling model, and default intent model in the semantic understanding model share the same common encoding module. Since the encoder in the common encoding module is the most time-consuming component, in this embodiment, the intent recognition model, slot filling model, and default intent model share the same common encoding module, which can reduce the time used to run the semantic understanding model.
[0092] To more clearly illustrate the solutions of the embodiments of this application, the display device can be controlled according to the semantic understanding model shown below.
[0093] In this embodiment, the semantic understanding model mainly uses the LaBSE model (Language-agnostic BERTsentence Embedding) and a classifier to construct an intent recognition model, a slot filling joint model, and a default intent model to perform intent recognition, parameter extraction, and default intent recognition. The LaBSE model is the common encoding module.
[0094] In this embodiment, the LABSE model is trained using hundreds of billions of monolingual corpora and hundreds of millions of bilingual corpora, enabling it to encode words and sentences in hundreds of languages, and possessing powerful multilingual sentence and word representation capabilities. The BERT model is used as a pre-trained model, ultimately training a LaBSE model that includes an encoder and an embedding layer.
[0095] See again Figure 8 For the input text, it is first tokenized (lexicalized). A [CLS] token (global feature aggregation) is added at the beginning of the sentence, which is generally used to represent the sentence information of the entire text. Multiple [SEP] tokens (separators) are added at the end of the sentence to make the input length of the semantic understanding model the same. Then, each token is mapped to a tokenID (lexical encoding) so that the display device can recognize the text.
[0096] The token vocabulary using the LaBSE model in this embodiment contains hundreds of thousands of tokens, and most text content can be represented by combinations of these tokens. After passing through the token layer, the feature representation corresponding to each token is read from the pre-trained vocabulary and then input into the LaBSE model. After multiple layers of encoder interaction, the encoder can be a transformer encoder, outputting the final encoding result, which is the feature of each token, i.e., the output result of the LaBSE model.
[0097] In this embodiment, the output features are connected to the classifier for subsequent tasks.
[0098] Specifically, firstly, since the feature corresponding to [CLS] represents the features of the entire content text sentence, an intent classifier is added after this feature to classify the intent and output different intents, such as media asset search or display device control. For example, the intent classifier can be a softmax classification model.
[0099] The BIO labeling system is used for slot prediction, where B represents the beginning of a label, I represents the innermost part of a label, and O represents no label (other). Sample data is labeled using this BIO labeling system, and the model's prediction results are then made consistent with the BIO labels. Therefore, during slot prediction, a slot prediction classifier is added to the features of each remaining token for BIO label classification. For example, the slot prediction classifier can be a softmax classification model.
[0100] In this embodiment, after encoding the user request, the sentence representation vector [CLS] is concatenated with a default intent classifier to classify the content text into search intent and non-search intent. For example, the default intent classifier can be a softmax binary classifier layer.
[0101] The foregoing describes the structure of the semantic understanding model in detail. In this embodiment, the semantic understanding model needs to be trained before use. The specific method for training the semantic understanding model is as follows:
[0102] In some embodiments, the semantic understanding model is trained using sample data. In some embodiments, the sample data includes first sample data (D original ) and second sample data (D default The first sample data contains detailed grammatical information and explicit intent instructions. For example, "search for movie Spider-Man" contains grammatical results, clearly indicating the intent to search for a movie. The second sample data consists of the default intent output by the semantic understanding model, as well as data that cannot output execution instructions for controlling the display device. This data may lack syntactic meaning, be a meaningful named entity, or have an ambiguous intent that differs significantly from the current business intent. For example, if a user outputs the voice command "Jillian and Addie," and the display device shows "This problem is too difficult, Xiaox is still learning," then "Jillian and Addie" can be included as part of the second sample data. Additionally, the output data corresponding to "Jillian and Addie" can be included as another part of the second sample data, and this output data is manually labeled. In this embodiment, the output data can be labeled as either a search intent or a non-search intent.
[0103] In this embodiment, during the semantic understanding model training process, it is necessary to train the LaBSE model, the intent classifier, the slot prediction classifier, and the default intent classifier. Since both the first and second sample data update the parameters of the LaBSE model, a reasonable training method is needed to integrate the first and second sample data to make the trained model more closely resemble the real situation.
[0104] For the datasets containing the first and second sample data, the process of training a semantic understanding model can include the following three methods:
[0105] The first approach involves training a semantic understanding model by mixing the first and second sample data together.
[0106] When training by mixing the two datasets, the first and second sample data are mixed and fed into the semantic understanding model for training. The model parameters in the LaBSE model, intent classifier, slot prediction classifier, and default intent classifier are updated according to the error.
[0107] The second approach involves cross-training the first and second sample datasets separately, with the first sample dataset as the primary training data and the second sample dataset as a secondary data source. The training process follows this cross-training method: first, the LaBSE model, intent classifier, and slot prediction classifier are trained using the first sample dataset; then, the second sample dataset is used for fine-tuning. Because the LaBSE model uses a common encoding module, its learning rate should be set relatively small. Fine-tuning these parameters primarily involves modifying the parameters of the default intent classifier. This process is then repeated on both the first and second sample datasets until the model performs well on both datasets.
[0108] The third approach involves training the first and second sample data separately. The LaBSE model is trained strictly using the first sample data, with the parameters of the semantic understanding LaBSE model remaining unchanged. A fully functional LaBSE model, intent recognition model, and slot-filling model are trained on the first sample data. The parameters of the common encoding LaBSE model, intent recognition model, and slot-filling model are then frozen and not updated. Finally, the semantic understanding model is trained using the second sample data, updating only the parameters of the default intent classifier to obtain the default intent classifier.
[0109] In the training process of the semantic understanding model, the first method mixes the first sample data and the second sample data. However, because the amount of the first sample data and the second sample data are different, the difficulty of the classification task is different, and the relative size of the error is also different when calculating the error, the convergence speed of the semantic understanding model is slow. In the end, the intention recognition model and slot filling model in the semantic understanding model do not perform well in the final test.
[0110] In the second approach, by reducing the learning rate and performing five rounds of cross-training, the semantic understanding model's errors converged on both training sets. However, the intent recognition and slot-filling models performed poorly, showing slightly lower performance on the first dataset and overfitting on the second dataset. Analysis suggests that the intent recognition model requires 50 categories, while the slot-filling model, based on the BIO labeling system, requires 102 categories. In contrast, the default intent classifier only requires binary classification, making its task complexity far lower than the intent recognition and slot-filling models, leading to overfitting.
[0111] In the third approach, the performance of the intent recognition model and the slot filling model is not affected, and the default intent classifier is also suitable for binary classification tasks, ultimately achieving relatively ideal results. On the first sample dataset, with the LaBSE model parameters unchanged, the accuracy of the intent classifier is 95.89%, and the accuracy of the slot filling classifier is 91.02%. On the second sample dataset, the accuracy of the default intent classifier is 94.02%, and good generalization performance was observed during testing.
[0112] Therefore, in this embodiment, the intent recognition model and slot filling model are trained using the first sample data to determine and lock the parameters in the common encoding module, intent recognition classifier, and slot filling classifier; the default intent model of the common encoding module, including the locked parameters, is trained using the second sample data to determine the parameters of the default intent classifier. This ensures that the semantic understanding model can accurately and quickly identify the user's intent.
[0113] In this embodiment, the encoder is the main time-consuming part when running the speech understanding model. Therefore, the intent recognition model, slot filling model and default intent model use the same common encoding module, which includes the encoder. This can shorten the time spent using the speech understanding model and further determine the user's default intent, thereby improving the user experience.
[0114] In some embodiments, when the semantic understanding model is used, the intent recognition model, the slot filling model, and the default intent model output corresponding results and confidence scores, respectively. The results to be used include a first confidence score and a first result corresponding to the intent recognition model, a second confidence score and a second result corresponding to the slot filling model, and a third confidence score and a third result corresponding to the default intent model.
[0115] S300. Compare the first confidence level with the first preset confidence level, and compare the second confidence level with the second preset confidence level. In this embodiment, the confidence level can reflect the credibility of the model output result. The higher the confidence level, the higher the approximation of the output result with the true result; the lower the confidence level, the lower the approximation of the output result with the true result.
[0116] S400. If the first confidence level is greater than the first preset confidence level and the second confidence level is greater than the second preset confidence level, then the encapsulation result is determined based on the first result and the second result, and the corresponding operation is performed based on the encapsulation result.
[0117] In this embodiment, if the first confidence level is greater than the first preset confidence level, and the second confidence level is greater than the second preset confidence level, it indicates that the first result output by the intent recognition model and the true intent, as well as the second result output by the slot filling model and the true parameters, are very close. When both are close to the real situation, the first result and the second result can be directly determined as the encapsulation result, and the corresponding operation can be performed based on the encapsulation result. For example, if the second result is "{datetime:tomorrow,city:Berlin}" and the first result is "weather.query", then the search for and display of tomorrow's Berlin weather can be performed directly. Figure 9 An exemplary illustration shows yet another user interface diagram provided according to some embodiments, in Figure 9 The screen displays information about the Berlin weather.
[0118] In some embodiments, the result encapsulation process involves encapsulating the output of the semantic understanding model into a unified format to facilitate command execution by downstream display devices. For example, the intent is converted into an execution command for the display device, and the intent parameters are converted into a unified format. For instance, some fixed settings of the display device need to be converted into the display device's language, and time formats in different languages are converted into numeric formats that the terminal can parse, such as converting 2hours in English into {h:2,m:0,s:0}.
[0119] In some embodiments, the first preset reliability and the second preset reliability can be the same value, for example, 0.9. Of course, the first preset reliability and the second preset reliability can also be different values.
[0120] S500. If the first confidence level is not greater than the first preset confidence level, and / or the second confidence level is not greater than the second preset confidence level, then the encapsulation result is determined based on the third confidence level and the third result, and corresponding operations are performed based on the encapsulation result.
[0121] In this embodiment of the application, if the confidence level of the first result is not greater than the first preset confidence level and the confidence level of the second result is not greater than the second preset confidence level, it means that the operation performed by the display device may be inaccurate after directly using the first result and the second result as the encapsulation result. Therefore, in order to avoid the display device performing inaccurate operations, the encapsulation result is determined by using the third confidence level and the third result. This allows for further analysis of the voice commands corresponding to operations that the display device cannot provide to meet user needs, and determination of the intent corresponding to the voice commands. This can reduce the situation where the display device displays a default fallback response.
[0122] The default intent model in this application embodiment performs further analysis on the user intent in the voice command when the confidence of the first and second results is not high, thereby improving the user experience.
[0123] In some embodiments, Figure 10 An exemplary flowchart illustrates another control method for a display device according to some embodiments. S500, the step of determining a packaging result based on a third confidence level and a third result, and performing corresponding operations based on the packaging result, includes:
[0124] The third result includes either search intent or non-search intent.
[0125] S501. Determine whether the third result is a search intent or not. In this embodiment, since at least one confidence level corresponding to the intent recognition model and the slot filling model cannot meet the preset confidence level requirement, in order to further understand the user's needs, the third result output by the default intent model is used to determine whether it is a search intent. According to research, most display devices cannot provide voice commands corresponding to operations that meet user needs. Most of the time, users want to use the display device to search for text content in voice commands. For example, if a user says the name of an actor, it is highly likely that the user wants the display device to search for the actor's name. This embodiment utilizes this feature and uses the default intent model to determine whether the third result corresponding to the current text content is a search intent. If the third result is a search intent, based on the third confidence level, it is further determined how the display device should perform the search.
[0126] S502. When the third result is not a search intent, the encapsulation result is determined to be a default intent, and the operation of displaying default information is performed. In this embodiment, the default intent instructs the display device to display default information. For example, the default information could be "This problem is too difficult, Xiaox is still learning." In this embodiment, when the third result is not a search intent, the text content is determined to be unsuitable for searching using the display device to meet user needs.
[0127] S503. When the third result represents a search intent, detect whether the third confidence level is greater than a third preset confidence level. In this embodiment, if the third result represents a search intent, the third confidence level is detected to determine the reliability of the third result. When the relationship between the third confidence level and the third preset confidence level is different, it is determined that the display device performs the search in a different manner.
[0128] S504. If the third confidence level is less than or equal to the third preset confidence level, then the encapsulation result is determined to be the default intent, and the operation of displaying default information is performed. In this embodiment, if the third confidence level is not high, it indicates that the credibility of the third result as a search intent is not high, then the encapsulation result is determined to be the default intent, and the display device displays default information.
[0129] In addition, in some embodiments, when the third confidence level is less than or equal to the third preset confidence level, the content text is saved for subsequent manual analysis to uncover new business needs, etc. In this embodiment, the semantic understanding model can be continuously trained with a large amount of sample data. When the third confidence level is less than or equal to the third preset confidence level, the content text corresponding to the semantic understanding model is saved, and the content text can be analyzed to obtain new sample data for training the semantic understanding model.
[0130] S505. If the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent. In this embodiment, errors in the first and second confidence levels may occur, preventing the display device from directly performing corresponding operations using the first and second results. However, if the third confidence level is greater than the third preset confidence level, it indicates that the third result is highly credible as a search intent. To avoid errors in the first and second confidence levels, it is possible to further detect whether the first result is a media asset search intent. If it is a media asset search intent, the second result can be used for media asset search.
[0131] In some embodiments, a media asset search intent can be preset. For example, the preset media asset search intent can be video search (video.search) or video playback (video.play), etc. When the first result is a preset media asset search intent, the first result is determined to be a media asset search intent.
[0132] S506. If the first result is a media asset search intent, then determine that the encapsulated result includes a search intent and a second result, and perform a search operation using the second result as the search object. For example, the second result is today's weather, and the display device performs a search using today's weather as the search object.
[0133] It should be noted that the search scope is not limited in this embodiment of the application; it can be a local database, an external database, or other scopes.
[0134] S507. If the first result is not a media asset search intent, then determine that the encapsulated result includes a search intent and content text, and perform a search operation using the content text as the search object. For example, the content text could be xxx (person's name), and the display device searches using xxx as the search object.
[0135] This application embodiment also provides a display device, including: a display for displaying a user interface; a user interface for receiving input signals; and a controller connected to the display and the user interface respectively, configured to: when receiving a voice command input by a user, recognize the content text in the voice command; input the content text into a semantic understanding model, and output a result to be used, wherein the semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model utilizing a common encoding module; the result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model; if the first confidence level is not greater than a first preset confidence level, and / or the second confidence level is not greater than a second preset confidence level, then a packaging result is determined based on the third confidence level and the third result, so as to perform a corresponding operation based on the packaging result.
[0136] In some embodiments, the controller, in performing the step of determining the encapsulation result based on a third confidence level and a third result to perform an operation corresponding to the encapsulation result, is further configured to: the third result includes a search intent or a non-search intent; when the third result is a search intent, detect whether the third confidence level is greater than a third preset confidence level; if the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent; if the first result is a media asset search intent, determine that the encapsulation result includes a search intent and a second result to perform a search operation using the second result as the search object.
[0137] In some embodiments, the controller is further configured to: determine the encapsulation result as a default intent when the third result is not a search intent, or when the third confidence level is less than or equal to a third preset confidence level, so as to perform the operation of displaying default information.
[0138] In some embodiments, the controller is further configured to: if the first result is not a media asset search intent, determine that the encapsulated result includes a search intent and content text, and perform a search operation with the content text as the search object.
[0139] In some embodiments, the controller is further configured to save the content text when the third confidence level is less than or equal to a third preset confidence level.
[0140] In some embodiments, the controller is further configured to:
[0141] If the first confidence level is greater than the first preset confidence level, and the second confidence level is greater than the second preset confidence level, then based on the first result and the second result, a packaging result is determined, and corresponding operations are performed based on the packaging result.
[0142] In some embodiments, the controller is further configured to train the semantic understanding model using sample data.
[0143] In some embodiments, the controller, performing the training of the semantic understanding model using sample data, is further configured to: the sample data include first sample data and second sample data; the intent recognition model further includes an intent recognition classifier; the slot filling model further includes a slot filling classifier; the default intent model further includes a default intent classifier; the intent recognition model and the slot filling model are trained using the first sample data to determine and lock the parameters in the common encoding module, the intent recognition classifier, and the slot filling classifier; the default intent model, including the locked parameters, is trained using the second sample data to determine the parameters of the default intent classifier.
[0144] In the above embodiments, the display device and the control method for the display device use a semantic understanding model composed of an intent recognition model, a slot filling model, and a default intent model, which can more accurately determine the user's intent and improve the user experience. The method includes: when a user inputs a voice command, recognizing the content text in the voice command; inputting the content text into the semantic understanding model, and outputting a result to be used, wherein the semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model utilizing a common encoding module; the result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model; if the first confidence level is not greater than a first preset confidence level, and / or the second confidence level is not greater than a second preset confidence level, then based on the third confidence level and the third result, a packaging result is determined, and a corresponding operation is performed based on the packaging result.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
[0146] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, include: A monitor is used to display the user interface. User interface, used to receive input signals; The controllers, which are connected to the display and the user interface respectively, are configured as follows: When a user inputs a voice command, the system recognizes the text content within the voice command. The content text is input into a semantic understanding model, and the output is the result to be used. The semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model. The intent recognition model, the slot filling model, and the default intent model have a common encoding module. The result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model. If the first confidence level is not greater than the first preset confidence level, and / or the second confidence level is not greater than the second preset confidence level, then the encapsulation result is determined based on the third confidence level and the third result, and corresponding operations are performed based on the encapsulation result. The encapsulation result represents a data structure with a unified format formed after the result output by the semantic understanding model is encapsulated.
2. The display device according to claim 1, characterized in that, The third result includes search intent or non-search intent; the controller, which performs the step of determining the encapsulation result based on the third confidence level and the third result, and then performs the operation corresponding to the encapsulation result, is further configured to: When the third result is a search intent, it is detected whether the third confidence level is greater than the third preset confidence level; If the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent; If the first result is a media asset search intent, then the encapsulated result is determined to include a search intent and a second result, so as to perform a search operation that uses the second result as the search object.
3. The display device according to claim 2, characterized in that, The controller is also configured to: When the third result is not a search intent, or when the third confidence level is less than or equal to the third preset confidence level, the encapsulation result is determined to be a default intent, and the operation of displaying default information is performed.
4. The display device according to claim 2, characterized in that, The controller is also configured to: If the first result is not a media asset search intent, then the encapsulated result is determined to include a search intent and content text, in order to perform a search operation with the content text as the search object.
5. The display device according to claim 2, characterized in that, The controller is also configured to: When the third confidence level is less than or equal to the third preset confidence level, the content text is saved.
6. The display device according to claim 1, characterized in that, The controller is also configured to: If the first confidence level is greater than the first preset confidence level, and the second confidence level is greater than the second preset confidence level, then the encapsulation result is determined based on the first result and the second result, and corresponding operations are performed based on the encapsulation result.
7. The display device according to claim 1, characterized in that, The controller is also configured to: Obtain sample data; The semantic understanding model is trained using sample data.
8. The display device according to claim 7, characterized in that, The sample data includes first sample data and second sample data; the intent recognition model further includes an intent recognition classifier; the slot filling model further includes a slot filling classifier; the default intent model further includes a default intent classifier; the controller, which executes training of the semantic understanding model using the sample data, is further configured to: The intent recognition model and slot filling model are trained using the first sample data to determine and lock the parameters in the common encoding module, the intent recognition classifier, and the slot filling classifier. The default intent model, which includes a common encoding module with locked parameters, is trained using the second sample data to determine the parameters of the default intent classifier.
9. A control method for a display device, characterized in that, include: When a user inputs a voice command, the system recognizes the text content within the voice command. The content text is input into a semantic understanding model, and the output is the result to be used. The semantic understanding model includes an intent recognition model, a slot filling model, and a default intent model. The intent recognition model, the slot filling model, and the default intent model have a common encoding module. The result to be used includes a first confidence level and a first result corresponding to the intent recognition model, a second confidence level and a second result corresponding to the slot filling model, and a third confidence level and a third result corresponding to the default intent model. If the first confidence level is not greater than the first preset confidence level, and / or the second confidence level is not greater than the second preset confidence level, then based on the third confidence level and the third result, a packaging result is determined, and corresponding operations are performed based on the packaging result. The packaging result represents a data structure with a unified format formed after packaging the output of the semantic understanding model.
10. The control method according to claim 9, characterized in that, The step of determining the packaging result based on the third confidence level and the third result, and then performing the operation corresponding to the packaging result, includes: The third result includes search intent or non-search intent; When the third result is a search intent, it is detected whether the third confidence level is greater than the third preset confidence level; If the third confidence level is greater than the third preset confidence level, detect whether the first result is a media asset search intent; If the first result is a media asset search intent, then the encapsulated result is determined to include a search intent and a second result, so as to perform a search operation that uses the second result as the search object.
Citation Information
Patent Citations
Semantic recognition method and device
CN110309514A
Natural language understanding processing
US11335346B1