Display device and specification query method

By implementing voice interaction and semantic understanding models on the display device, identifying the user's voice query intention and generating answers, the problem of high threshold for use of traditional electronic manual query methods is solved, and higher query accuracy and interactivity are achieved.

CN120011530APending Publication Date: 2025-05-16VIDAA (NETHERLANDS) INT HLDG LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411887234.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional electronic manual query methods rely on text search or directory navigation, resulting in a high threshold for use, especially for users who are not familiar with professional terms, visual impairment or low operating proficiency.

Method used

Through a display device and corresponding query method, using voice interaction and semantic understanding models, users can receive voice inquiries, identify query intentions, locate and generate query answers, and display them to users.

Benefits of technology

It simplifies user input operations, reduces the burden on both eyes and hands, realizes automatic connection between fuzzy input and accurate search, improves the accuracy and interactivity of query results, and lowers the threshold for use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011530A_ABST
    Figure CN120011530A_ABST
Patent Text Reader

Abstract

The invention relates to a display device and a specification query method. The display equipment comprises a display, a sound collector, an audio output device and a controller, the controller is configured to control the sound collector to receive user inquiry voice; an intention recognition result corresponding to the user inquiry voice is obtained, and the intention recognition result is obtained by conducting specification inquiry intention recognition on the user inquiry voice through a semantic understanding model; on the basis of the intention recognition result, under the condition that it is determined that the target specification query intention is recognized, an inquiry answer corresponding to the user inquiry voice is obtained, and the inquiry answer is obtained by locating target text content from a preset specification text based on the target specification query intention through a specification text processing model. Generating based on the target text content and the user inquiry voice; and controlling at least one of the display and the audio output device to display the inquiry answer. By adopting the display equipment, the use threshold of the specification query function can be improved and reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device and a method for querying instructions. Background Art

[0002] With the rapid development of digital technology, traditional paper manuals have gradually been replaced by electronic manuals due to their shortcomings such as large size, inconvenience in updating, and difficulty in retrieval.

[0003] In traditional technology, in order to facilitate users to read electronic manuals, many layout designs are made for the display page of the electronic manual, such as refined directory design, introduction of search elements, etc., so that users can quickly find the required information through keywords.

[0004] However, traditional electronic manual query methods still rely on text search or directory navigation, which can still be difficult for users who are unfamiliar with professional terminology, unfamiliar with interface operations, or have visual impairments, and the threshold for use is high. Summary of the Invention

[0005] The present application provides a display device and a manual query method to solve the problem of a high threshold for using the manual query function.

[0006] In a first aspect, some embodiments provide a display device comprising: a display, a sound collector, an audio output device, and a controller. The display is configured to display a user interface; the sound collector is configured to receive external sound; the audio output device is configured to output sound; and the controller is configured to:

[0007] a display configured to display a user interface;

[0008] A sound collector is configured to receive external sound;

[0009] an audio output device configured to output sound;

[0010] The controller is configured as:

[0011] Control the sound collector to receive user inquiry voice;

[0012] Obtain the intent recognition result corresponding to the user's inquiry voice. The intent recognition result is obtained by performing instruction manual query intent recognition on the user's inquiry voice through a semantic understanding model;

[0013] If the target manual query intent is determined based on the intent recognition result, the query answer corresponding to the user's query voice is obtained. The query answer is generated based on the target text content and the user's query voice by locating the target text content from the preset manual text based on the target manual query intent through the manual text processing model;

[0014] At least one of a display and an audio output device is controlled to display an answer to the query.

[0015] Technical effects: First, by receiving user inquiry voice, the operation of searching for text content that meets one's own needs from dense instruction manual text is reduced, the user input operation is simplified, the input device is freed from the constraints of the hands, the burden on the eyes is reduced, and the threshold for using the instruction manual query function is lowered; further, through the semantic understanding of the user inquiry voice, fuzzy input on the user side can be realized, and then the fuzzy input on the user side is converted into an instruction manual query intention that can be used for instruction manual text search, thereby realizing the automatic connection between the fuzzy input on the user side and the precise search of the instruction manual text. In this way, it can not only ensure that the inquiry answer has a high accuracy, but also reduce the user's memory of some professional terms in the instruction manual, further lowering the threshold for using the instruction manual query function; further, based on the user inquiry voice, the target text content is converted into the inquiry answer and the inquiry answer is displayed, which can not only further filter out the information in the target text content that is irrelevant to the user inquiry voice and improve the accuracy of the feedback information, but also because the question and answer format is closer to the user's natural communication habits, it can also improve the interactivity of the instruction manual query function.

[0016] In a second aspect, some embodiments further provide a method for querying a specification, the method comprising:

[0017] Receive user inquiry voice;

[0018] Obtain the intent recognition result corresponding to the user's inquiry voice. The intent recognition result is obtained by performing instruction manual query intent recognition on the user's inquiry voice through a semantic understanding model;

[0019] If the target manual query intent is determined based on the intent recognition result, the query answer corresponding to the user's query voice is obtained. The query answer is generated based on the target text content and the user's query voice by locating the target text content from the preset manual text based on the target manual query intent through the manual text processing model;

[0020] Display the answers to the inquiries.

[0021] Technical effects: First, by receiving user inquiry voice, the operation of searching for text content that meets one's own needs from dense instruction manual text is reduced, the user input operation is simplified, the input device is freed from the constraints of the hands, the burden on the eyes is reduced, and the threshold for using the instruction manual query function is lowered; further, through the semantic understanding of the user inquiry voice, fuzzy input on the user side can be realized, and then the fuzzy input on the user side is converted into an instruction manual query intention that can be used for instruction manual text search, thereby realizing the automatic connection between the fuzzy input on the user side and the precise search of the instruction manual text. In this way, it can not only ensure that the inquiry answer has a high accuracy, but also reduce the user's memory of some professional terms in the instruction manual, further lowering the threshold for using the instruction manual query function; further, based on the user inquiry voice, the target text content is converted into the inquiry answer and the inquiry answer is displayed, which can not only further filter out the information in the target text content that is irrelevant to the user inquiry voice and improve the accuracy of the feedback information, but also because the question and answer format is closer to the user's natural communication habits, it can also improve the interactivity of the instruction manual query function. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;

[0024] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;

[0025] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;

[0026] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;

[0027] Figure 5 A timing diagram of implementing a manual query function through a display device provided in some embodiments of the present application;

[0028] Figure 6 A timing diagram for implementing the instruction manual query function through a display device and a server in some embodiments of the present application;

[0029] Figure 7A schematic diagram of a scenario illustrating the interaction between a display device, a cloud platform, and an algorithm service provided in some embodiments of the present application;

[0030] Figure 8 A flowchart of a streaming text data transmission process provided in some embodiments of the present application;

[0031] Figure 9 A timing diagram of implementing a manual query function through a display device provided in some other embodiments of the present application;

[0032] Figure 10 A flowchart of a method for implementing a specification query is provided for some embodiments of the present application. DETAILED DESCRIPTION

[0033] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.

[0034] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0035] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0036] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0037] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.

[0038] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0039] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.

[0040] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling connection and communication via a network communication protocol, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.

[0041] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0042] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.

[0043] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.

[0044] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0045] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.

[0046] In some embodiments, the display 260 includes a display component for displaying images and a driver component for driving image display. The display 260 is configured to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces.

[0047] In some embodiments, the communication device 220 is a component for communicating with the first external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.

[0048] The communication device 220 can establish a communication connection between the display device 200 and the first external device or server 400 via a wireless or wired connection. A wired connection can use components such as a data cable and an interface to display the display device 200 and the personalized recommendation reasons. A wireless connection can use wireless signals or a wireless network to display the display device 200 and the personalized recommendation reasons. The display device 200 can establish a direct connection with the first external device or an indirect connection via a gateway, router, or connection device.

[0049] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200.

[0050] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0051] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).

[0052] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.

[0053] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.

[0054] Figure 3 Some embodiments of this application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0055] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .

[0056] In some embodiments, the control device 100 may be a smart device. For example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.

[0057] In some embodiments, as Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .

[0058] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and external and internal data processing functions.

[0059] Under the control of the controller 110, the communication interface 130 communicates control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, or other near field communication modules.

[0060] The user input / output interface 140 includes at least one of a microphone 141 , a touch panel 142 , a sensor 143 , a button 144 and other input interfaces.

[0061] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, to encode user input commands via the WiFi protocol, Bluetooth protocol, or NFC protocol and transmit them to the display device 200.

[0062] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0063] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0064] To facilitate user interaction, in some embodiments, display device 200 may run an operating system. An operating system is a computer program used to manage and control the hardware and software resources of display device 200. The operating system can provide a user interface (control the display device), allow users to interact with display device 200, and support the running of various application programs.

[0065] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for display devices.

[0066] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, Figure 4 Some embodiments of this application provide Figure 1 Schematic diagram of software configuration in the display device. In some embodiments, the system of the display device 200 can be divided into three layers, namely, application layer, middleware layer and hardware layer from top to bottom.

[0067] The application layer mainly includes commonly used applications on the TV and the application framework. Common applications are mainly browser-based applications, such as HTML5 apps, and native apps.

[0068] The Application Framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interfaces of these functions (toolbars, status bars, menus, dialog boxes).

[0069] Native apps can support online or offline, message push or local resource access.

[0070] The middleware layer includes various television protocols, multimedia protocols, and system components. Middleware uses the basic services (functions) provided by system software to connect various parts of application systems or different applications on the network, enabling resource and function sharing.

[0071] The hardware layer primarily includes the Hardware Abstraction Layer (HAL) interface, hardware, and drivers. The HAL interface serves as a unified interface for all TV chipsets, while the specific logic is implemented by each chip. Drivers primarily include audio drivers, display drivers, Bluetooth drivers, camera drivers, Wi-Fi drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, and pressure sensors), and power drivers.

[0072] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.

[0073] With the rapid development of digital technology, traditional paper manuals are gradually being replaced by electronic manuals due to their bulk, inconvenience in updating, and difficulty in searching. Traditionally, to facilitate user access to electronic manuals, numerous layout designs have been implemented for the display pages of electronic manuals, such as refined table of contents and the introduction of search elements, allowing users to quickly find the information they need by using keywords.

[0074] However, on the one hand, with the diversification of the functions of display devices, more and more content needs to be elaborated in the electronic manual. The difficulty of searching from the massive information is inherently high, and the electronic manual uses a large number of professional terms or technical language, which may be difficult for non-professionals to understand, further increasing the difficulty of information search; on the other hand, different models of products and different versions of systems have corresponding electronic manuals that are different. The format, interface and operation method of the electronic manual may be different, and users need to adapt to different usage methods. At present, many display devices have integrated online update functions. The display device itself and the electronic manual will be updated more frequently. Users need to adapt to and understand new functions more frequently. The frequency of use of the electronic manual will also increase accordingly, and the complexity and inconvenience of the operation of the electronic manual will become more prominent. Especially for the elderly or users who are not technically proficient, operating a complex electronic manual may cause difficulties.

[0075] In some embodiments, a display device is provided, the display device comprising:

[0076] a display configured to display a user interface;

[0077] A sound collector is configured to receive external sound;

[0078] an audio output device configured to output sound;

[0079] The controller is configured as:

[0080] Control the sound collector to receive user inquiry voice;

[0081] Obtain the intent recognition result corresponding to the user's inquiry voice. The intent recognition result is obtained by performing instruction manual query intent recognition on the user's inquiry voice through a semantic understanding model;

[0082] If the target manual query intent is determined based on the intent recognition result, the query answer corresponding to the user's query voice is obtained. The query answer is generated based on the target text content and the user's query voice by locating the target text content from the preset manual text based on the target manual query intent through the manual text processing model;

[0083] At least one of a display and an audio output device is controlled to display an answer to the query.

[0084] Among them, the instruction manual may refer to a document provided by the manufacturer to the user, which aims to help the user understand and correctly use the display device, usually including one or more instructions for the installation and use of each hardware structure of the display device, instructions for the use of each software function of the display device, common problems and their solutions, and maintenance instructions.

[0085] A manual query intent refers to a user's desire to obtain specific information from the manual when interacting with the system. Different manual query intents represent different information users desire from the manual. A manual query intent at least includes a clear query target, which represents the user's query task orientation. The query target can be used to match and locate target text content within the pre-set manual text that meets the manual query intent.

[0086] User inquiry speech may refer to the verbal expression used by a user to ask a question or request to a display device via voice interaction. It is understood that scenarios in which a user interacts with a display device through voice include, but are not limited to, manual querying. Therefore, the user inquiry speech may or may not contain the intent to query for manuals. When the user inquiry speech contains the intent to query for manuals, the user inquiry speech may include at least one of query request information and problem scenario description information. Query request information may refer to the user's expressed need for specific information, including at least one of the desired content, desired level of detail, format, and presentation method. Problem scenario description information may refer to a collection of information describing the problematic situation, problematic event, or problematic usage of the display device. If the user's purpose is clear, they can instruct the display device to query for manuals by transmitting query request information. If the purpose is unclear, they can also instruct the display device to query for manuals by transmitting problem scenario description information. The user inquiry speech may or may not contain the same or similar keywords as the manual text.

[0087] Intent recognition results refer to the output of a semantic understanding model after parsing a specific instruction manual query intent from a user's voice query. Intent recognition results can include no instruction manual query intent found, a target instruction manual query intent found, or multiple target instruction manual query intents found. They can also include model predictions such as the confidence level of the target instruction manual query intent. The target instruction manual query intent refers to the instruction manual query intent parsed from the user's voice query.

[0088] A semantic understanding model can refer to a natural language processing model used to identify the intent of a manual query. It aims to use computer algorithms and machine learning methods to analyze the surface structure of words or sentences, as well as the semantic level behind text or speech, to capture the speaker's true intent, emotion, and contextual information, thereby inferring the manual query intent contained in the user's query speech. The semantic understanding model can be obtained by collecting training sample data based on actual needs and training the natural language processing model. For example, historical query speech and the intent recognition labels corresponding to each historical query speech can be collected. Then, the natural language processing model can be initialized and the historical query speech can be input into the natural language processing model to perform manual query intent recognition and obtain intent recognition training results. Based on the difference between the intent recognition training results and the intent recognition labels, the model loss is calculated. Based on the model loss, the natural language processing model is iteratively optimized using backpropagation until the calculated model loss converges or the number of iterations reaches a preset maximum number. The trained natural language processing model is then determined as a semantic understanding model. The semantic understanding model can be deployed on a display device, cloud platform, or server to perform manual query intent recognition on user query speech and obtain intent recognition results.

[0089] The inquiry answer can refer to the response or solution provided based on the instruction text in response to the question or request raised in the user's inquiry voice. Specifically, the instruction text processing model can be used to locate the target text content from the preset instruction text based on the target instruction query intent, and generate it based on the target text content and the user's inquiry voice.

[0090] The target text content may refer to a portion of the text content in the preset specification text that matches the target specification query intent.

[0091] A manual text processing model may refer to a pre-trained language model used to generate answers to queries. The manual text processing model can be pre-trained on a large general corpus and then fine-tuned using a large amount of manual text data. The fine-tuning process may include collecting manual text data from smart devices in various languages ​​and annotating the data. The annotated information may include device functions, setting options, and corresponding scenario descriptions. The manual text data may then be cleaned and pre-processed. The cleaning and pre-processing process may include at least one of noise removal, word segmentation, stop word removal, and lemmatization. The pre-processed manual text data may then be divided into a training set and a test set. The training set is used to fine-tune the pre-trained language model to learn how to map manual query intent to specific text content. The test set is used to evaluate the performance of the pre-trained language model. Evaluation metrics include at least one of accuracy, recall, and F1 score (the harmonic mean of precision and recall). Based on the evaluation results, model parameters and training strategies are adjusted to further improve the accuracy and robustness of the model. The fine-tuned pre-trained language model is then determined as the manual text processing model. The instruction text processing model can be deployed on a display device, cloud platform or server to locate the target text content from the preset instruction text based on the target instruction query intent, and generate a query answer corresponding to the user query voice based on the target text content and the user query voice.

[0092] For example, Figure 5 As shown, when a user uses a display device, he or she can interact with the display device by voice according to his or her needs. During the voice interaction, the controller of the display device can control the sound collector to collect the user's inquiry voice, and input the user's inquiry voice into the semantic understanding model. The semantic understanding model is used to identify the instruction manual query intention of the user's inquiry voice to obtain the intention recognition result. Then, the target instruction manual query intention can be extracted from the intention recognition result. When the target instruction manual query intention is extracted, the target instruction manual query intention is input into the instruction manual text processing model. The instruction manual text processing model is used to match and locate the target text content corresponding to the target instruction manual query intention from the preset instruction manual text, and generate the inquiry answer corresponding to the user's inquiry voice based on the target text content. Then, at least one output device such as a display and an audio output device can be controlled to display the inquiry answer corresponding to the user's inquiry voice to the user.

[0093] Among them, the semantic understanding model and the instruction text processing model can be deployed on the display device or the server.

[0094] In this embodiment, first, by receiving the user's inquiry voice, the operation of searching for text content that meets one's own needs from the dense instruction manual text is reduced, the user input operation is simplified, the input device is freed from the constraints of the hands, the burden on the eyes is reduced, and the threshold for using the instruction manual query function is lowered; further, by understanding the semantics of the user's inquiry voice, fuzzy input on the user side can be realized, and then the fuzzy input on the user side is converted into an instruction manual query intention that can be used for instruction manual text search, thereby realizing the automatic connection between the fuzzy input on the user side and the precise search of the instruction manual text. In this way, it can not only ensure that the inquiry answer has a high accuracy, but also reduce the user's memory of some professional terms in the instruction manual, further lowering the threshold for using the instruction manual query function; further, based on the user's inquiry voice, the target text content is converted into an inquiry answer and the inquiry answer is displayed, which can not only further filter out information in the target text content that is irrelevant to the user's inquiry voice and improve the accuracy of the feedback information, but also improve the interactivity of the instruction manual query function because the question and answer format is closer to the user's natural communication habits.

[0095] In some embodiments, the intent recognition result includes multiple alternative specification query intents and corresponding confidence prediction values ​​for each alternative specification query intent; before obtaining the query answer corresponding to the user's query voice, the controller is further configured to:

[0096] Get the current scene information of the display device;

[0097] According to the current scenario information and each confidence prediction value, the target specification query intent is determined from each alternative specification query intent.

[0098] It should be noted that after training, the semantic understanding model can, to a certain extent, more accurately identify the instruction manual query intent contained in the user's inquiry voice. However, due to the diversity of product models, software versions, system modes, usage scenarios, etc., every function or every operation may not be possible in every model, every version, every mode, and every usage scenario. The semantic understanding model infers the target instruction manual query intent based on the user's inquiry voice. In this way, if the query answer corresponding to the identified target instruction manual query intent cannot be achieved in the current scenario, the user still cannot solve his actual problem after knowing the query answer.

[0099] Among them, the current scene information may at least include current usage scene information and current device scene information, wherein the current usage scene information may refer to a set of information used to describe the scenario in which the user interacts with the display device, such as a live viewing scenario, a children's mode scenario, a customer service inquiry scenario, etc.; the current device scene information may refer to a combination of information used to characterize the status of the display device itself, such as the software version, the start and stop status of the functional module, etc.

[0100] The alternative instruction manual query intent may refer to the instruction manual query intent parsed by the semantic understanding model from the user's query speech. The confidence prediction value may refer to the probability that the semantic understanding model determines that the user's query speech contains the alternative instruction manual query intent.

[0101] For example, after obtaining the intent recognition result, if the intent recognition result contains multiple alternative specification query intentions, each alternative specification query intention and the confidence prediction value of each alternative specification query intention can be extracted from the intent recognition result, and each alternative specification query intention can be sorted from high to low according to the confidence prediction value; then the current scene information of the display device is obtained, and each alternative specification query intention can be filtered out and re-sorted according to the current scene information. From the alternative specification query intentions that have not been filtered out, the one with the highest ranking priority is selected as the target specification query intention.

[0102] In this embodiment, by further screening the model prediction results in combination with the current scenario information, the target specification query intention that is suitable for the current actual scenario can be selected, the output of invalid information can be reduced, and the effectiveness of the query answer can be improved.

[0103] In some embodiments, in the process of determining the target specification query intent from each candidate specification query intent based on the current scenario information and each confidence prediction value, the controller is further configured to:

[0104] According to the correspondence between the preset scenario and the instruction manual query intent, determining the alternative instruction manual query intent that does not match the current scenario information;

[0105] Filter out the alternative instruction manual query intents that do not match the current scenario information to obtain the primary instruction manual query intent;

[0106] According to the current scenario information and each confidence prediction value, the target instruction manual query intent is determined from each preliminary instruction manual query intent.

[0107] It should be noted that due to the universality of model training, it is difficult to adapt to the actual needs of different scenarios. During the manual query process, the user's actual manual query intention or the manual query intention predicted by the model may not be feasible in the current scenario. For example, the user wants to inquire about how to set the screen saver function, but the display device or the current version does not support the screen saver function. The manual query intention that is not available in the current scenario may not be able to query the corresponding target text content, or may query incorrect target text content, or even if the correct target text content is queried, the user will ultimately be unable to complete the corresponding operation on the display device. Therefore, it will not only lead to unnecessary waste of resources, but may also provide invalid or erroneous information to the user, resulting in a poor user experience.

[0108] For example, after obtaining the current scene information, the correspondence between the preset scenes and the specification query intent can be queried to determine whether each alternative specification query intent matches the current scene information, retain the alternative specification query intent that matches the current scene information, and filter out the alternative specification query intent that does not match the current scene information, and determine the retained alternative specification query intent as the preliminary specification query intent; then, the various alternative specification query intents can be re-sorted and processed according to the current scene information, and then, from the preliminary specification query intent, the one with the highest priority ranking is selected as the target specification query intent.

[0109] In this embodiment, by filtering out alternative specification query intentions that do not match the current scenario, it is possible to effectively avoid performing subsequent specification query operations according to specification query intentions that cannot be realized in the current scenario, reduce resource waste, and avoid situations where the query answers cannot be realized in the current scenario, thereby reducing the output of invalid information and improving the effectiveness of the query answers.

[0110] In some embodiments, in the process of determining the target specification query intent from the preliminary specification query intents based on the current scenario information and the confidence prediction values, the controller is further configured to:

[0111] Based on the current scenario information, query the intent weight corresponding to each preliminary instruction manual query intent;

[0112] According to the weight of each intention, the confidence prediction value corresponding to each preliminary specification query intention is adjusted to determine the confidence target value corresponding to each preliminary specification query intention;

[0113] The preliminary specification query intent corresponding to the highest confidence target value is determined as the target specification query intent.

[0114] It should be noted that, due to the universality of model training, it is difficult to adapt to the actual needs of different scenarios. During the manual query process, there may be user inquiry voices that can represent multiple similar manual query intentions, but these similar manual query intentions are applicable to different actual scenarios. For example, the user's inquiry voice is how to optimize the sound. For display devices with a recap function, you can query the method for automatically setting the surround sound recap. For display devices that do not have a recap function but can adjust the sound mode, you should query the method for setting the sound mode. Based on the manual query intention that is not applicable to the current scenario, the user may not be able to complete the corresponding operation on the display device for the target text content that is retrieved. Therefore, it will not only lead to unnecessary waste of resources, but may also provide users with invalid or erroneous information, resulting in a poor user experience.

[0115] Among them, the intention weight is used to characterize the importance of each instruction manual query intention in the current scenario. It can be determined based on the correlation between the instruction manual query intention and the current scenario information, or it can be determined in advance based on actual conditions or test results, etc. This embodiment does not limit this.

[0116] For example, based on the current scenario information, the intention weight corresponding to each preliminary specification query intention can be queried first, and then, for each preliminary specification query intention, its corresponding intention weight is multiplied by its corresponding confidence prediction value to obtain the confidence target value of the preliminary specification query intention; further, after calculating the confidence target value corresponding to each preliminary specification query intention, the confidence target value with the highest numerical value is determined, and the preliminary specification query intention corresponding to the highest numerical confidence target value is determined as the target specification query intention.

[0117] As an example, based on the current scene information, querying the intention weights corresponding to each preliminary specification query intention may include: determining the intention weight data set corresponding to the current scene information by querying the correspondence between the preset scene and the intention weight data set, the intention weight data set may refer to a data set composed of intention weights, and then, from the intention weight data set corresponding to the current scene information, querying the intention weights corresponding to each preliminary specification query intention.

[0118] As another example, based on the current scene information, querying the intention weight corresponding to each preliminary specification query intention may include: detecting the correlation coefficient between each preliminary specification query intention and the current scene information, and linearly mapping each correlation coefficient to an intention weight, wherein the correlation coefficient detection method can adopt at least one of an artificial intelligence model or a mathematical model, etc., and this embodiment does not limit this.

[0119] In this embodiment, by reordering the alternative specification query intentions to adapt to the current actual scenario, it is possible to effectively avoid performing subsequent specification query operations according to specification query intentions that cannot be realized in the current scenario, reduce resource waste, and avoid situations where the query answers cannot be realized in the current scenario, thereby reducing the output of invalid information and improving the effectiveness of the query answers.

[0120] In some embodiments, the semantic understanding model and the instruction text processing model are deployed on the server; the display device further includes a communication device configured to communicate with the server;

[0121] In the process of obtaining the intent recognition result corresponding to the user's query voice, the controller is further configured to:

[0122] Controlling the communication device to send the user's inquiry voice to the server, so that the semantic understanding model deployed on the server can perform instruction manual query intent recognition on the user's inquiry voice and obtain an intent recognition result;

[0123] Control the communication device to receive the intent recognition result returned by the server in response to the user's query voice.

[0124] It should be noted that the semantic understanding model and the instruction manual text processing model have high demands on computing resources. When the display device has sufficient computing power, deploying the two on the display device can effectively save communication resources between the display device and the server, reduce communication time, and thus improve the efficiency of instruction manual query. However, when the display device has insufficient computing power, not only will the semantic understanding model and the instruction manual text processing model have slower computing speeds or even fail to run, but it will also affect the normal operation of other functions on the display device.

[0125] For example, Figure 6 As shown, when the semantic understanding model is deployed on the server, after receiving the user inquiry voice, the display device can control the communication device deployed on the display device, send the user inquiry voice and the instruction manual query intention recognition instruction to the server, and receive the intention recognition result returned by the server based on the user inquiry voice, wherein the instruction manual query intention recognition instruction is used to instruct the server to perform instruction manual query intent recognition on the user inquiry voice through the semantic understanding model, obtain the intent recognition result, and return the obtained intent recognition result to the display device; then, the display device can receive the intention recognition result returned by the server for the user inquiry voice.

[0126] In the process of obtaining the query answer corresponding to the user's query voice, the controller is further configured to:

[0127] Controlling the communication device to transmit the target manual query intent to the server, so that the manual text processing model deployed on the server locates the target text content from the preset manual text based on the target manual query intent, and generates a query answer based on the target text content and the user's query voice;

[0128] The communication device is controlled to receive a query answer returned by the server in response to the query intention of the target specification.

[0129] For example, Figure 6 As shown, when the instruction manual text processing model is deployed on the server, after determining the target instruction manual query intention, the display device can control the communication device deployed on the display device, send the user inquiry voice, intention recognition result and inquiry answer generation instruction to the server, and receive the inquiry answer returned by the server based on the user inquiry voice and intention recognition result, wherein the inquiry answer generation instruction is used to instruct the server to locate the target text content from the preset instruction manual text based on the target instruction manual query intention through the instruction manual text processing model, and generate an inquiry answer based on the target text content and the user inquiry voice, and return the generated inquiry answer to the display device; then, the display device can receive the inquiry answer returned by the server for the user inquiry voice and the target instruction manual query intention.

[0130] In this embodiment, by deploying the semantic understanding model and the instruction manual text processing model on the server side, on the one hand, the computing burden of the display device can be reduced and the normal operation of each function on the display device can be ensured; on the other hand, the reasoning speed of the model can be improved, thereby improving the efficiency of instruction manual query.

[0131] In some embodiments, the communication device is configured to establish a communication connection with a cloud platform and communicate with a server through the cloud platform.

[0132] It should be noted that different models may be deployed on the same or different servers. Due to the different nature of data requirements, the interaction methods used between display devices and different servers or different models are different. For example, Figure 7As shown, the user inquiry voice is first converted into user inquiry voice text through the ASR (Automatic Speech Recognition) service, and then the instruction manual query intent is identified through the NLP (Natural Language Processing) service. The inquiry answer is then generated through the instruction manual text processing model. The semantic understanding model can be an NLP model, and an NLP model is deployed on the NLP service. The ASR service requires full-duplex communication, so WebSocket (a full-duplex communication protocol) can be used. The NLP service only requires lightweight single-time communication, so the HTTP (HyperText Transfer Protocol) request / response mode can be used. For display devices, other communications with higher security levels can be used.

[0133] Among them, cloud platform can refer to the infrastructure, platform or software that provides computing resources and services through the Internet, allowing display devices to access and use these resources on demand without the need to install and maintain corresponding software or systems on local hardware.

[0134] For example, the cloud platform can be used as a transfer point, and the display devices interact with the cloud platform in a unified manner. The cloud platform then processes the data and forwards it to the same or different servers. The data returned by the server is also first returned to the cloud platform, processed by the cloud platform, and then returned to the server.

[0135] In some feasible implementations, the screening process for the target specification query intent can be implemented by the cloud platform. That is, after the semantic understanding model on the server side recognizes the specification query intent from the user's query voice, it obtains an initial recognition result and returns the initial recognition result to the cloud platform. The initial recognition result includes multiple alternative specification query intents and the confidence prediction value corresponding to each alternative specification query intent; the cloud platform obtains current scene information from the display device, and then determines the target specification query intent from the alternative specification query intents based on the current scene information and the confidence prediction value, and then returns the target specification query intent as the intent recognition result to the display device.

[0136] In this embodiment, the cloud platform can smooth out the differences in communication methods at each end, and the terminals can interact with the cloud platform in a unified manner, which can ensure the uniqueness of data entry and exit, and can also process common data such as signature verification and device information verification, thereby reducing the coupling between other algorithm services and terminals.

[0137] In some embodiments, the semantic understanding model includes a speech recognition sub-model and a semantic understanding sub-model; in controlling the communication device to send the user inquiry speech to the server, so that the semantic understanding model deployed on the server can perform instruction manual query intent recognition on the user inquiry speech and obtain the intent recognition result, the controller is further configured to:

[0138] In the process of receiving the user's inquiry voice by the sound collector, if it is detected that the duration of the user's inquiry voice exceeds the preset duration threshold, voice segments with a duration equal to the preset duration threshold are sequentially intercepted from the user's inquiry voice in chronological order;

[0139] When a voice segment is intercepted, controlling the communication device to send the intercepted voice segment to the server, so that the speech recognition sub-model deployed on the server converts the voice segment into voice segment text;

[0140] Controlling the communication device to receive the voice segment text returned by the server in response to the user's voice inquiry;

[0141] controlling at least one of a display and an audio output device to display a text of the speech segment;

[0142] When the user's inquiry voice is received, the communication device is controlled to send the unsent voice segment to the server, so that the voice recognition sub-model deployed on the server converts the voice segment into voice segment text, and the semantic understanding sub-model deployed on the server performs instruction manual query intent recognition on the voice text to obtain the intent recognition result, wherein the voice text is spliced ​​by the voice segment text.

[0143] The speech recognition sub-model can refer to a model that converts speech into text. The semantic understanding sub-model can refer to a model that parses the speaker's intent from speech text. The semantic understanding model can also include a intonation recognition sub-model and a tone recognition sub-model. The recognition results of the intonation recognition sub-model and the tone recognition sub-model can also be input into the semantic understanding sub-model along with the speech text to parse the instruction manual query intent from the speech text.

[0144] The user inquiry voice may refer to all voice information collected by the sound collector from the time the user voice input operation is triggered to the time the voice input operation ends. In some embodiments, a voice input key is provided on the user interface or input device. The user can press the voice input key to trigger the voice input operation, and then the user starts speaking. After the user finishes speaking, the user releases the voice input key to trigger the end of the voice input operation. Alternatively, the user automatically triggers the voice input operation after waking up the voice assistant. The user can start speaking, and when the speaking interruption exceeds a preset time threshold, the end of the voice input operation is automatically triggered. In the case where the user's inquiry voice is long and contains a large amount of information, the amount of data in the user inquiry voice is large. The one-time transmission and processing can easily cause a large network burden and server computing burden, resulting in reduced overall resource utilization and processing efficiency.

[0145] For example, in the process of receiving the user's inquiry voice by the sound collector, the timing can be started from the time when the voice input operation is triggered. When the timing reaches the preset time threshold and the voice input end operation is not triggered, it is determined that the duration of the user's inquiry voice exceeds the preset time threshold. The earliest voice segment with a duration equal to the preset time threshold can be intercepted from the user's inquiry voice in chronological order and the timing is reset. Then, the communication device is controlled to send the intercepted voice segment together with the voice recognition instruction to the server. After resetting the timing, each time the timing reaches the preset time threshold and the voice input end operation is not triggered, the voice segment can be intercepted and sent to the server for voice recognition until the user triggers the voice input end operation. When the user triggers the voice input end operation, the display device can control the communication device to send the remaining voice segments that have not been sent to the server together with the voice recognition instruction and the instruction manual query intention recognition instruction. The voice recognition instruction is used to instruct the server to perform voice recognition on the voice segment through the voice recognition submodel to obtain the voice segment text, and return the obtained voice segment text to the display device. The instruction for manual query intent recognition is used to instruct the server to concatenate all the voice fragment texts into voice text in sequence, input the voice text into the semantic understanding model, perform manual query intent recognition on the voice text through the semantic understanding model, obtain the intent recognition result, and return the obtained intent recognition result to the display device.

[0146] During the process of sending the voice segment to the server, the display device may also control the communication device to receive the voice segment text returned by the server based on the voice segment, and control at least one of the display and the audio output device to display the voice segment text.

[0147] In this embodiment, by uploading the user's inquiry voice in segments, recognizing it in segments, and displaying the corresponding voice fragment text in segments, it can not only reduce the network burden and the computing burden of the server, but also provide more timely feedback on the voice recognition results and enhance the interactive experience; then, after forming a complete voice text, semantic understanding is performed to ensure the accuracy of semantic understanding.

[0148] In some embodiments, the query answer is composed of a plurality of streamed text data, each of which is sent by the server to a preset data queue after being generated. In the process of controlling the communication device to receive the query answer returned by the server for the target specification query intent, the controller is further configured to:

[0149] Controlling the communication device to sequentially obtain the streaming text data returned by the server in response to the target specification query intent from a preset data queue until the data queue is empty;

[0150] In controlling at least one of the display and the audio output device to display the answer to the query, the controller is further configured to:

[0151] Control at least one of a display and an audio output device to sequentially display the acquired streaming text data.

[0152] It should be noted that when the query and answer contain a lot of content, the amount of data in the query and answer is large, and one-time transmission is likely to cause a large network burden, which may cause feedback delay.

[0153] Streaming text data may refer to text information generated continuously and in real time. When the answer to a question is generated continuously and in real time, each word in the answer to the question may be generated sequentially to form streaming text data. The streaming text data may be returned one by one, first stored in a preset data queue, and then retrieved one by one by the display device from the preset data queue.

[0154] Exemplarily, after the server sends at least one streaming text data to the preset data queue, the preset data queue can send a data receiving instruction to the display device. After receiving the data receiving instruction, the controller of the display device can control the communication device to sequentially obtain the streaming text data returned by the server for the target specification query intention from the preset data queue until the data queue is empty.

[0155] In some possible implementations, such as Figure 8As shown, the server and the display device communicate through the cloud platform. After receiving the streaming text data sent by the server, the cloud platform parses the received streaming text data, stores the parsed streaming text data in the data queue, and sends a sync (Synchronize DiskCache) instruction to the display device. After receiving the sync instruction, the display device can control the communication device to obtain the streaming text data written by the cloud platform from the data queue in sequence until the data queue is empty, and wait for the termination instruction or the next sync instruction. If the termination instruction is received, the process ends. If the next sync instruction is received, the streaming text data continues to be extracted from the data queue.

[0156] In the process of receiving the streaming text data, the display device may control at least one of the display and the audio output device to display the acquired streaming text data each time a preset amount of streaming text data is received.

[0157] In this embodiment, during the query answer generation process, the streaming text data transmission and display method can be used to respond to the user's instruction manual query request earlier and more promptly, enhance the interactive experience, reduce the data transmission burden, and improve data transmission efficiency.

[0158] In some embodiments, after obtaining the query answer corresponding to the user's query voice, the controller is further configured to:

[0159] Detecting at least one of a derived operation query intent corresponding to a target specification query intent and a derived operation text content corresponding to a target text content;

[0160] Obtaining derivative operation prompt information corresponding to at least one of a derivative operation query intent and a derivative operation text content; wherein the derivative operation prompt information is generated based on the derived text content by locating the derived operation text content from a preset instruction manual text based on the derived operation query intent through an instruction manual text processing model, and / or the derived operation prompt information is generated based on the derived text content through the instruction manual text processing model;

[0161] Control at least one of a display and an audio output device to display derivative operation prompt information.

[0162] Among them, the derived operation query intent can refer to the query intent of the problem handling operation related to the target specification query intent. For example, when the target specification query intent is to inquire about the cause of the abnormal situation, the derived operation query intent can be to inquire about the solution to the abnormal situation. If only the cause of the abnormal situation is fed back to the user, but the user still does not know the solution, there is a high probability that they will continue to ask about the solution. If the solution is fed back together with the cause of the abnormal situation, the interactive operation can be simplified. The derived operation query intent related to each target specification query intent can be determined in advance based on actual conditions or test results, etc., and this embodiment does not limit this.

[0163] The derived operation text content may refer to the instruction text content for the problem handling operation related to the target text content. For example, if the target text content is the cause of an abnormal situation, the derived operation text content may be the solution to the abnormal situation. The derived operation text content related to each target text content can be determined in advance based on actual conditions or test results, and this embodiment does not limit this.

[0164] Derived operation prompt information can be used to guide the user to the next step to resolve the abnormal situation raised in the user's inquiry voice. Derived operation prompt information can be generated based on the derived operation text content located from the preset instruction manual text based on the derived operation query intent through the instruction manual text processing model, or it can be directly generated based on the derived text content through the instruction manual text processing model.

[0165] Exemplarily, after generating the query answer, the correspondence between the preset instruction manual query intention and the derived operation query intention can be queried to determine the derived operation query intention corresponding to the target instruction manual query intention; the correspondence between the preset text content and the derived operation text content can also be queried to determine the derived operation text content corresponding to the target text content; then at least one of the obtained derived operation query intention and derived operation text content is input into the instruction manual text processing model to generate derived operation prompt information; after obtaining the derived operation prompt information, at least one of the display and the audio output device can be controlled to display the derived operation prompt information.

[0166] In this embodiment, after generating the answer to the inquiry, while ensuring timely feedback of the answer to the inquiry, idle resources are further used to supplement the derived operation prompt information, thereby reducing the user's further query for solutions and simplifying the operation steps for manual query.

[0167] In some embodiments, as Figure 9As shown, when the user is using the display device, he can press the voice button according to his own needs to interact with the display device by voice. The display device responds to the voice button trigger operation, establishes a communication connection with the cloud platform and sends the current scene information to the cloud platform. After confirming that the communication connection with the cloud platform is established, the cloud platform shows the user a sound reception animation. When the cloud platform establishes a communication connection with the display device, it establishes a communication connection with the ASR service. Then, after seeing the sound reception animation, the user issues a user inquiry voice. The display device uploads the collected user inquiry voice to the cloud platform in a segmented upload manner, and the cloud platform transmits the voice segment sound file to the ASR service The ASR service performs speech recognition on the received voice clips, converts the voice clips into voice clip text, and sends it back to the cloud platform. The cloud platform returns the voice clip text sent back by the ASR service to the display device. The display device displays the received voice clip text to the user in real time. During the speech recognition process, the user's query voice is continuously uploaded and speech recognized in segments until the user releases the voice key. The display device responds to the voice key release operation, instructing the cloud platform to end speech recognition. The cloud platform disconnects the ASR service and establishes a communication connection with the natural language processing service. Then, the cloud platform sends a semantic understanding to the natural language processing service. The semantic understanding request carries the speech text composed of all the speech fragments. The natural language processing service parses the instruction manual query intent from the received speech text. The cloud platform further filters the instruction manual query intent based on the current scene information of the display device, and adjusts the confidence by weight. It re-sorts the query intent according to the adjusted confidence, and sends the highest ranked query intent as the best intent to the display device. The display device can show the best intent to the user for confirmation of whether it is correct. At the same time, it can send a query answer generation request to the cloud platform. The cloud platform establishes a communication connection with the instruction manual text processing model based on the query answer generation request, and sends the instruction manual text to the cloud platform. The processing model sends the target manual query intent and voice text. The manual text processing model generates a query answer based on the target manual query intent and voice text, and returns the query answer to the cloud platform in the form of streaming text data. After the cloud platform parses the streaming text data returned by the manual text processing model, it writes it into the preset data queue. The display device listens to the streaming text data written in the preset data queue and displays it to the user in real time. After generating a complete query answer, the manual text processing model disconnects the communication connection with the cloud platform, and the cloud platform synchronously feeds back to the display device. The display device can display the generated query answer information to the user.

[0168] In some embodiments, a method for querying a specification is provided, such as Figure 10 As shown, the method includes:

[0169] Step 1002: receiving a user's voice inquiry;

[0170] Step 1004: Obtain the intent recognition result corresponding to the user's inquiry voice. The intent recognition result is obtained by performing instruction manual query intent recognition on the user's inquiry voice using a semantic understanding model.

[0171] Step 1006: If the target manual query intent is determined based on the intent recognition result, a query answer corresponding to the user's voice query is obtained. The query answer is generated based on the target text content and the user's voice query by locating the target text content from the preset manual text based on the target manual query intent using the manual text processing model.

[0172] Step 1008: Display the answer to the inquiry.

[0173] In some embodiments, the intent recognition result includes multiple alternative specification query intents and corresponding confidence prediction values ​​for each alternative specification query intent; before obtaining the query answer corresponding to the user's query voice, the method further includes:

[0174] Get the current scene information of the display device;

[0175] According to the current scenario information and each confidence prediction value, the target specification query intent is determined from each alternative specification query intent.

[0176] In some embodiments, determining a target specification query intent from each candidate specification query intent based on current scenario information and each confidence prediction value includes:

[0177] According to the correspondence between the preset scenario and the instruction manual query intent, determining the alternative instruction manual query intent that does not match the current scenario information;

[0178] Filter out the alternative instruction manual query intents that do not match the current scenario information to obtain the primary instruction manual query intent;

[0179] According to the current scenario information and each confidence prediction value, the target instruction manual query intent is determined from each preliminary instruction manual query intent.

[0180] In some embodiments, determining a target specification query intent from each of the preliminary specification query intents based on the current scenario information and each confidence prediction value includes:

[0181] Based on the current scenario information, query the intent weight corresponding to each preliminary instruction manual query intent;

[0182] According to the weight of each intention, the confidence prediction value corresponding to each preliminary specification query intention is adjusted to determine the confidence target value corresponding to each preliminary specification query intention;

[0183] The preliminary specification query intent corresponding to the highest confidence target value is determined as the target specification query intent.

[0184] In some embodiments, the semantic understanding model and the instruction manual text processing model are deployed on the server side; the process of obtaining the intent recognition result corresponding to the user's query voice includes:

[0185] The user's query voice is sent to the server, so that the semantic understanding model deployed on the server can identify the instruction manual query intent of the user's query voice and obtain the intent recognition result;

[0186] Receive the intent recognition results returned by the server for the user's voice query;

[0187] In the process of obtaining the query answer corresponding to the user's query voice, the controller is further configured to:

[0188] Send the target manual query intent to the server, so that the manual text processing model deployed on the server can locate the target text content from the preset manual text based on the target manual query intent, and generate the query answer based on the target text content and the user's query voice;

[0189] Receive the query answer returned by the server for the target specification query intent.

[0190] In some embodiments, the semantic understanding model includes a speech recognition sub-model and a semantic understanding sub-model; the user query voice is sent to the server, so that the semantic understanding model deployed on the server performs instruction manual query intent recognition on the user query voice, and the intent recognition results obtained include:

[0191] In the process of receiving the user's inquiry voice by the sound collector, if it is detected that the duration of the user's inquiry voice exceeds the preset duration threshold, voice segments with a duration equal to the preset duration threshold are sequentially intercepted from the user's inquiry voice in chronological order;

[0192] When a voice segment is captured, the captured voice segment is sent to the server, so that the speech recognition sub-model deployed on the server converts the voice segment into voice segment text;

[0193] Receive the voice fragment text returned by the server in response to the user's voice query;

[0194] Display the text of the speech clip;

[0195] When the user's inquiry voice is received, the unsent voice segment is sent to the server, so that the voice recognition sub-model deployed on the server can convert the voice segment into voice segment text, and the semantic understanding sub-model deployed on the server can perform instruction query intent recognition on the voice text to obtain the intent recognition result, wherein the voice text is spliced ​​by the voice segment text.

[0196] In some embodiments, the query answer is composed of multiple streamed text data. After each streamed text data is generated, it is sent by the server to a preset data queue. The query answer returned by the server for the target specification query intent includes:

[0197] Sequentially obtain the streaming text data returned by the server for the target specification query intent from the preset data queue until the data queue is empty;

[0198] Answers to display inquiries include:

[0199] The obtained streaming text data is displayed in sequence.

[0200] In some embodiments, after obtaining the query answer corresponding to the user's query voice, the method further includes:

[0201] Detecting at least one of a derived operation query intent corresponding to a target specification query intent and a derived operation text content corresponding to a target text content;

[0202] Obtaining derivative operation prompt information corresponding to at least one of a derivative operation query intent and a derivative operation text content; wherein the derivative operation prompt information is generated based on the derived text content by locating the derived operation text content from a preset instruction manual text based on the derived operation query intent through an instruction manual text processing model, and / or the derived operation prompt information is generated based on the derived text content through the instruction manual text processing model;

[0203] Displays derivative operation prompt information.

[0204] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the method of the above embodiment when executing the computer program.

[0205] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method of the above embodiment is implemented.

[0206] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method of the above embodiment is implemented.

[0207] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0208] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0209] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0210] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A display device, characterized in that: include: a display configured to display a user interface; A sound collector is configured to receive external sound; an audio output device configured to output sound; The controller is configured as: Controlling the sound collector to receive user inquiry voice; Obtaining an intent recognition result corresponding to the user inquiry voice, wherein the intent recognition result is obtained by performing instruction manual query intent recognition on the user inquiry voice through a semantic understanding model; In the case where it is determined based on the intention recognition result that the target manual query intention is recognized, obtaining a query answer corresponding to the user query voice, wherein the query answer is generated based on the target text content and the user query voice by locating the target text content from the preset manual text based on the target manual query intention through the manual text processing model; Controlling at least one of the display and the audio output device to display the answer to the inquiry.

2. The display device according to claim 1, characterized in that The intention recognition result includes a plurality of alternative specification query intentions and confidence prediction values ​​corresponding to each of the alternative specification query intentions; Before obtaining the inquiry answer corresponding to the user inquiry voice, the controller is further configured to: Acquire current scene information of the display device; According to the current scene information and each confidence prediction value, a target specification query intent is determined from each of the alternative specification query intents.

3. The display device according to claim 2, characterized in that In the process of determining the target specification query intent from each of the candidate specification query intents according to the current scene information and each of the confidence prediction values, the controller is further configured to: According to the correspondence between the preset scenario and the specification query intention, determining the alternative specification query intention that does not match the current scenario information; Filter out the alternative specification query intents that do not match the current scene information to obtain the primary specification query intent; According to the current scene information and each confidence prediction value, a target specification query intent is determined from each of the preliminary specification query intents.

4. The display device according to claim 3, characterized in that In the process of determining the target specification query intent from each of the preliminary specification query intents according to the current scene information and each of the confidence prediction values, the controller is further configured to: Based on the current scenario information, query the intent weights corresponding to the respective initial specification query intents; According to each of the intention weights, the confidence prediction values ​​corresponding to each of the preliminary specification query intentions are adjusted to determine the confidence target values ​​corresponding to each of the preliminary specification query intentions; The preliminary specification query intent corresponding to the highest confidence target value is determined as the target specification query intent.

5. The display device according to claim 1, characterized in that The semantic understanding model and the instruction text processing model are deployed on the server; the display device further includes a communication device, which is configured to communicate with the server; In the process of obtaining the intention recognition result corresponding to the user inquiry voice, the controller is further configured to: Controlling the communication device to send the user inquiry voice to the server, so that the semantic understanding model deployed on the server performs instruction manual query intent recognition on the user inquiry voice to obtain an intent recognition result; Controlling the communication device to receive the intention recognition result returned by the server in response to the user's inquiry voice; In the process of obtaining the inquiry answer corresponding to the user inquiry voice, the controller is further configured to: Controlling the communication device to send the target manual query intent to the server, so that the manual text processing model deployed on the server locates the target text content from the preset manual text based on the target manual query intent, and generates a query answer based on the target text content and the user query voice; The communication device is controlled to receive the query answer returned by the server in response to the target specification query intention.

6. The display device according to claim 5, characterized in that The communication device is configured to establish a communication connection with the cloud platform and communicate with the server through the cloud platform.

7. The display device according to claim 5, characterized in that The semantic understanding model includes a speech recognition sub-model and a semantic understanding sub-model; in the process of controlling the communication device to send the user inquiry speech to the server so that the semantic understanding model deployed on the server performs instruction manual query intent recognition on the user inquiry speech and obtains the intent recognition result, the controller is further configured to: In the process of receiving the user inquiry voice by the sound collector, when it is detected that the duration of the user inquiry voice exceeds a preset duration threshold, voice segments with a duration equal to the preset duration threshold are sequentially intercepted from the user inquiry voice in chronological order; When a voice segment is intercepted, controlling the communication device to send the intercepted voice segment to the server, so that the speech recognition sub-model deployed on the server converts the voice segment into a voice segment text; Controlling the communication device to receive the voice segment text returned by the server in response to the user's inquiry voice; controlling at least one of the display and the audio output device to display the speech segment text; When the user inquiry voice is received, the communication device is controlled to send the unsent voice segment to the server, so that the voice recognition sub-model deployed on the server converts the voice segment into voice segment text, and the semantic understanding sub-model deployed on the server performs instruction manual query intent recognition on the voice text to obtain an intent recognition result, wherein the voice text is spliced ​​by the voice segment text.

8. The display device according to claim 5, characterized in that The query answer is composed of a plurality of streaming text data, and each streaming text data is sent by the server to a preset data queue after being generated; In the process of controlling the communication device to receive the query answer returned by the server in response to the target specification query intention, the controller is further configured to: Controlling the communication device to sequentially obtain the streaming text data returned by the server for the target specification query intent from the preset data queue until the data queue is empty; In the process of controlling at least one of the display and the audio output device to display the answer to the inquiry, the controller is further configured to: Control at least one of the display and the audio output device to display the acquired streaming text data in sequence.

9. The display device according to any one of claims 1 to 8, characterized in that: After obtaining the inquiry answer corresponding to the user inquiry voice, the controller is further configured to: Detecting at least one of a derivative operation query intent corresponding to the target specification query intent and a derivative operation text content corresponding to the target text content; Obtaining derivative operation prompt information corresponding to at least one of the derivative operation query intent and the derivative operation text content; wherein the derivative operation prompt information is located from a preset instruction manual text based on the derivative operation query intent through an instruction manual text processing model, and is generated based on the derived text content, and / or the derivative operation prompt information is generated based on the derived text content through an instruction manual text processing model; Control at least one of the display and the audio output device to display the derived operation prompt information.

10. A method for querying instructions, characterized in that: The method comprises: Receive user inquiry voice; Obtaining an intent recognition result corresponding to the user inquiry voice, wherein the intent recognition result is obtained by performing instruction manual query intent recognition on the user inquiry voice through a semantic understanding model; In the case where it is determined based on the intention recognition result that the target manual query intention is recognized, obtaining a query answer corresponding to the user query voice, wherein the query answer is generated based on the target text content and the user query voice by locating the target text content from the preset manual text based on the target manual query intention through the manual text processing model; The answer to the query is displayed.

Citation Information

Cited By

  • Question and answer processing method and device

    CN120687181A