Display device, voice search method and storage medium

The user's voice is obtained through the user input interface of the display device, the media name and type are identified, the target knowledge graph library is used to correct the deviation of user intention, and the media type and verb are accurately identified, which solves the search difficulties caused by confusion in media names and improves the user experience.

CN115862615BActive Publication Date: 2025-10-10HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211428652.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-10-10
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

When users search for media resources through voice, it is often difficult to accurately identify the intention due to confusion in the media resource names. Existing technologies are unable to accurately identify user intentions, which affects the user experience.

Method used

The user voice is obtained through the user input interface of the display device, the controller identifies the media name, type and verb, uses the target knowledge graph library to determine the candidate media, makes a prediction based on the media name and type, determines and predicts, identifies the target media, provides a method ...

Benefits of technology

It achieves accurate identification of user intent and provides matching media resource search results when the media resource name does not match the type and/or verb, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862615B_ABST
    Figure CN115862615B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a display device, a voice search method and a storage medium, and relates to the technical field of voice interaction. The display device comprises: a user input interface configured to obtain a user voice; a controller configured to: identify the user voice, obtain a keyword in the user voice, the keyword comprising a media asset name, a media asset type and / or a verb; determine candidate media assets from a target knowledge graph database according to the media asset name, the candidate media assets comprising a first media asset indicated by the media asset name and to-be-fed-back media assets having an association relationship with the first media asset; determine target media assets matching the media asset type and / or the verb from the candidate media assets, and control a display to display search results corresponding to the target media assets. The embodiments of the present disclosure are used to solve the problem that the existing voice search method is difficult to accurately identify user intent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Bluetooth technology, and in particular to a display device, a voice search method, and a storage medium. Background Art

[0002] Currently, users are increasingly using voice to search for popular TV series, songs, and other content. As electronic products become increasingly intelligent, users' demands for AI comprehension capabilities are also increasing. In real life, users may struggle to accurately describe the media they are searching for. For example, a user might confuse TV series title A with the title of the TV series' theme song, title B. They might voice-instruct, "I want to watch TV series theme song B," when their actual intention is to watch the TV series. Based on the title of the TV series theme song, title B, included in the voice instruction, the TV will then display the user's music video (MV) for the theme song. This voice search method struggles to accurately identify the user's true intent, deviating from their actual needs and impacting their user experience. Summary of the Invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a display device, a voice search method and a storage medium, which can accurately identify user intentions and enhance user experience.

[0004] In order to achieve the above objectives, the technical solutions provided by the embodiments of the present disclosure are as follows:

[0005] In a first aspect, the present disclosure provides a display device, comprising:

[0006] The user input interface is configured to: obtain user voice;

[0007] The controller is configured to: recognize user speech and obtain keywords in the user speech, where the keywords include media asset names, media asset types, and / or verbs;

[0008] Determine candidate media assets from the target knowledge graph library according to the media asset name, where the candidate media assets include a first media asset indicated by the media asset name and media assets to be fed back that are associated with the first media asset;

[0009] From the candidate media assets, a target media asset that matches the media asset type and / or verb is determined, and the display is controlled to display the search results corresponding to the target media asset.

[0010] In a second aspect, the present disclosure provides a voice search method, comprising:

[0011] Get user voice;

[0012] Recognize user speech and obtain keywords in the user speech, including media asset name, media asset type and / or verbs;

[0013] Determine candidate media assets from the target knowledge graph library according to the media asset name, where the candidate media assets include a first media asset indicated by the media asset name and media assets to be fed back that are associated with the first media asset;

[0014] From the candidate media assets, a target media asset that matches the media asset type and / or verb is determined, and the display is controlled to display the search results corresponding to the target media asset.

[0015] In a third aspect, the present disclosure provides a computer-readable storage medium, comprising: a computer program stored on the computer-readable storage medium, and when the computer program is executed by a processor, the voice search method shown in the second aspect is implemented.

[0016] In a fourth aspect, the present disclosure provides a computer program product, which includes a computer program. When the computer program is run on a computer, the computer implements the voice search method as shown in the second aspect.

[0017] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art:

[0018] The disclosed embodiment provides a display device, a voice search method, and a storage medium, wherein the controller of the display device recognizes the user voice obtained by the user input interface, obtains the keywords therein, including the media asset name, and the media asset type and / or verb, and then first determines the candidate media from the target knowledge graph library based on the media asset name, the first media asset indicated by the media asset name in the candidate media, and the media asset to be fed back that has an associated relationship with the first media asset, and then determines the target media asset that matches the media asset type and / or verb from the candidate media, and controls the display to display the search results corresponding to the target media asset. In this way, even if the media asset name indicated by the user voice does not match the media asset type and / or verb, the matching media asset search results can still be accurately fed back to the user, thereby accurately analyzing and understanding the user voice, identifying the user's true intention, and improving the user's usage experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0020] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1Schematic diagrams of scenarios in some embodiments provided in the present disclosure;

[0022] Figure 2 exemplarily shows a configuration block diagram of the control device 100 according to an exemplary embodiment;

[0023] Figure 3 FIG. 2 shows a hardware configuration block diagram of a display device 200 according to an exemplary embodiment;

[0024] Figure 4 Schematic diagram of software configuration in the display device 200 according to one or more embodiments of the present disclosure;

[0025] Figure 5 A schematic diagram of a system architecture of a display device provided in an embodiment of the present disclosure;

[0026] Figure 6 A schematic diagram of a voice interaction network architecture provided by an embodiment of the present disclosure;

[0027] Figure 7 A flowchart of a voice search method is provided in an embodiment of the present disclosure;

[0028] Figure 8 Schematic diagram of the user interface of the voice search provided by the embodiment of the present disclosure Figure 1 ;

[0029] Figure 9 A schematic diagram of the audio-visual knowledge graph library provided in an embodiment of the present disclosure;

[0030] Figure 10 Schematic diagram of the user interface of the voice search provided by the embodiment of the present disclosure Figure 2 ;

[0031] Figure 11 A schematic structural diagram of a display device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0033] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0034] The terms "first", "second", "third", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to the specific circumstances. In addition, in the description of the present disclosure, unless otherwise specified, "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0035] At present, more and more users are using voice to search for TV dramas, songs and other media content on TV. With the widespread application of smart TVs, users have higher and higher requirements for the understanding ability of artificial intelligence. In real life, users often confuse the names of TV dramas and TV drama theme songs. They want to search for TV dramas but use the name of TV drama theme songs for voice search, or want to search for TV drama theme songs but use the name of TV drama for voice search.

[0036] For example, a user inputs the voice command "I want to watch the TV series "Borrow Another Five Hundred Years from Heaven"" to the TV, but "Borrow Another Five Hundred Years from Heaven" is not the name of the TV series, but the theme song of the TV series "Kangxi Dynasty", which means that the user's real intention is "I want to watch the TV series "Kangxi Dynasty""; however, it is difficult for the TV to accurately understand the user's intention. After recognizing the command "I want to watch the TV series "Borrow Another Five Hundred Years from Heaven", it will feedback the MV of the song "Borrow Another Five Hundred Years from Heaven" to the user, which deviates from the user's actual need to watch the TV series "Kangxi Dynasty" and affects the user's user experience.

[0037] There are many types of media resources. The names of the media resources are the same or similar, but the corresponding media content and media types are not exactly the same. This also increases the difficulty for users to search for media resources by voice, making it difficult for users to accurately search for the desired media resources through voice.

[0038] To solve the above technical problems, the display device, the voice search method and the storage medium are provided, wherein the display device comprises a user input interface and a controller, the user input interface is used to obtain a user voice, and the controller is used to: firstly, recognize the obtained user voice to obtain a media asset name, a media asset type and / or a verb and other keywords included in the user voice; secondly, determine candidate media assets from a target knowledge graph library according to the media asset name, wherein the candidate media assets comprise a first media asset indicated by the media asset name and a to-be-feedback media asset having an association relationship with the first media asset; thirdly, determine a target media asset matching the media asset type and / or the verb from the candidate media assets; and fourthly, control a display to display a search result corresponding to the target media asset. Thus, in the case that the user voice does not match the media asset name, the media asset type and / or the verb, the matching media asset search result can be accurately fed back to the user, the user voice can be accurately analyzed and understood, the real intention of the user can be recognized, and the user experience is improved.

[0039] Figure 1 The scene schematic diagram in some embodiments provided by the embodiments of the present disclosure is shown in FIG. 1. Figure 1 As shown in FIG. 1, Figure 1 The display device 200, the smart device 300 and the server 400 are included in the control device 100. The user can operate the display device 200 through the smart device 300 or the control device 100 to play an audio / video resource on the display device 200.

[0040] Taking the user operating the display device 200 through the control device 100 as an example, the user opens a user input interface, such as a microphone, of the display device 200 through the control device 100, so that the display device 200 obtains a user voice. The user expects to control the display device 200 to play a media asset through voice, the user input interface of the display device 200 receives the user voice, and then the controller of the display device 200 recognizes the user voice to obtain keywords included in the user voice: a media asset name, a media asset type and / or a verb. Then, the controller determines candidate media assets from a target knowledge graph library according to the media asset name, wherein the candidate media assets comprise a first media asset indicated by the media asset name and a to-be-feedback media asset having an association relationship with the first media asset. Then, the controller determines a target media asset matching the media asset type and / or the verb from the candidate media assets, and further controls a display to display a search result corresponding to the target media asset.

[0041] Compared with the prior art that searches only based on media asset names, the present disclosure first determines candidate media assets from the target knowledge graph library based on the media asset names and media asset types and / or verbs included in the user's voice, so as to correct the deviation indicated by the user's voice and narrow the range of media assets. Then, the target business parameters are calculated based on the media asset types and / or verbs, thereby determining the business type of the target media assets that the user actually expects, accurately identifying the user's true intentions, and obtaining search results for target media assets that match the user's expected media asset names and media asset types and / or verbs. Voice search is more accurate, more in line with the user's actual needs, and improves the user's usage experience.

[0042] In some embodiments, the control device 100 may be a remote controller, and communication between the remote controller and the display device may include infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, etc. In some embodiments, a mobile terminal, tablet computer, computer, laptop computer, or other smart device may also be used to control the display device 200.

[0043] In some embodiments, the smart device 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the smart device 300 can also be transmitted to the display device 200 to achieve a synchronous display function. The display device 200 also communicates data with the server 400 through a variety of communication methods. The display device 200 can be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN) and other networks. The server 400 can provide various content and interactions to the display device 200. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing a broadcast receiving television function, the display device 200 can also provide an intelligent network TV function that provides computer support functions.

[0044] Figure 2 Schematically shows a block diagram of the configuration of the control device 100 according to an exemplary embodiment. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, acting as an intermediary for interaction between the user and the display device 200. The communication interface 130 is used for external communication and includes at least one of a Wi-Fi chip, a Bluetooth module, NFC, or an alternative module. The user input / output interface 140 includes at least one of a microphone, a touchpad, a sensor, a button, or an alternative module.

[0045] Figure 3 FIG. 2 shows a hardware configuration block diagram of a display device 200 according to an exemplary embodiment. Figure 3 The display device 200 shown includes a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a user input interface 280, a memory, a power supply, and the like. The controller 250 includes a central processing unit (CPU), a video processor, an audio processor, a graphics processor, RAM, ROM, and first through nth interfaces for input / output. The display 260 can be at least one of a liquid crystal display (LCD), an OLED display, a touchscreen display, and a projection display, and can also be a projection device and projection screen. The tuner / demodulator 210 receives broadcast television signals via wired or wireless reception, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or external interactions. The controller 250 and tuner / demodulator 210 can be located in separate devices, i.e., the tuner / demodulator 210 can be located in a device external to the main device where the controller 250 resides, such as an external set-top box.

[0046] In some embodiments, the above-mentioned display device is a terminal device with a display function, such as a television, a mobile phone, a computer, a learning machine, etc.

[0047] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. The user can enter user commands through the graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI). Alternatively, the user can enter user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0048] an output interface (display 260 and / or audio output interface 270 ), configured to output user interaction information;

[0049] Communicator 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip, a near-field communication protocol chip, and an infrared receiver. Display device 200 can establish communication with server 400 via communicator 220 to send and receive control signals and data signals.

[0050] The user input interface 280 can be used to receive external control signals.

[0051] Detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or detector 230 includes an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or detector 230 includes a sound collector, such as a microphone, for receiving external sounds.

[0052] The sound collector can be a microphone, also known as a "microphone" or "microphone", which can be used to receive the user's voice and convert the sound signal into an electrical signal. The display device 200 can be provided with at least one microphone. In other embodiments, the display device 200 can be provided with two microphones, which can not only collect sound signals but also implement a noise reduction function. In other embodiments, the display device 200 can also be provided with three, four, or more microphones to collect sound signals, reduce noise, identify the sound source, implement a directional recording function, etc.

[0053] In addition, the microphone may be built into the display device 200, or the microphone may be connected to the display device 200 by wire or wireless means. Of course, the embodiment of the present application does not limit the position of the microphone on the display device 200. Alternatively, the display device 200 may not include a microphone, that is, the microphone is not provided in the display device 200. The display device 200 may be connected to an external microphone (also referred to as a microphone) through an interface (such as a USB interface 130). The external microphone may be fixed to the display device 200 by an external fixing member (such as a camera bracket with a clip).

[0054] An embodiment of the present disclosure provides a display device 200, which includes:

[0055] The user input interface 280 is configured to: obtain user voice;

[0056] The controller 250 is configured to: recognize user speech and obtain keywords in the user speech, where the keywords include media asset names, media asset types, and / or verbs;

[0057] Determine candidate media assets from the target knowledge graph library according to the media asset name, where the candidate media assets include a first media asset indicated by the media asset name and media assets to be fed back that are associated with the first media asset;

[0058] From the candidate media assets, a target media asset that matches the media asset type and / or verb is determined, and the display 260 is controlled to display the search result corresponding to the target media asset.

[0059] The display device 200 corrects the possible deviation in the user voice by identifying the media name and the media type and / or the verb in the user voice, accurately identifies the user intention, avoids the situation that the user cannot obtain accurate media search results due to the confusion of the media name, improves the user friendliness, and improves the user experience.

[0060] In some embodiments, the target knowledge graph library is at least one of preset knowledge graph libraries; the preset knowledge graph libraries include: a first error correction knowledge graph library, a second error correction knowledge graph library, and a video and audio knowledge graph library; the first error correction knowledge graph library includes media whose media names have a pronunciation similarity greater than a first similarity threshold but different media types; the second error correction knowledge graph library includes media whose media names have a pronunciation similarity greater than a second similarity threshold but the same media content, and the second similarity threshold is greater than the first similarity threshold; the video and audio knowledge graph library includes video media and music media, and the video media and the music media have a corresponding relationship.

[0061] In some embodiments, the number of target media is multiple;

[0062] The controller 250 controls the display 260 to display the search results corresponding to the target media, and is configured to: acquire historical search records, and determine a first sorting weight of the multiple target media according to the historical search records; acquire resource heat parameters of the multiple target media, and determine a second sorting weight of the multiple target media according to the resource heat parameters of the multiple target media; calculate a target sorting weight according to the first sorting weight and the second sorting weight; and control the display 260 to display the search results corresponding to the target media according to the target sorting weight.

[0063] In some embodiments, the controller 250 determines the target media matching the media type and / or the verb from the candidate media, and is configured to: determine a target business parameter by calculating a business parameter corresponding to the media type and / or the verb and a business parameter corresponding to the actual media type of the first media, the target business parameter being the business parameter corresponding to the user voice; and determine the target media from the candidate media according to the target business parameter.

[0064] In some embodiments, the controller 250 determines the target media from the candidate media according to the target business parameter, and is configured to: determine a second media from the to-be-feedback media according to the target business parameter; if the media type of the second media is the same as the media type of the first media, control the display 260 to display the search results corresponding to the first media and the search results corresponding to the second media; and if the media type of the second media is different from the media type of the first media, control the display 260 to display the search results corresponding to the second media.

[0065] In some embodiments, the target knowledge graph library is a second error correction knowledge graph library; the second error correction knowledge graph includes media assets whose pronunciation similarity of the media asset names is greater than a second similarity threshold but whose media asset contents are different;

[0066] Before determining the service parameters corresponding to the media asset type and / or verb and the service parameters corresponding to the actual media asset type of the first media asset to calculate the target service parameters, the controller 250 is further configured to: determine whether the media asset type of the first media asset is the same as that of the media asset to be fed back;

[0067] The controller 250 determines the business parameters corresponding to the media asset type and / or verb, as well as the business parameters corresponding to the actual media asset type of the first media asset, to calculate the target business parameters, and is configured to: when the first media asset is different from the media asset type of the media asset to be fed back, determine the business parameters corresponding to the media asset type and / or verb, as well as the business parameters corresponding to the actual media asset type of the first media asset, to calculate the target business parameters.

[0068] In some embodiments, the controller 250, after recognizing the user voice and obtaining the keywords in the user voice, is further configured to determine whether the media asset name and the media asset type correspond, and / or whether the media asset name and the verb correspond; if the media asset name and the media asset type do not correspond, and / or the media asset name and the verb do not correspond, then determine the target knowledge graph library based on the media asset name before determining the candidate media asset from the target knowledge graph library based on the media asset name.

[0069] like Figure 4 As shown, Figure 4 FIG. 1 is a schematic diagram of software configuration in the display device 200 according to one or more embodiments of the present disclosure. Figure 4 As shown in Figure 1, the system is divided into four layers: from top to bottom, the Applications layer (referred to as the "Application layer"), the Application Framework layer (referred to as the "Framework layer"), the Android runtime and system library layer (referred to as the "System Runtime Library layer"), and the Kernel layer. The Kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, Wi-Fi driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.

[0070] In some examples, the operating system of the smart device is an Android system, for example, Figure 5 As shown, Figure 5This is a schematic diagram of the system architecture of a display device provided in an embodiment of the present disclosure. The display device 200 can be logically divided into an application layer (abbreviated as "application layer") 21, a kernel layer 22 and a hardware layer 23.

[0071] Among them, such as Figure 5 As shown, the hardware layer may include Figure 3 The controller 250, communicator 220, detector 230, etc. are shown. The application layer 21 includes one or more applications. The applications can be system applications or third-party applications. For example, the application layer 21 includes a voice recognition application that can provide a voice interaction interface and services for connecting the display device 200 with the server 400.

[0072] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21 and is used to manage and control hardware and software resources.

[0073] In some examples, the kernel layer 22 includes a detector driver, which is used to send the voice data collected by the detector 230 to the voice recognition application. For example, when the voice recognition application in the display device 200 is started and a communication connection is established between the display device 200 and the server 400, the detector driver is used to send the user input voice data collected by the detector 230 to the voice recognition application. The voice recognition application then sends query information containing the voice data to the intent recognition module 202 in the server. The intent recognition module 202 is used to input the voice data sent by the display device 200 into the intent recognition model.

[0074] To clearly illustrate the embodiments of the present disclosure, Figure 6 A speech recognition network architecture provided by an embodiment of the present disclosure is described.

[0075] See also Figure 6 , Figure 6 A schematic diagram of a voice interaction network architecture provided in an embodiment of the present disclosure. Figure 6In the embodiment, the display device is used to receive input information and output the processing result of the information. The Automatic Speech Recognition (ASR) module is deployed with a speech recognition service for recognizing audio as text; the Natural Language Understanding (NLU) module is deployed with a semantic understanding service for semantically parsing text; the Dialog Manager (DM) module is deployed with a business instruction management service for providing business instructions; the language generation module is deployed with a language generation service (NLG) for converting instructions for the display device to execute into text language; the speech synthesis module is deployed with a text to speech synthesis (TTS) service for processing the text language corresponding to the instruction and sending it to the speaker for broadcasting. In one embodiment, Figure 6 The illustrated architecture may include multiple physical service devices deployed with different business services, or one or more physical service devices may integrate one or more functional services.

[0076] In some embodiments, the following Figure 6 The process of processing the information input to the display device in the illustrated architecture is described by way of example, taking the case where the information input to the display device is a voice command inputted via voice as an example:

[0077] [Voice Recognition] After receiving a voice command input via voice, the display device can perform noise reduction processing and feature extraction on the audio of the voice command. The noise reduction processing here may include steps such as removing echoes and ambient noise.

[0078] [Semantic Understanding] Utilizes acoustic models and language models to perform natural language understanding on the identified candidate text and associated contextual information, parsing the text into structured, machine-readable information, including business domain, intent, word slots, and other information to express semantics. The executable intent is obtained and the intent confidence score is determined. The semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.

[0079] [Dialogue Management] The semantic understanding module sends execution instructions to the corresponding business management module based on the semantic analysis results of the voice command text to perform the operation corresponding to the voice command, complete the user's request for this operation, and provide feedback on the execution results of the operation corresponding to the voice command.

[0080] In order to explain this solution in more detail, the following will be combined with the following examples. Figure 6 To explain, it is understandable that Figure 7The steps involved may include more steps or fewer steps in actual implementation, and the order of these steps may also be different, so as to be able to implement the voice search method provided in the embodiment of the present disclosure.

[0081] like Figure 7 As shown, Figure 7 A flow chart of a voice search method is provided in an embodiment of the present disclosure. The method includes the following steps S701 to S704:

[0082] S701: Acquire user voice.

[0083] In some embodiments, the display device obtains the user voice through the user input interface, and may also obtain the user voice input by the user through a voice device externally connected to the user input interface.

[0084] In some embodiments, after the user input interface of the display device acquires the user voice, the user voice is preprocessed. The preprocessing includes but is not limited to at least one of the following: denoising and voice extraction, which is not limited in the present disclosure.

[0085] In some embodiments, the user opens the user input interface of the display device through the control device or smart device, and the display device displays the user interface of the voice search to remind the user to start speaking, such as Figure 8 As shown, Figure 8 Schematic diagram of the user interface of the voice search provided by the embodiment of the present disclosure Figure 1 , a microphone icon is shown in the figure to remind the user that he can start talking to the display device. Figure 8 This is only an exemplary illustration, and the present disclosure does not specifically limit the user interface of the voice search.

[0086] S702: Recognize user voice and obtain keywords in the user voice.

[0087] Among them, the keywords include media asset name, media asset type and / or verb; it is understandable that the keywords may include media asset name and media asset type, or include media asset name and verb, or include media asset name, media asset type and verb.

[0088] For example, the media asset name may be "Borrow Another Five Hundred Years from Heaven", "Ashes of Love", etc., the media asset type may be "TV series", "song", "movie", "variety show", etc., and the verb may be "play", "watch", "listen", or "introduce".

[0089] In some embodiments, during the process of recognizing the user's voice, the user's voice is first converted into text characters, and the text characters are segmented to obtain keywords.

[0090] Taking Table 1 as an example, Table 1 shows the user's voice and the keywords included therein.

[0091] Table 1

[0092]

[0093] In some embodiments, the display device can be connected to a server via a communication module and transmit the captured user voice to the server for recognition by the server. The display device then receives the recognition result of the user voice from the server. Of course, the recognition of the user voice can be performed by the display device, such as in step S702 above and any of the embodiments therein, or only a portion of the voice information required for server processing can be transmitted to the server. This disclosure is not limited to this.

[0094] In the above embodiment, the media asset name, media asset type and / or verb included in the user's voice are obtained by performing voice recognition on the user's voice, so as to accurately understand the true intention contained in the user's voice based on the above keywords.

[0095] S703: Determine candidate media assets from the target knowledge graph library according to the media asset name.

[0096] The target knowledge graph library is at least one of the preset knowledge graph libraries, which include a first error correction knowledge graph library, a second error correction knowledge graph library, and an audio-visual knowledge graph library.

[0097] The first error correction knowledge graph includes media assets of different types whose pronunciation similarity exceeds a first similarity threshold. The first similarity threshold is a pre-set threshold used to distinguish whether the pronunciations of media asset names are similar, typically set to 60%. Pronunciation similarity refers to similar nasal sounds and similar tones. For example, the TV series "Deep Love" and the song "Deep Love" belong to the same first error correction knowledge graph.

[0098] The second error correction knowledge graph includes media assets whose pronunciation similarity of the media asset names exceeds a second preset similarity threshold, but whose media asset content differs. The second similarity threshold is a pre-set similarity threshold used to distinguish whether the pronunciations of media asset names are identical. The second similarity threshold is greater than the first similarity threshold and is typically set to 100%. For example, "Sound of Life" and "Life and Life" are a variety show and music respectively.

[0099] The audio-visual knowledge graph includes both film and television media assets and music media assets, and these assets correspond to each other. For example, the film and television asset "Kangxi Dynasty" has its corresponding music asset "Borrowing Another 500 Years from Heaven," and the film and television asset "Ashes of Love" has its corresponding music assets "Left Hand Pointing to the Moon" and "Unstained." It should be emphasized that film and television media assets include, but are not limited to, TV series, movies, variety shows, documentaries, and operas, and this disclosure does not specifically limit these assets.

[0100] Each knowledge graph in the above-mentioned preset knowledge graph library uses each media asset as a node and the relationship between the media assets as an edge.

[0101] For example, Figure 9 Show, Figure 9 This is a schematic diagram of the audio-visual knowledge graph library provided in an embodiment of the present disclosure. The nodes in the graph include film and television media assets: "Kangxi Dynasty" and "Ashes of Love", and music media assets: "Borrowing Five Hundred Years from Heaven", "Left Hand Pointing to the Moon", and "Unstained". Among them, "Kangxi Dynasty" and "Borrowing Five Hundred Years from Heaven" have a corresponding relationship, with "Kangxi Dynasty" and "Borrowing Five Hundred Years from Heaven" as nodes, and the corresponding relationship between the two is an edge. The length of the edge between the nodes here is related to the strength of the corresponding relationship. Please refer to the prior art for details, which will not be elaborated in this disclosure; "Ashes of Love" and "Left Hand Pointing to the Moon" and "Unstained" have a corresponding relationship, with "Ashes of Love", "Left Hand Pointing to the Moon", and "Unstained" as nodes, and the corresponding relationship between the three is an edge.

[0102] The candidate media assets include the first media asset indicated by the media asset name and the to-be-feedback media assets that are associated with the first media asset.

[0103] In some embodiments, after identifying the user's voice and obtaining the media asset name contained therein, the quantized flag value corresponding to the media asset name is first determined. If the flag value is 0, it indicates that the first media asset indicated by the media asset name is relatively simple in terms of media asset type and media asset content, and there are no other media assets that can be confused with the first media asset. In this case, the present embodiment provides an implementation method that does not call a preset knowledge graph library, but directly uses the vocabulary annotation and text logic reasoning in the main dictionary to obtain the first media asset indicated by the media asset name, and the display device controls the display to display the search results for the first media asset. The search results for the first media asset include, but are not limited to: the first media asset's label, detailed information about the first media asset, and the first media asset's media content.

[0104] For example, if the flag value corresponding to the media asset name in the user's voice is 0, the user is directly fed back the search results for the first media asset indicated by the media asset name. The display shows the poster and episode number of the first media asset, and the poster is an image link to the first media asset's details page. It is understood that by clicking on the first media asset's poster, the user jumps to the first media asset's details page to view the first media asset's detailed content.

[0105] When the flag value is 1, it means that there are other media assets with similar pronunciation to the media asset name, or there are other media assets corresponding to the media asset name. The embodiment of the present disclosure provides an implementation method. When the flag value corresponding to the media asset name is 1, it is determined that the target knowledge graph library is the first error correction knowledge graph library, or the target knowledge graph library is the audio-visual knowledge graph library, or the target knowledge graph library is the first error correction knowledge graph library and the audio-visual knowledge graph library.

[0106] Furthermore, the first media asset indicated by the media asset name in the target knowledge graph library and the media asset to be fed back that is associated with the first media asset are used as candidate media assets. Optionally, the first media asset indicated by the media asset name in the first error correction knowledge graph library and the media asset to be fed back that has a pronunciation similarity with the first media asset greater than a first similarity threshold but is of a different media type are used as candidate media assets. Alternatively, the candidate media assets include: the first media asset indicated by the media asset name in the audio-visual knowledge graph library and the media asset to be fed back, wherein, if the first media asset is a film and television media asset, the media asset to be fed back is the music media asset corresponding to the first media asset; if the first media asset is a music media asset, the media asset to be fed back is the film and television media asset corresponding to the first media asset. Alternatively, the candidate media assets include: the first media asset indicated by the media asset name in the first error correction knowledge graph library, and the media asset to be fed back having a pronunciation similarity with the first media asset greater than a first similarity threshold but a different media asset type, as well as the first media asset indicated by the media asset name in the audio-visual knowledge graph library and the media asset to be fed back that has a corresponding relationship with the first media asset, wherein, if the first media asset is a film and television media asset, the media asset to be fed back is the music media asset corresponding to the first media asset; if the first media asset is a music media asset, the media asset to be fed back is the film and television media asset corresponding to the first media asset.

[0107] When the flag value is 2, it indicates that there are other media assets with the same pronunciation as the media asset name. The embodiment of the present disclosure provides an implementation method. When the flag value is 2, the target knowledge graph library is determined to be the second error correction knowledge graph library, and the first media asset indicated by the media asset name in the second error correction knowledge graph, and the media to be fed back whose pronunciation similarity with the media asset name is greater than the second similarity threshold but whose media content is different, are taken as candidate media assets.

[0108] In some embodiments, after recognizing the user voice and obtaining the keywords included therein, if the keywords include the media asset name and the media asset type, it is determined whether the media asset name and the media asset type correspond; or if the keywords include the media asset name and the verb, it is determined whether the media asset name and the verb correspond; or, if the keywords include the media asset name, the media asset type and the verb, it is determined whether the media asset name corresponds to the media asset type and the verb.

[0109] If the media asset name and the media asset type do not correspond, it means that there is a deviation between the content described by the user's voice and the media asset that the user actually expects to search for. The embodiment of the present disclosure provides an implementation method. When the media asset name and the media asset type do not correspond, the target knowledge graph library is determined from the preset knowledge graph library based on the media asset name to obtain candidate media assets from the target knowledge graph library.

[0110] If the media asset name and the verb do not correspond, or the media asset name and the media asset type and the verb do not correspond, the implementation method is the same as the implementation method where the media asset name and the media asset type do not correspond, and this disclosure will not elaborate on it here.

[0111] For example, the user voice is "I want to watch the TV series "Borrow Another Five Hundred Years from Heaven"", the media resource name obtained is "Borrow Another Five Hundred Years from Heaven", the media resource type is "TV series", and the verb is "watch", but "Borrow Another Five Hundred Years from Heaven" is the name of a song, and the media resource type is "song", which means that the media resource name in the user voice does not correspond to the media resource type, and the content described by the user voice does not match the media resource they expect to obtain. Then, according to the media resource name, the audio-visual knowledge graph library is queried from the preset knowledge graph library to obtain candidate media resources from the audio-visual knowledge graph library.

[0112] S704: Determine a target media asset that matches the media asset type and / or verb from the candidate media assets, and control the display to display search results corresponding to the target media asset.

[0113] In some embodiments, service parameters of media asset types and verbs are pre-set, wherein the service parameters include but are not limited to: video service parameters, music service parameters, and encyclopedia service parameters, which are not specifically limited in this disclosure.

[0114] As shown in Table 2, Table 2 shows some of the pre-set service parameters corresponding to media asset types and service parameters corresponding to verbs.

[0115] Table 2

[0116]

[0117] It should be noted that Table 2 does not show the values ​​of the various service parameters corresponding to the media asset names. Searching based on the media asset name can determine the actual media asset type of the first media asset indicated by the media asset name. For example, the actual media type of the media asset named "Borrow Five Hundred Years from Heaven" is song. For details on determining the actual media type of the first media asset based on the media name, reference can be made to existing technologies and will not be elaborated upon in this disclosure.

[0118] In some embodiments, based on the pre-set media type and business parameters of the verb, the business parameters corresponding to the media type (hereinafter described as the "first media type", to be distinguished from the actual media type of the first media) included in the user voice and / or the business parameters corresponding to the verb, as well as the business parameters corresponding to the actual media type of the first media are determined, and then the target business parameters are calculated based on the business parameters corresponding to the first media type and / or the business parameters corresponding to the verb, as well as the business parameters corresponding to the actual media type, and the target business parameters include at least one type business parameter.

[0119] For example, the user voice is "I want to watch the TV series "Borrow Another Five Hundred Years from Heaven", where the video service parameter corresponding to the first media asset type "TV series" is 0.5, the video service parameter corresponding to the verb "watch" is 0.5, the video service parameter corresponding to the actual media asset type song of "Borrow Another Five Hundred Years from Heaven" is 0, and its corresponding music service parameter is 0.5, and then the target service parameters include a video service parameter of 1 and a music service parameter of 0.5, indicating that the user expects to watch the video, and the media asset name included in the voice indicates that the user expects to listen to the song. After comparing the video service parameters and the music service parameters included in the target service parameters, it is determined that the user's real intention is to watch the video, indicating that the media asset name included in the user's voice does not match the media asset actually expected by the user.

[0120] Furthermore, after calculating the target service parameters, a target media asset is determined from the candidate media assets based on the target service parameters. Optionally, if the target service parameters include at least one service parameter, the target media asset can be determined from the candidate media assets based on the larger of the service parameters. The target media asset may include the second media asset determined from the media assets to be fed back, or may also include the first media asset.

[0121] The embodiment of the present disclosure provides an implementation, and the second media asset is determined from the to-be-feedback media asset according to the target service parameter. It can be understood that the target service parameter includes at least one service parameter, for example, the target service parameter includes a video service parameter and a music service parameter. The to-be-feedback media asset is screened according to the at least one target service parameter, and the to-be-feedback media asset matching the target service type is taken as the second media asset, wherein the target service type is the type corresponding to the target service parameter, for example, a video service or a music service. Further, it is compared whether the media asset type of the second media asset is the same as the media asset type of the first media asset, if yes, it indicates that the first media asset and the second media asset are both target media assets, and the search result corresponding to the first media asset and the search result corresponding to the second media asset are controlled to be displayed. If not, the search result corresponding to the second media asset is displayed, indicating that the second media asset is the target media asset, which is the media asset that the user really expects to search.

[0122] In the above example, after the target service parameter is calculated, it is determined that the media asset name included in the user voice does not match the media asset that the user actually expects to search, and there is a deviation. The to-be-feedback media asset associated with <Tianzai Again Borrow Five Hundred Years> is determined from the candidate media asset obtained in step S603: <Kangxi Dynasty> as the second media asset. Since the media asset type of <Kangxi Dynasty> is a TV series and the media asset type of <Tianzai Again Borrow Five Hundred Years> is a song, the media asset types of the two are different, <Kangxi Dynasty> is taken as the target media asset and is fed back to the user, and the search result corresponding to <Kangxi Dynasty> is displayed.

[0123] In some embodiments, in step 703, the target knowledge graph library is determined to be a second error correction knowledge graph library according to the media asset name, and the second error correction knowledge graph library includes media assets with a pronunciation similarity greater than a second similarity threshold but different media asset contents. Optionally, the second error correction knowledge graph library includes media assets with the same pronunciation but different media asset contents, for example, the TV series <Guyu> and the encyclopedia entry “Guyu”. The pronunciations of the two are the same, both being “guyu”, but the media asset contents of the two are different. The embodiment of the present disclosure provides an implementation, and before the service parameter corresponding to the media asset type and / or the verb is determined, it is judged whether the media asset types of the first media asset and the to-be-feedback media asset are the same, if yes, the to-be-feedback media asset and the first media asset are taken as the target media asset and are fed back to the user; if not, it indicates that the media asset type included in the user voice may have a deviation, and the service parameter corresponding to the media asset type and / or the verb included in the user voice and the service parameter corresponding to the actual media asset type of the first media asset need to be further determined, and then the target service parameter is calculated according to the above service parameters, so as to determine the second media asset from the to-be-feedback media assets with the same pronunciation but different media asset contents, so as to meet the actual needs of the user.

[0124] In some embodiments, after determining the target media that matches the media type and / or verb from the candidate media, historical search records are obtained, and a first sorting weight of multiple target media assets is determined based on the historical search records; the historical search records can be associated with user information, and after recognizing the user voice in step S702, the user information corresponding to the user voice is determined. It can be understood that the voiceprint information contained in the user voice makes the user information unique, and unique user information can be determined based on the user voice. Binding the user information to the historical search records is conducive to analyzing user preferences.

[0125] In the process of displaying the search results corresponding to the target media asset, the historical search records corresponding to the user information are obtained, and the first ranking weight of the target media asset is determined based on the historical search records. The resource heat parameters of multiple target media assets are obtained, and the second ranking weights of multiple target media assets are determined based on the resource heat parameters of the multiple target media assets; the resource heat parameters are used to indicate the popularity of the target media asset, and are calculated and processed by the server based on public opinion data, wherein the public opinion data includes but is not limited to the number of plays, the number of likes, the number of comments, the number of reposts, the number of shares, etc., and this disclosure does not limit this. It can be understood that the greater the number of plays of the target media asset, the greater its resource heat parameter, and accordingly, the greater the second ranking weight.

[0126] Furthermore, a target ranking weight is calculated based on the first ranking weight and the second ranking weight. Optionally, the target ranking weight is the average of the first ranking weight and the second ranking weight. The display is controlled to display search results corresponding to the target media asset according to the target ranking weight, and the search results for the target media asset are ranked based on user preferences and media asset popularity.

[0127] For example, Figure 10 As shown, Figure 10 Schematic diagram of the user interface of the voice search provided by the embodiment of the present disclosure Figure 2 After determining that the target media resources include "Bu Ran" and "Zuo Zhi Zhi Yue", obtain the historical search records. In the historical search records, the number of plays of "Bu Ran" exceeds "Zuo Zhi Zhi Yue". The first ranking weight corresponding to "Bu Ran" is greater than the first ranking weight corresponding to "Zuo Zhi Zhi Yue". And according to the resource heat parameter, it is determined that the second ranking weight corresponding to "Bu Ran" is greater than the second ranking weight corresponding to "Zuo Zhi Zhi Yue". It is calculated that the target ranking weight corresponding to "Bu Ran" is greater than the target ranking weight corresponding to "Zuo Zhi Zhi Yue". According to the order of the target ranking weights, refer to Figure 10 As shown, the search result 11 corresponding to "Unstained" is displayed above the search result 12 corresponding to "Left Hand Pointing to the Moon".

[0128] In summary, the disclosed embodiment provides a voice search method, which recognizes the acquired user voice to obtain keywords therein, including media asset names, media asset types and / or verbs, and then first determines candidate media assets from the target knowledge graph library based on the media asset names, the first media asset indicated by the media asset name in the candidate media, and the media assets to be fed back that have an associated relationship with the first media asset, and then determines the target media asset that matches the media asset type and / or verb from the candidate media, and controls the display to display the search results corresponding to the target media asset. This method connects the contextual information in the user voice and accurately understands the user's true intentions, so that even if the media asset name indicated by the user voice does not match the media asset type and / or verb, the matching media asset search results can still be accurately fed back to the user, thereby improving the user's usage experience.

[0129] like Figure 11 As shown, Figure 11 This is a schematic diagram of the structure of a display device provided in an embodiment of the present disclosure. The display device includes a processor 1101, a memory 1102, and a computer program stored in the memory 1102 and executable on the processor 1101. When executed by the processor 1101, the computer program implements the various processes of the voice search method in the above-mentioned method embodiment. The same technical effects can be achieved, and to avoid repetition, they are not described here.

[0130] An embodiment of the present disclosure provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the above-mentioned voice search method and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0131] The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0132] The present disclosure provides a computer program product, which includes a computer program. When the computer program is run on a computer, the computer is enabled to implement the above-mentioned voice search method.

[0133] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion of some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that: include: The user input interface is configured to: obtain user voice; The controller is configured to: recognize the user voice and obtain keywords in the user voice, wherein the keywords include a media asset name, a media asset type and / or a verb; Determining candidate media assets from a target knowledge graph library according to the media asset name, the candidate media assets including a first media asset indicated by the media asset name and media assets to be fed back that are associated with the first media asset; Determine, from the candidate media assets, a target media asset that matches the media asset type and / or the verb, and control a display to display search results corresponding to the target media asset; The controller determines the target media asset that matches the media asset type and / or the verb from the candidate media assets, and is configured to: determine the business parameters corresponding to the media asset type and / or the verb, and the business parameters corresponding to the actual media asset type of the first media asset, so as to calculate the target business parameters, wherein the target business parameters are the business parameters corresponding to the user voice; and determine the target media asset from the candidate media assets based on the target business parameters.

2. The display device according to claim 1, wherein The target knowledge graph library is at least one knowledge graph library in the preset knowledge graph library; the preset knowledge graph library includes: a first error correction knowledge graph library, a second error correction knowledge graph library, and an audio-visual knowledge graph library; The first error correction knowledge graph library includes media assets whose pronunciation similarity of the media asset names is greater than a first similarity threshold but whose media asset types are different; The second error correction knowledge graph includes media assets with the same media asset names but different media asset contents, the pronunciation similarity of which is greater than a second similarity threshold, and the second similarity threshold is greater than the first similarity threshold; The audio-visual knowledge graph includes film and television media assets and music media assets, and there is a corresponding relationship between the film and television media assets and the music media assets.

3. The display device according to claim 1, wherein There are multiple target media assets; The controller controls the display to display the search results corresponding to the target media asset, and is configured to: Acquire historical search records, and determine first ranking weights of the plurality of target media assets according to the historical search records; Acquire resource heat parameters of the plurality of target media assets, and determine second ranking weights of the plurality of target media assets according to the resource heat parameters of the plurality of target media assets; Calculating a target ranking weight according to the first ranking weight and the second ranking weight; The display is controlled to display the search results corresponding to the target media asset according to the target ranking weight.

4. The display device according to claim 1, wherein The controller determines the target media asset from the candidate media assets according to the target service parameter, and is configured to: determining a second media asset from the media assets to be fed back according to the target service parameter; If the media asset type of the second media asset is the same as the media asset type of the first media asset, controlling the display to display the search results corresponding to the first media asset and the search results corresponding to the second media asset; If the media asset type of the second media asset is different from the media asset type of the first media asset, the display is controlled to display the search result corresponding to the second media asset.

5. The display device according to claim 1, wherein The target knowledge graph library is a second error correction knowledge graph library; the second error correction knowledge graph includes media assets whose pronunciation similarity of the media asset names is greater than a second similarity threshold but whose media asset contents are different; Before the controller determines the service parameters corresponding to the media asset type and / or the verb, and the service parameters corresponding to the actual media asset type of the first media asset to calculate the target service parameters, the controller is further configured to: Determining whether the first media asset and the media asset to be fed back are of the same media asset type; The controller determines the service parameters corresponding to the media asset type and / or the verb, and the service parameters corresponding to the actual media asset type of the first media asset to calculate the target service parameters, and is configured to: When the first media asset and the media asset to be fed back are of different media asset types, the service parameters corresponding to the media asset type and / or the verb and the service parameters corresponding to the actual media asset type of the first media asset are determined to calculate the target service parameters.

6. The display device according to claim 1, wherein The controller, after recognizing the user voice and obtaining keywords in the user voice, and before determining candidate media assets from a target knowledge graph library based on the media asset names, is further configured to: Determining whether the media asset name corresponds to the media asset type, and / or whether the media asset name corresponds to the verb; If the media asset name and the media asset type do not correspond, and / or the media asset name and the verb do not correspond, the target knowledge graph library is determined according to the media asset name.

7. A voice search method, characterized in that: include: Get user voice; Recognizing the user's voice and obtaining keywords in the user's voice, wherein the keywords include a media asset name, a media asset type, and / or a verb; Determining candidate media assets from a target knowledge graph library according to the media asset name, the candidate media assets including a first media asset indicated by the media asset name and media assets to be fed back that are associated with the first media asset; Determine, from the candidate media assets, a target media asset that matches the media asset type and / or the verb, and control a display to display search results corresponding to the target media asset; Determine the target media asset that matches the media asset type and / or the verb from the candidate media assets, including: determining the business parameters corresponding to the media asset type and / or the verb, and the business parameters corresponding to the actual media asset type of the first media asset, to calculate the target business parameters, wherein the target business parameters are the business parameters corresponding to the user voice; and determine the target media asset from the candidate media assets based on the target business parameters.

8. The method according to claim 7, characterized in that The target knowledge graph library is at least one knowledge graph library in the preset knowledge graph library; the preset knowledge graph library includes: a first error correction knowledge graph library, a second error correction knowledge graph library, and an audio-visual knowledge graph library; The first error correction knowledge graph library includes media assets whose pronunciation similarity of the media asset names is greater than a first similarity threshold but whose media asset types are different; The second error correction knowledge graph includes media assets with the same media asset names but different media asset contents, the pronunciation similarity of which is greater than a second similarity threshold, and the second similarity threshold is greater than the first similarity threshold; The audio-visual knowledge graph includes film and television media assets and music media assets, and there is a corresponding relationship between the film and television media assets and the music media assets.

9. A computer-readable storage medium, characterized in that include: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the voice search method according to any one of claims 7 to 8.

Citation Information

Patent Citations

  • Program recommendation method and device, equipment and computer storage medium

    CN112333477A

  • Display device and media asset display method

    CN115150673A