Display device and voice recognition method

By acquiring the entities and sentence templates to be translated from the display device, translating them into the target language and generating a speech recognition model, the problem of inaccurate translation caused by literal translation is solved, and accurate recognition of user speech is achieved, thus improving the user experience.

CN115602167BActive Publication Date: 2026-06-02VIDAA INT HLDG (NETHERLANDS) CO

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIDAA INT HLDG (NETHERLANDS) CO
Filing Date
2022-09-29
Publication Date
2026-06-02

Smart Images

  • Figure CN115602167B_ABST
    Figure CN115602167B_ABST
Patent Text Reader

Abstract

Some embodiments of the present application provide a display device and a voice recognition method. The display device obtains a to-be-translated sentence based on a to-be-translated entity and a pre-stored sentence template, and translates the to-be-translated sentence into a target sentence in a preset language. The display device obtains a target entity corresponding to the to-be-translated entity based on the target sentence, and generates a voice recognition model according to the target entity and the sentence template. For the collected user voice, the display device recognizes the user voice based on the voice recognition model, obtains a control instruction corresponding to the user voice, and executes the control instruction. For user voice in different languages, the display device can recognize the voice itself instead of using a direct translation method, so that the voice of the user can be accurately recognized, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and more particularly to a display device and a voice recognition method. Background Technology

[0002] Display devices refer to terminal devices capable of outputting specific display images. With the rapid development of display devices, their functions are becoming increasingly rich and their performance increasingly powerful, enabling two-way human-computer interaction and integrating multiple functions such as audio-visual, entertainment, and data processing to meet diverse and personalized user needs. Users can also utilize the display device's voice recognition function to control the device via voice.

[0003] When a user controls a display device using voice, the device needs to recognize the user's voice to determine the control commands. Generally, the display device may store corpus information corresponding to the local language, such as sentence templates and word entities. Based on this corpus information, it can recognize some speech sounds corresponding to the local language. However, users may have a need to control the display device in other languages. In this case, the display device can translate the stored corpus information into the user's current language and use the translated corpus information to recognize the user's voice.

[0004] However, display devices often use literal translation when translating expected information. This method may result in a translation that deviates from the original meaning of the expected information, leading to inaccurate translations and an inability to accurately recognize the user's speech, which seriously affects the user experience. Summary of the Invention

[0005] This application provides a display device and a speech recognition method in some embodiments. This addresses the problem in related technologies where literal translation of speech data leads to inaccurate translations, resulting in an inability to accurately recognize user speech and severely impacting the user experience.

[0006] In a first aspect, some embodiments of this application provide a display device, including a display, a sound acquisition unit, and a controller. The sound acquisition unit is configured to acquire user-input voice; the controller is configured to perform the following steps:

[0007] The sentence to be translated is obtained based on the sentence templates pre-stored in the entity to be translated and the display device;

[0008] Translate the statement to be translated into a target statement in a preset language;

[0009] Based on the target statement, obtain the target entity corresponding to the entity to be translated;

[0010] A speech recognition model is generated based on the target entity and the pre-stored sentence template;

[0011] In response to the user's voice collected by the sound collector, the user's voice is recognized based on the speech recognition model to obtain the control command corresponding to the user's voice and execute the control command.

[0012] Secondly, some embodiments of this application provide a speech recognition method applied to a display device, including:

[0013] The sentence to be translated is obtained based on the sentence templates pre-stored in the entity to be translated and the display device;

[0014] Translate the statement to be translated into a target statement in a preset language;

[0015] Based on the target statement, obtain the target entity corresponding to the entity to be translated;

[0016] A speech recognition model is generated based on the target entity and the pre-stored sentence template;

[0017] In response to the user's voice collected by the sound collector, the user's voice is recognized based on the speech recognition model to obtain the control command corresponding to the user's voice and execute the control command.

[0018] As can be seen from the above technical solutions, some embodiments of this application provide a display device and a speech recognition method. The display device obtains a sentence to be translated based on the entity to be translated and a pre-stored sentence template, and translates the sentence into a target sentence in a preset language. The display device obtains the target entity corresponding to the entity to be translated based on the target sentence, and generates a speech recognition model based on the target entity and the sentence template. For the collected user speech, the display device recognizes the user speech based on the speech recognition model, obtains the control commands corresponding to the user speech, and executes them. For user speech in different languages, the display device can recognize it automatically, rather than using a direct translation method, thus accurately recognizing the user's speech and improving the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This illustrates a use case of a display device according to some embodiments;

[0021] Figure 2A hardware configuration block diagram of a control device according to some embodiments is shown;

[0022] Figure 3 A hardware configuration block diagram of a display device according to some embodiments is shown;

[0023] Figure 4 A software configuration diagram in a display device according to some embodiments is shown;

[0024] Figure 5 A schematic diagram of the application panel in some embodiments is shown;

[0025] Figure 6 A schematic diagram of the voice interaction network architecture of the display device is shown in some embodiments;

[0026] Figure 7 A schematic diagram of the system settings UI interface in some embodiments is shown;

[0027] Figure 8 This diagram illustrates the display of voice recognition mode confirmation information on a screen in some embodiments;

[0028] Figure 9 The diagrams show the interaction flowcharts of various components of the display device in some embodiments;

[0029] Figure 10 A schematic diagram of the language selection interface in some embodiments is shown;

[0030] Figure 11 The illustrations show some scenarios of users interacting with display devices via voice in some embodiments;

[0031] Figure 12 A schematic diagram of a display device showing a search interface is shown in some embodiments;

[0032] Figure 13 The diagram shows some examples of prompting information;

[0033] Figure 14 Flowcharts of speech recognition methods in some embodiments are shown. Detailed Implementation

[0034] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0035] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0036] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0037] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0038] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.

[0039] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.

[0040] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0041] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.

[0042] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.

[0043] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.

[0044] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.

[0045] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0046] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.

[0047] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.

[0048] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.

[0049] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.

[0050] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.

[0051] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).

[0052] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0053] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.

[0054] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.

[0055] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0056] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can execute operations related to the object selected by the user command.

[0057] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0058] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0059] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0060] like Figure 4 As shown, the display device system is divided into three layers, from top to bottom: the application layer, the middleware layer, and the hardware layer.

[0061] The application layer mainly includes commonly used applications on TVs, as well as the application framework. The commonly used applications are mainly browser-based applications, such as HTML5 apps, and native apps.

[0062] An application framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interface for these functions (toolbar, status bar, menu, dialog box).

[0063] Native apps can support online or offline access, push notifications, or access to local resources.

[0064] The middleware layer includes various television protocols, multimedia protocols, and system components. Middleware can use the basic services (functions) provided by system software to connect different parts of application systems or different applications on the network, achieving resource sharing and function sharing.

[0065] The hardware layer mainly includes the HAL interface, hardware, and drivers. The HAL interface is a unified interface for all TV chips, with the specific logic implemented by each chip. The drivers mainly include: audio drivers, display drivers, Bluetooth drivers, camera drivers, Wi-Fi drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers.

[0066] In some embodiments of this application, the display 260 may display a user interface. The user interface may contain a specific target image, such as various media resources obtained from a network signal source, including videos, images, and other content. The user interface may also be some UI interface of the display device, such as a system recommendation page.

[0067] Display devices can have a variety of functions, such as playing media, playing games, and video chatting, thus providing users with a wide range of services.

[0068] In some embodiments, when the user controls the display device to power on, the controller 250 can control the display 260 to display the user interface. The user interface may include a "My Applications" control. The user can click the "My Applications" control to input a display command for the application panel page, thereby triggering entry into the corresponding application panel. It should be noted that the user can also use other methods to input a selection operation on the function control to trigger entry into the application panel. For example, using voice control or search functions, etc., to control entry into the application panel page.

[0069] Users can view the applications installed on the display device through the application panel, which represents the functions supported by the display device. Users can select and open one of the applications to perform its functions. It should be noted that the applications installed on the display device can be system applications or third-party applications. By opening an application, users control the display device to perform the corresponding functions of that application. Figure 5 Schematic diagrams of the application panel in some embodiments are shown. For example... Figure 5As shown, the application panel includes three controls: "Player," "Cable TV," and "Video Chat." Users can click the "Player" control to open the player application on the display device. Users can perform various operations within the player, such as searching for media assets. Users can click the "Cable TV" control to watch various media channels on the display device, including programs provided by cable TV providers. Users can click the "Video Chat" control to engage in video chat on the display device.

[0070] Users can use control devices, such as remote controls or mobile terminals, to input commands into the display device to control it and perform various functions. Users can use the control device to move the focus on the monitor to select different controls and open them. Users can also use the control device to input text into the display device, such as entering the name of a media asset when searching for it.

[0071] In some embodiments, considering the user experience, the display device has a voice recognition function, so that the user can input control commands to the display device by voice input to realize voice interaction.

[0072] Figure 6 A schematic diagram of the voice interaction network architecture of a display device is shown in some embodiments. For example... Figure 6 As shown, the display device 200 is used to receive input information such as sound and output the processing results of that information. The speech recognition module deploys an Automatic Speech Recognition (ASR) service to recognize audio as text; the semantic understanding module deploys a Natural Language Understanding (NLU) service to perform semantic parsing on the text; the business management module deploys a business instruction management service such as Dialog Management (DM) to provide business instructions; the language generation module deploys a Natural Language Understanding (NLG) service to convert instructions to be executed by the display device into text language; and the speech synthesis module deploys a Text-to-Speech (TTS) service to process the text language corresponding to the instructions and send it to the speaker for playback. The voice interaction network architecture can contain multiple entity service devices deploying different business services, or one or more entity service devices can combine one or more functional services.

[0073] In some embodiments, the following describes the basis Figure 6The process of processing information input to the display device 200 in the architecture shown is described with an example, taking a query statement input via voice as an example:

[0074] Voice recognition: After receiving a query statement input by voice, the display device 200 can perform noise reduction processing and feature extraction on the audio of the query statement. The noise reduction processing may include steps such as removing echoes and environmental noise.

[0075] Semantic understanding: Natural language understanding is performed on the identified candidate text and associated contextual information. The text is parsed into structured, machine-readable information, business domain, intent, slots, etc., to express semantics, and an executable intent is determined. An intent confidence score is obtained, and the semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.

[0076] Business Management: Based on the semantic parsing results of the query statement text, the semantic understanding module sends query instructions to the corresponding business management module to obtain the query results provided by the business service, as well as the actions required to "complete" the user's final request, and feeds back the device execution instructions corresponding to the query results.

[0077] Language generation: This is configured to generate spoken text from information or instructions. Specifically, it can be divided into casual conversation, task-oriented, knowledge-based question-answering, and recommendation-oriented systems. In casual conversation, NLG (Natural Language Generation) performs intent recognition and sentiment analysis based on context, then generates open-ended responses. In task-oriented conversations, learned strategies are used to generate responses, typically including clarifying needs, guiding the user, asking questions, confirming, and closing remarks. In knowledge-based question-answering conversations, the required knowledge (knowledge, entities, fragments, etc.) is generated based on question type identification and classification, information retrieval, or text matching. In recommendation-oriented conversation systems, user interests are matched, candidate recommendations are ranked, and then recommended content is generated for the user.

[0078] Speech synthesis: The voice output configured to be presented to the user. The speech synthesis processing module synthesizes speech output based on text provided by the digital assistant. For example, the generated dialogue response is in the form of a text string. The speech synthesis module converts the text string into audible speech output.

[0079] It should be noted that, Figure 6 The architecture shown is merely an example and is not intended to limit the scope of protection of this application. Other architectures can also be used to achieve similar functions in the embodiments of this application. For example, all or part of the above process can be performed by the display device 200, which will not be elaborated here.

[0080] In some embodiments, the voice recognition function can be implemented by a sound acquisition device and a controller 250 set on the display device, and the semantic function can be implemented by the controller 250 of the display device.

[0081] Users can use control devices, such as remote controls, to control the display device 200. For example, with a smart TV, users can use the remote control to control the TV to play media or adjust the volume, thereby controlling the smart TV.

[0082] In some embodiments, the display device 200 has a voice recognition function. When the display device 200 enables the voice recognition function, the user can send voice commands to the display device 200 via voice input, thereby enabling the display device 200 to perform the corresponding functions. Therefore, the display device 200 may be equipped with a voice recognition mode.

[0083] In some embodiments, a user can send a voice recognition mode command to the display device by operating a designated button on the remote control. In practical applications, a pre-defined correspondence is established between the voice recognition mode command and the remote control button. For example, a voice recognition mode button is provided on the remote control. When the user touches this button, the remote control sends a voice recognition mode command to the controller 250, at which point the controller 250 controls the display device to enter voice recognition mode. When the user touches the button again, the controller 250 can control the display device to exit voice recognition mode.

[0084] In some embodiments, when a user controls a display device using a smart device, such as a mobile phone, they can also send a voice recognition mode command to the display device. In practical applications, a control can be set in the mobile phone, which allows the user to select whether to enter voice recognition mode, thereby sending a voice recognition mode command to the controller 250. At this time, the controller 250 can control the display device to enter voice recognition mode.

[0085] In some embodiments, when a user controls a display device using a mobile phone, they can issue a continuous click command to the phone. A continuous click command refers to the user clicking the same area of ​​the phone's touchscreen more than a preset threshold number of times within a preset period. For example, if a user clicks a certain area of ​​the phone's touchscreen three times consecutively within 1 second, it is considered one continuous click command. After receiving the continuous click command, the mobile phone can send a voice recognition mode command to the display device, so that the controller 250 controls the display device to enter voice recognition mode.

[0086] In some embodiments, when a user controls a display device using a mobile phone, the mobile phone can be configured to send a voice recognition mode command to the display device when the user's touch pressure on a certain area of ​​the mobile phone touchscreen exceeds a preset pressure threshold.

[0087] You can also set a voice recognition mode option in the UI of the display device. When the user clicks this option, they can control the display device to enter or exit voice recognition mode. Figure 7 Schematic diagrams of the system settings UI interface in some embodiments are shown. For example... Figure 7 As shown, the system settings include screen settings, sound settings, voice recognition settings, network settings, and factory reset. Users can click the voice recognition control to control the display device 200 to enter or exit voice recognition mode.

[0088] In some embodiments, to prevent users from accidentally triggering the voice recognition mode, when the controller 250 receives a voice recognition mode instruction, it can control the display 260 to display voice recognition mode confirmation information, thereby allowing the user to make a secondary confirmation as to whether to control the display device to enter the voice recognition mode. Figure 8 A schematic diagram showing voice recognition mode confirmation information displayed on a display in some embodiments is shown.

[0089] In some embodiments, after the display device 200 enters voice recognition mode, the user can directly send commands to the display device 200 via voice input. After receiving the user's voice input, the display device 200 can recognize the user's voice, determine the user's control command, and execute the corresponding operation to achieve the function required by the user.

[0090] In some embodiments, the controller 250 can control a sound acquisition device, which can be a microphone, to acquire the user's input voice signal. After the sound acquisition device acquires the user's voice, the controller 250 can parse the user's voice to obtain the corresponding speech text. The controller can perform semantic analysis on the speech text to determine the user's control commands.

[0091] In some embodiments, the display device 200 may further include a third-party voice recognition interface. Upon receiving voice input from a user, the controller 250 can send the voice data to the third-party voice recognition interface, using a third-party voice recognition device or the like to convert the user's voice into speech-to-text. After obtaining the speech-to-text, the controller 250 can parse the speech-to-text and execute the voice command.

[0092] In some embodiments, the controller 250 may also send voice commands to a server. The server can generate voice text based on the voice commands and then send the voice text back to the display device.

[0093] When a display device recognizes user speech, it can utilize stored predictive information to perform semantic analysis on the corresponding speech text. The display device can store high-quality sentence templates and word entities that reflect frequently input speech patterns. This predictive information can be used as training data to build a speech recognition model, which can then be used to recognize user speech.

[0094] It should be noted that since display devices are typically used in a fixed area, the pre-defined information stored in them may only be in the local language or the language corresponding to the device's manufacturing location. However, users may use different languages ​​and therefore have a need to control the display device in other languages. The language corresponding to the user's voice may differ from the language corresponding to the pre-defined information stored on the display device; therefore, the pre-defined information stored on the display device alone cannot recognize the user's voice. To address this, the display device can translate the stored speech data into the user's current language, thereby using the translated speech data to recognize the user's voice.

[0095] Display devices can use existing translation software to translate intended information. However, this software may rely on literal translation. This approach deviates from the original meaning of the intended information, resulting in inaccurate translations. For example, in the phrase "Wake me up at twelve o'clock," the word "twelve o'clock" clearly indicates time. A literal translation of "twelve o'clock" might yield "twelve points," representing twelve points, not the time itself. Such translations fail to accurately recognize the user's speech, severely impacting the user experience.

[0096] Therefore, the display device provided in this application embodiment can accurately translate expected information, such as various entities, thereby accurately recognizing user voice.

[0097] In some embodiments, the display device may pre-translate stored pre-defined information, such as various word entities, into different languages ​​before the user uses the display device's speech recognition function. Alternatively, the user may specify one or more languages, and the display device may then translate the pre-defined information into the corresponding languages.

[0098] The display device can generate a speech recognition model based on the translated expected information, and use the speech recognition model to recognize the user's voice in order to respond to the user's control commands.

[0099] Figure 9 The diagram illustrates the interaction flowcharts of various components of the display device in some embodiments. For example... Figure 9As shown, it includes the following steps:

[0100] S101, Controller 250 obtains the statement to be translated based on the statement templates pre-stored in the entity to be translated and the display device.

[0101] S102, Controller 250 translates the statement to be translated into a target statement in a preset language.

[0102] S103, Controller 250 obtains the target entity corresponding to the entity to be translated based on the target statement.

[0103] S104, Controller 250 generates a speech recognition model based on the target entity and pre-stored sentence templates.

[0104] S105. In response to the user's voice collected by the sound collector, the controller 250 recognizes the user's voice based on the voice recognition model, obtains the control command corresponding to the user's voice, and executes the control command.

[0105] In some embodiments, the display device may pre-store pre-defined information, including information such as statement templates and word entities.

[0106] A manual acquisition method can be used to first obtain some high-frequency and high-quality sentence templates and entities. For example, the text corresponding to the most frequently input voice commands by different users when using voice-controlled display devices can be statistically analyzed, and some sentence templates and entities can be determined based on these texts. Taking the language of the region where the display device is located as Chinese as an example, when users use voice control to control the display device, they may be searching for media materials, playing media materials, or setting alarms. In this embodiment, the alarm function is used as an example to describe the process of translating expected information.

[0107] Users can use voice control to set alarm clock functions on the display device. Based on user experience, some statement templates can be summarized, with the corresponding language being Chinese. The statement templates can contain fixed entities and slots to be filled. Fixed entities refer to the instruction text used by the user to control the display device to perform the corresponding function. For example, when controlling the display device to perform the alarm clock function, the fixed entity could be "Wake me up" or "Wake me up." Slots to be filled are the slots in the statement templates that await the filling of function parameters. Different types of entities can be filled in different slots, such as time-type entities or date-type entities.

[0108] In this embodiment, the fixed entity in the statement template corresponding to the Chinese language is referred to as the first fixed entity. The statement template is described in detail below:

[0109] Statement template A: '@{StartDate}@{StartTime} wake me up'. Here, "wake me up" is the first fixed entity, "@{StartDate}" is a slot to be filled for date type entities, and "@{StartTime}" is another slot to be filled for time type entities.

[0110] Statement Template B: 'Set an alarm for @{StartDate}@{StartTime}'. Here, "Set an alarm for..." is the first fixed entity, and "@{StartDate}" and "@{StartTime}" are the slots to be filled for date type entities and time type entities, respectively.

[0111] Statement template C: '@{StartDate}@{StartTime}Please wake me up with @{Song}'. Here, "Please wake me up with" is the first fixed entity, "@{StartDate}" and "@{StartTime}" are the slots to be filled for date type entities and time type entities, respectively, and "@{Song}" is the slot to be filled for song type entities.

[0112] Statement template D: '@{StartDate}@{StartTime} Set an alarm clock with @{Singer}'. Here, "use", "of", and "set an alarm clock" are the first fixed entities; "@{StartDate}" and "@{StartTime}" are the slots to be filled for date type entities and time type entities, respectively; "@{Singer}" is the slot to be filled for singer type entities; and "@{Song}" is the slot to be filled for song type entities.

[0113] When a user inputs the voice corresponding to one of the four statement templates mentioned above into the display device, the display device can determine the statement template corresponding to the user's voice and the entity in the slot to be filled. For example, if the user's voice is "Wake me up at four o'clock tomorrow," then the display device can determine that the user's voice corresponds to statement template A, and the entity information in "@{StartDate}" is "tomorrow," and the entity information in "@{StartTime}" is "four o'clock," thus determining that the user needs to activate the alarm clock function on the display device at four o'clock tomorrow.

[0114] In this embodiment of the application, the above-mentioned statement template is also called a seed template, which can be a high-quality template obtained manually and can reflect the voice templates that users use more frequently.

[0115] In some embodiments, for the slots to be filled in the above statement template, some entities corresponding to each slot can be pre-determined. These entities can be determined manually, using high-frequency and high-quality data. It should be noted that each slot can match one entity type; for example, "@{StartDate}" and "@{StartTime}" match date and time types, respectively. For each slot, some entities can be pre-obtained, referred to as seed entities in this embodiment. For example, the seed entity corresponding to "@{StartDate}" could be "tomorrow," the seed entity corresponding to "@{StartTime}" could be "morning," the seed entity corresponding to "@{Singer}" could be "Eagles," and the seed entity corresponding to "@{Song}" could be "hotel."

[0116] In some embodiments, considering the user's need to control the display device using other languages, the display device needs to be able to recognize user speech in other languages. Therefore, the seed template and seed entity described above can be pre-translated into other languages. This application embodiment uses the translation of Chinese into English as an example for illustration.

[0117] After translating the seed template into English, we can obtain the corresponding English seed template, including:

[0118] Statement template A: 'wake me up at@{StartTime}@{StartDate}'.

[0119] Statement template B: 'set an alarm for@{StartTime}@{StartDate}'.

[0120] Statement template C: 'wake me up with @{Song}at@{StartTime}@{StartDate}'.

[0121] Statement template D: 'set an alarm clock for @{StartTime}@{StartDate}with@{Singer}'s@{Song}'.

[0122] It should be noted that manual translation can be used when translating seed templates. After translation, the slots to be filled will not change, but the fixed entities will be translated into English. Both seed templates and seed entities can be translated.

[0123] In some embodiments, a large amount of entity information may be needed as training data to build a speech recognition model. Considering the large workload of manually translating all entities, a number of entities can be manually selected as seed entities for translation. The remaining entities are used as training data, which can then be automatically translated into other languages ​​by the display device. To avoid problems caused by using existing translation software for literal translation of entities, the display device in this embodiment can automatically translate the entities to be translated. This embodiment uses the translation of Chinese entities into English entities as an example to illustrate the entity translation process.

[0124] The controller 250 of the display device can first obtain the statement to be translated based on the entity to be translated and the statement template pre-stored in the display device.

[0125] Specifically, a database can be pre-configured in the display device to store specific data. Once seed templates, seed entities, and translated seed templates and seed entities are obtained, this corpus information can be stored in the pre-configured database for subsequent applications.

[0126] The controller 250 can obtain the target entity type of the entity to be translated. For example, for the entity "12 o'clock", since it is the entity information corresponding to the user's control of the display device, the entity type of the entity "12 o'clock" is time type, that is, the target entity type is time.

[0127] After determining the target entity type, the controller 250 can obtain the seed template corresponding to the target entity type.

[0128] The database contains several pre-stored statement templates, also known as seed templates. The controller 250 can filter these templates based on the target entity type to obtain the statement template corresponding to that target entity type, referred to as the target statement template in this embodiment. It should be noted that during this filtering process, the initial language corresponding to the entity to be translated, such as Chinese, needs to be determined first, along with all the corresponding seed templates. Then, the target statement template corresponding to the target entity type is obtained by filtering among these Chinese seed templates.

[0129] In some embodiments, when filtering statement templates, the controller 250 can filter statement templates in the database based on preset filtering conditions.

[0130] Since we need to obtain the seed template corresponding to the target entity type, the seed template needs to contain the slots to be filled corresponding to the target entity type.

[0131] In the statement templates corresponding to the initial language, these templates include a first fixed entity and several slots to be filled. There is a matching relationship between the slots to be filled and the entity types; that is, the slot to be filled corresponding to the target entity type matches the target entity type. For example, the time type matches the "@{StartTime}" slot, and the date type matches the "@{StartDate}" slot.

[0132] Therefore, the preset filtering criteria can be set as follows: if at least one unfilled slot in a statement template matches the target entity type, then the statement template can be determined as a target statement template. That is, the controller 250 can obtain all statement templates containing unfilled slots matching the target entity type from the statement templates corresponding to the initial language in the database, and use them as target statement templates.

[0133] In some embodiments, after obtaining the target statement template, the controller 250 can obtain the statement to be translated based on the entity to be translated and the target statement template. Each target statement template corresponds to one statement to be translated.

[0134] The controller 250 can fill the slots in the target statement template to obtain a complete statement. Therefore, the controller 250 can first analyze the slots in each target statement template to determine the slots corresponding to the entity to be translated and other slots.

[0135] The controller 250 can determine the slot corresponding to the entity to be translated in the target statement template, referred to as the target slot to be filled in this embodiment. The controller 250 can also determine the slots other than the target slot to be filled, referred to as the remaining slots to be filled in this embodiment. The controller 250 can fill the target slot and the remaining slots to be filled to obtain a complete statement.

[0136] When filling the slots to be filled, the entity to be translated needs to be filled into the corresponding target slot, and other related entities need to be filled into the remaining slots to be filled. Therefore, the controller 250 needs to obtain the seed entity corresponding to the remaining slots to be filled.

[0137] In some embodiments, the controller 250 can retrieve seed entity pairs corresponding to the remaining slots to be filled from a database. A seed entity pair refers to the different representations of the same seed entity in two languages. A seed entity pair may include a first seed entity in the initial language corresponding to the entity to be translated, and a second seed entity in a preset language. The initial language is the language corresponding to the entity to be translated itself, and the preset language is the target language into which the entity to be translated needs to be translated. For example, when translating a Chinese entity to be translated into English, the initial language is Chinese, and the preset language is the target language.

[0138] The controller 250 can determine the entity type corresponding to all remaining slots to be filled in the target statement template, and obtain the seed entity of the corresponding entity type from the database. It can obtain two representations of the seed entity in the initial language and the preset language, that is, obtain the first seed entity and the second seed entity, and obtain the seed entity pair corresponding to each slot to be filled.

[0139] The controller 250 can fill the target slot with the entity to be translated, and at the same time fill the remaining slots with the first seed entity to obtain the complete sentence, which is referred to as the sentence to be translated in this embodiment.

[0140] For each target statement template, the corresponding statement to be translated can be obtained. Taking statement template A as an example, the statement to be translated after filling the slots is: Wake me up at noon tomorrow.

[0141] In some embodiments, after obtaining the statements to be translated, the controller 250 can translate all the statements to be translated into a target statement in a preset language.

[0142] Taking the sentence to be translated, "Wake me up at twelve o'clock tomorrow," as an example, the target sentence after translation could be "Wake me up at twelve o'clock tomorrow."

[0143] For statement template B, the corresponding statement to be translated is "Set an alarm for 12 o'clock tomorrow". During the translation process, the same entity in different statements, such as the entity to be translated "12 o'clock", may yield different translation results. Therefore, the target statement after translating the statement "Set an alarm for 12 o'clock tomorrow" could be "set an alarm for 12 o'clock tomorrow". That is, the entity to be translated "12 o'clock" can be translated as "twelve o'clock", possibly as "12 o'clock", or even as the purely numerical representation "12" in different statements.

[0144] The controller 250 can obtain the target entity corresponding to the entity to be translated based on the target statement. The target entity is the representation of the entity to be translated in a preset language.

[0145] In some embodiments, the controller 250 may first perform word segmentation on the target statement to obtain the word segmentation result. Each target statement will have one word segmentation result.

[0146] Taking the target sentence "Wake me up at twelve o’clock tomorrow" as an example, the corresponding word segmentation result is: wake, me, up, at, twelve, o’clock, tomorrow.

[0147] The controller 250 can obtain the candidate entity corresponding to each target sentence based on the word segmentation result. Each target sentence contains a candidate entity, which is the representation of the entity to be translated in the target sentence.

[0148] In some embodiments, when obtaining the candidate entity, the controller 250 can screen the word segmentation result to determine the candidate entity corresponding to the entity to be translated.

[0149] The controller 250 can obtain the representation form of the target sentence template corresponding to the entity to be translated in the preset language in the prior database, which is called the matching sentence template in the embodiments of the present application. That is, the target sentence template and the matching sentence template are two representation forms in the initial language and the preset language respectively.

[0150] For the sentence template A, its Chinese form is '@{StartDate}@{StartTime} wake me up', and its English form is 'wake me up at@{StartTime}@{StartDate}'. The above two templates are the corresponding situations of the target sentence template and the matching sentence template.

[0151] In the matching sentence template, it contains both fixed entities and several slots to be filled. Only the fixed entities in the matching sentence template are in the preset language, while the fixed entities in the target sentence template are in the initial language. In the embodiments of the present application, the fixed entity in the matching sentence template is called the second fixed entity, and the second fixed entity is the entity corresponding to the first fixed entity in the preset language.

[0152] The controller 250 can screen the word segmentation result. By obtaining the remaining word segments in the word segmentation result except for the second fixed entity and the second seed entity, the remaining word segments are determined as the candidate entities corresponding to each target sentence. For example, for the word segmentation result "wake, me, up, at, twelve, o’clock, tomorrow", where "wake, me, up, at" are the second fixed entities and "tomorrow" is the second seed entity, then "twelve o’clock" is the candidate entity.

[0153] In some embodiments, the controller 250 can extract the candidate entity corresponding to the entity to be translated based on the reward and punishment mechanism.

[0154] The controller 250 can assign an initial score to all entities in the word segmentation result, and the initial score can be 1.

[0155] Therefore, the initial scores for all entities in the word segmentation results are: wake:1, me:1, up:1, at:1, twelve:1, o'clock:1, tomorrow:1.

[0156] Controller 250 can penalize fixed entities, i.e., the second fixed entity in the matching statement template, by deducting 1 point from each. At this point, the scores for all entities are: wake:0, me:0, up:0, at:0, twelve:1, o'clock:1, tomorrow:1.

[0157] Controller 250 can continue to penalize the seed entity, namely the second seed entity "tomorrow", by deducting 1 point. It should be noted that the minimum score can be set to 1; when the score is already 0, the score after penalty remains 0. At this point, the scores of all entities are: wake:0, me:0, up:0, at:0, twelve:1, o'clock:1, tomorrow:0.

[0158] The controller 250 can acquire all entities whose scores have not changed (score is 1) and determine candidate entities. Therefore, the candidate entity is twelve o'clock.

[0159] In some embodiments, candidate entities corresponding to each target statement can be obtained by filtering the word segmentation results corresponding to all target statements. Since the same entity in different statements may have different results after translation, the candidate entities corresponding to different target statements may be the same or different. The controller 250 can organize the candidate entities corresponding to all target statements to form a candidate entity set. In some embodiments, the candidate entity set is ['12', '12o'clock', 'twelve o'clock', 'twelve o'clock']. It may contain multiple identical candidate entities.

[0160] The controller 250 can filter all candidate entities to obtain the candidate entity that is closest to the entity to be translated, which is then used as the final translation result.

[0161] The controller 250 can translate all candidate entities into corresponding entities in the initial language, referred to as initial language candidate entities in this embodiment. That is, each candidate entity is translated to obtain each initial language candidate entity.

[0162] The controller 250 can obtain the edit distance for each initial language candidate entity and the entity to be translated. Edit distance, also known as Levenshtein distance, is a quantitative measure of the difference between two strings (e.g., English characters). It measures the minimum number of processing steps required to transform one string into another. Edit distance can be used in natural language processing; for example, spell checking can determine which (or which characters) is more likely to be misspelled based on the edit distance between a misspelled character and other correct characters.

[0163] At the same time, the controller 250 can obtain the number of times each candidate entity appears in all candidate entities.

[0164] The controller 250 can calculate the translation score of each candidate entity based on the edit distance, the number of occurrences, and the preset weight coefficients, and determine the candidate entity with the highest translation score as the target entity corresponding to the entity to be translated.

[0165] Edit distance and occurrence frequency can be pre-assigned weight coefficients; for example, the weight coefficient for edit distance could be 0.3, and the weight coefficient for occurrence frequency could be 0.7.

[0166] The controller 250 can calculate the translation score for each candidate entity according to formula (1).

[0167] F = a*S + b*G (1)

[0168] in,

[0169] F represents the translation score, S represents the edit distance between the initial language candidate entity and the entity to be translated, and G represents the number of times the candidate entity appears in the candidate entity set.

[0170] a represents the weighting coefficient of edit distance, b represents the weighting coefficient of occurrence frequency, and a+b=1.

[0171] After calculating the translation score for each candidate entity, the candidate entity with the highest score can be identified as the target entity, that is, the entity to be translated in the preset language.

[0172] In some embodiments, after obtaining the target entity corresponding to the entity to be translated, the entity to be translated, the target entity, and the sentence template can be used as training data, and a speech recognition model can be generated based on this training data to recognize the user's speech.

[0173] In some embodiments, a corresponding speech recognition template can be generated for each language.

[0174] The controller 250 can first determine the initial language of the entity to be translated.

[0175] The controller 250 can retrieve all statement templates corresponding to the initial language from the database, referred to as the first statement template in this embodiment. It can also retrieve all statement templates corresponding to a preset language, referred to as the second statement template in this embodiment.

[0176] The controller 250 can generate a first speech recognition model based on the entity to be translated and a first sentence template. The first speech recognition model is used to recognize the user's speech corresponding to the initial language. The controller 250 can generate a second speech recognition model based on the target entity and a second sentence template. The second speech recognition model is used to recognize the user's speech corresponding to a preset language.

[0177] In some embodiments, the controller 250 can synthesize training data from all languages ​​to generate a general speech recognition model for recognizing user speech in all languages.

[0178] The controller 250 can generate a third speech recognition model based on the target entity, the entity to be translated, and all sentence templates. The third speech recognition model is used to recognize the user's speech corresponding to the preset language and the initial language.

[0179] In some embodiments, the controller 250 can determine a preset language according to user needs. For example, the user controls the display device to bring up a language selection interface, thereby setting the language that the display device can recognize. The controller 250 can then determine the user-set language as the preset language. The controller 250 then translates the entity to be translated into the target entity in the preset language, thereby generating a speech recognition model. Figure 10 Schematic diagrams of the language selection interface in some embodiments are shown. For example... Figure 10 As shown, the language selection interface can have several language controls, each representing a language supported by the display device, including English, French, German, Spanish, and Chinese. When a user selects a language, the controller 250 can translate the entity to be translated into the corresponding target entity in that language, thereby obtaining training data and building a speech recognition model. After the user inputs speech in that language, the controller 250 can recognize the user's speech.

[0180] In some embodiments, the controller 250 can recognize user speech based on a speech recognition model.

[0181] The controller 250 can control the sound acquisition device to collect the user's voice input. When the user inputs voice into the display device, the controller 250 can recognize the user's voice based on the voice recognition model, obtain the control command corresponding to the user's voice, and execute the control command to achieve the corresponding function.

[0182] The controller 250 can first call a third-party speech recognition interface to convert the user's speech into speech text, and then use a speech recognition model to determine the meaning of the speech text in order to generate control commands and execute them.

[0183] In some embodiments, the controller 250 may also prompt the user after executing the control command. Figure 11 The illustrations show scenarios of voice interaction between users and display devices in some embodiments. For example... Figure 11 As shown, the user inputs the voice command "Wake me up at 12 o'clock tomorrow," and the controller 250 can recognize and execute the user's voice. The controller 250 can set an alarm for 12 o'clock tomorrow and prompt the user with a voice message, "An alarm for 12 o'clock tomorrow has been set for you."

[0184] In some embodiments, a user may use voice control to search for media resources on the display device. For example, the user may enter the voice command "Search for XXX movie season 3". After searching for the desired media resources, the controller 250 may display the search interface and simultaneously prompt the user with a voice message, "Videos about XXX have been recommended for you." Figure 12 The diagram illustrates a search interface displayed on a display device in some embodiments. When a user selects a target media asset, the controller 250 can control the display 260 to show the media asset details page, allowing the user to play and watch the media asset.

[0185] If the controller 250 does not find any relevant media assets, it can also display a preset prompt message to inform the user that no relevant media assets have been found, and provide a voice prompt to the user. Figure 13 A schematic diagram of the prompt information in some embodiments is shown.

[0186] This application also provides a speech recognition method, such as... Figure 14 As shown, the method includes:

[0187] Step 1401: Obtain the statement to be translated based on the statement templates pre-stored in the entity to be translated and the display device.

[0188] Step 1402: Translate the statement to be translated into the target statement in the preset language.

[0189] Step 1403: Obtain the target entity corresponding to the entity to be translated based on the target statement.

[0190] Step 1404: Generate a speech recognition model based on the target entity and pre-stored sentence templates.

[0191] Step 1405: In response to the user's voice collected by the sound collector, the user's voice is recognized based on the speech recognition model to obtain the control command corresponding to the user's voice and execute the control command.

[0192] In some embodiments, obtaining the statement to be translated based on the entity to be translated and a statement template pre-stored in the display device further includes:

[0193] Obtain the target entity type of the entity to be translated;

[0194] In the pre-defined database, statement templates are filtered based on the target entity type to obtain the target statement templates corresponding to the target entity type; several statement templates are pre-stored in the database.

[0195] Based on the entity to be translated and the target statement template, obtain the statement to be translated.

[0196] In some embodiments, the statement template includes a first fixed entity and a plurality of slots to be filled, wherein there is a matching relationship between the slots to be filled and the entity type. Filtering the statement template based on the target entity type further includes:

[0197] The statement templates in the database are filtered based on preset filtering conditions. The preset filtering conditions are: if there is at least one unfilled slot in the statement template that matches the target entity type, then the statement template is determined as the target statement template.

[0198] In some embodiments, obtaining the sentence to be translated based on the entity to be translated and the target sentence template further includes:

[0199] Determine the target slot to be filled for the entity to be translated in the target statement template, as well as the remaining slots to be filled outside the target slot; in the database, obtain the seed entity pairs corresponding to the remaining slots to be filled, the seed entity pairs include the first seed entity of the initial language and the second seed entity of the preset language corresponding to the entity to be translated; fill the entity to be translated into the target slot, and fill the first seed entity into the remaining slots to be filled to obtain the statement to be translated.

[0200] In some embodiments, obtaining the target entity corresponding to the entity to be translated based on the target statement further includes:

[0201] The target sentences are segmented into words to obtain segmentation results; each target sentence corresponds to a segmentation result; candidate entities corresponding to each target sentence are obtained based on the segmentation results; the candidate entities are filtered to obtain the target entities corresponding to the entities to be translated.

[0202] In some embodiments, obtaining candidate entities corresponding to each target statement based on the word segmentation results further includes:

[0203] In the database, the matching statement template corresponding to the target statement template is obtained. The matching statement template includes a second fixed entity and several slots to be filled. The second fixed entity is the entity corresponding to the first fixed entity in the preset language. In the word segmentation results, the remaining words other than the second fixed entity and the second seed entity are obtained, and the remaining words are determined as candidate entities corresponding to each target statement.

[0204] In some embodiments, filtering candidate entities to obtain the target entity corresponding to the entity to be translated further includes:

[0205] The process involves translating candidate entities into their corresponding initial language candidate entities; obtaining the edit distance between the initial language candidate entities and the entity to be translated; and obtaining the frequency of occurrence of each candidate entity among all candidate entities. Based on the edit distance, frequency of occurrence, and preset weight coefficients, the translation score of the candidate entities is calculated, and the candidate entity with the highest translation score is determined as the target entity corresponding to the entity to be translated.

[0206] In some embodiments, generating a speech recognition model based on a target entity and a pre-stored statement template further includes:

[0207] Determine the initial language of the entity to be translated; obtain the first statement template corresponding to the initial language and the second statement template corresponding to the preset language; generate a first speech recognition model based on the entity to be translated and the first statement template, and generate a second speech recognition model based on the target entity and the second statement template; the first speech recognition model is used to recognize the user speech corresponding to the initial language, and the second speech recognition model is used to recognize the user speech corresponding to the preset language.

[0208] In some embodiments, generating a speech recognition model based on a target entity and a pre-stored statement template further includes:

[0209] A third speech recognition model is generated based on the target entity, the entity to be translated, and pre-stored sentence templates. The third speech recognition model is used to recognize the user's speech corresponding to the preset language and the initial language.

[0210] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.

[0211] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0213] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of embodiments suitable for specific application considerations.

Claims

1. A display device, characterized in that, include: monitor; The sound acquisition device is configured to capture the user's voice input. The controller is configured as follows: Obtain the target entity type of the entity to be translated; In a pre-defined database, if at least one unfilled slot in a statement template matches a target entity type, then the statement template is determined to be the target statement template. The database pre-stores several statement templates, each including a first fixed entity and several unfilled slots, wherein there is a matching relationship between the unfilled slots and the entity type. Determine the target slot to be filled corresponding to the entity to be translated in the target statement template, and the remaining slots to be filled other than the target slot; In the database, the seed entity pair corresponding to the remaining slots to be filled is obtained. The seed entity pair includes a first seed entity of the initial language and a second seed entity of the preset language corresponding to the entity to be translated. The entity to be translated is filled into the target slot, and the first seed entity is filled into the remaining slots to obtain the sentence to be translated. Translate the statement to be translated into a target statement in a preset language; The target statement is segmented into words to obtain segmentation results; each target statement corresponds to one segmentation result. In the database, a matching statement template corresponding to the target statement template is obtained. The matching statement template includes a second fixed entity and the plurality of slots to be filled. The second fixed entity is the entity corresponding to the first fixed entity in the preset language. In the word segmentation results, the remaining words other than the second fixed entity and the second seed entity are obtained, and the remaining words are determined as candidate entities corresponding to each target statement; Translate the candidate entities into the corresponding initial language candidate entities under the initial language; Obtain the edit distance between the initial language candidate entities and the entity to be translated; And, obtain the number of times each candidate entity appears among all candidate entities; Based on the edit distance, the number of occurrences, and the preset weighting coefficient, the translation score of the candidate entity is calculated, and the candidate entity with the highest translation score is determined as the target entity corresponding to the entity to be translated. A speech recognition model is generated based on the target entity and the pre-stored sentence template; In response to the user's voice collected by the sound collector, the user's voice is recognized based on the speech recognition model to obtain the control command corresponding to the user's voice and execute the control command.

2. The display device according to claim 1, characterized in that, The controller is configured to generate a speech recognition model based on the target entity and the pre-stored statement template, and is also configured to: Determine the initial language of the entity to be translated; Obtain the first statement template corresponding to the initial language and the second statement template corresponding to the preset language; A first speech recognition model is generated based on the entity to be translated and the first sentence template, and a second speech recognition model is generated based on the target entity and the second sentence template; the first speech recognition model is used to recognize the user speech corresponding to the initial language, and the second speech recognition model is used to recognize the user speech corresponding to the preset language.

3. The display device according to claim 1, characterized in that, The controller is configured to generate a speech recognition model based on the target entity and the pre-stored statement template, and is also configured to: Determine the initial language of the entity to be translated; A third speech recognition model is generated based on the target entity, the entity to be translated, and the pre-stored sentence template. The third speech recognition model is used to recognize user speech corresponding to the preset language and the initial language.

4. A speech recognition method applied to a display device, characterized in that, The method includes: Obtain the target entity type of the entity to be translated; In a pre-defined database, if at least one unfilled slot in a statement template matches a target entity type, then the statement template is determined to be the target statement template. The database pre-stores several statement templates, each including a first fixed entity and several unfilled slots, wherein there is a matching relationship between the unfilled slots and the entity type. Determine the target slot to be filled corresponding to the entity to be translated in the target statement template, and the remaining slots to be filled other than the target slot; In the database, the seed entity pair corresponding to the remaining slots to be filled is obtained. The seed entity pair includes a first seed entity of the initial language and a second seed entity of the preset language corresponding to the entity to be translated. The entity to be translated is filled into the target slot, and the first seed entity is filled into the remaining slots to obtain the sentence to be translated. Translate the statement to be translated into a target statement in a preset language; The target statement is segmented into words to obtain segmentation results; each target statement corresponds to one segmentation result. In the database, a matching statement template corresponding to the target statement template is obtained. The matching statement template includes a second fixed entity and the plurality of slots to be filled. The second fixed entity is the entity corresponding to the first fixed entity in the preset language. In the word segmentation results, the remaining words other than the second fixed entity and the second seed entity are obtained, and the remaining words are determined as candidate entities corresponding to each target statement; Translate the candidate entities into the corresponding initial language candidate entities under the initial language; Obtain the edit distance between the initial language candidate entities and the entity to be translated; and obtain the number of occurrences of each candidate entity among all candidate entities; Based on the edit distance, the number of occurrences, and the preset weighting coefficient, the translation score of the candidate entity is calculated, and the candidate entity with the highest translation score is determined as the target entity corresponding to the entity to be translated. A speech recognition model is generated based on the target entity and the pre-stored sentence template; In response to the user's voice collected by the sound acquisition device, the user's voice is recognized based on the speech recognition model to obtain the control command corresponding to the user's voice and execute the control command.