Display device and subtitle display method
By collecting audio signals in the display device for language recognition and analyzing the stability of the use of historical subtitle language, the problem of low reliability of subtitle language in complex language environments is solved, and the reliability and user experience of subtitle display are improved.
Patent Information
- Application Number
- CN202510570950.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional display devices have low reliability in complex language environments, making it difficult to adapt to scenarios with large flow of people, resulting in users not being able to understand subtitles.
Audio signals are collected through microphones for language recognition, combined with the confidence of environmental languages and historical subtitle language recording, analyze the stability of historical subtitle language usage, and filter out the most stable subtitle language.
Improve the reliability of subtitle language, ensure the display effect of subtitles, and improve the user experience.
Smart Images

Figure CN120583285A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a display device and a subtitle display method. Background Art
[0002] Most display devices currently have subtitles to assist users in watching videos. However, because traditional display devices typically use fixed subtitle languages, users may not be able to understand the subtitles in public places with high traffic, such as shopping malls and hotels.
[0003] Some solutions have been proposed to address this issue, such as adjusting the subtitle language based on the surrounding language. However, complex language environments often contain more noise interference, which can easily affect language recognition and make subtitles less reliable. Summary of the Invention
[0004] The present application provides a display device and a subtitle display method to solve the current problem of low reliability of subtitle language.
[0005] In a first aspect, some embodiments provide a display device, including:
[0006] a display configured to display a user interface;
[0007] a microphone configured to collect audio signals;
[0008] The controller is configured as:
[0009] When the display device is in an adaptive subtitle display mode, performing language recognition based on the audio signal to obtain an ambient language and a confidence level of the ambient language;
[0010] When the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, obtaining the language usage of each historical subtitle language in the historical subtitle language record;
[0011] Analyzing the usage stability of each of the historical subtitle languages based on the usage of each of the languages;
[0012] From the historical subtitle languages, the language whose usage stability satisfies the usage stability condition is selected as the subtitle language of the display device.
[0013] Technical effect: In this solution, when the display device turns on the adaptive subtitle display mode, it will first perform language recognition based on the audio signal of the environment in which the display device is located to obtain the environment language and the confidence of the environment language. When the confidence of the environment language does not meet the confidence condition, it means that the currently identified environment language has a low credibility. At this time, it can be further determined whether there is a historical subtitle language record for the display device in the historical period. If so, the language usage of each historical subtitle language in the historical subtitle language record is obtained to analyze the usage stability of each historical subtitle language, and the language that meets the stability condition is used as the current subtitle language of the display device. It can be understood that the language with the most stable performance can directly reflect the recent subtitle language display requirements of the display device in the current scenario. In this way, when the real-time language recognition confidence is insufficient, this solution can effectively improve the reliability of the subtitle language by analyzing the usage stability of the historical subtitle language and selecting the most stable language as the subtitle language, thereby ensuring the subtitle display effect and improving the user experience.
[0014] In a second aspect, some embodiments further provide a subtitle display method, which is applied to the display device provided in the first aspect, wherein the display device includes: a display, a microphone, and a controller, and the method includes:
[0015] When the display device is in an adaptive subtitle display mode, performing language recognition based on an audio signal collected by a microphone of the display device to obtain an ambient language and a confidence level of the ambient language;
[0016] When the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, obtaining the language usage of each historical subtitle language in the historical subtitle language record;
[0017] Analyzing the usage stability of each of the historical subtitle languages based on the usage of each of the languages;
[0018] From the historical subtitle languages, the language whose usage stability satisfies the usage stability condition is selected as the subtitle language of the display device.
[0019] Technical effect: In this solution, when the display device turns on the adaptive subtitle display mode, it will first perform language recognition based on the audio signal of the environment in which the display device is located to obtain the environment language and the confidence of the environment language. When the confidence of the environment language does not meet the confidence condition, it means that the currently identified environment language has a low credibility. At this time, it can be further determined whether there is a historical subtitle language record for the display device in the historical period. If so, the language usage of each historical subtitle language in the historical subtitle language record is obtained to analyze the usage stability of each historical subtitle language, and the language that meets the stability condition is used as the current subtitle language of the display device. It can be understood that the language with the most stable performance can directly reflect the recent subtitle language display requirements of the display device in the current scenario. In this way, when the real-time language recognition confidence is insufficient, this solution can effectively improve the reliability of the subtitle language by analyzing the usage stability of the historical subtitle language and selecting the most stable language as the subtitle language, thereby ensuring the subtitle display effect and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;
[0022] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0023] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;
[0024] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0025] Figure 5 A schematic diagram of the interaction between various modules in the subtitle language matching process provided in some embodiments of the present application;
[0026] Figure 6 A flowchart of a subtitle display method provided in some embodiments of the present application;
[0027] Figure 7 A schematic diagram of the subtitle language matching process provided in some embodiments of the present application;
[0028] Figure 8 Another schematic diagram of a process for subtitle language matching provided in some embodiments of the present application;
[0029] Figure 9 A schematic diagram of a subtitle language display process provided in a specific embodiment of the present application;
[0030] Figure 10 A flowchart of real-time subtitle language update provided in a specific embodiment of the present application. DETAILED DESCRIPTION
[0031] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0032] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0033] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0034] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0035] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.
[0036] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0037] Figure 1This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0038] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling connection and communication via a network communication protocol, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.
[0039] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0040] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.
[0041] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0042] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0043] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.
[0044] In some embodiments, the display 260 includes a display component for displaying images and a driver component for driving image display. The display 260 is configured to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces.
[0045] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.
[0046] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.
[0047] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200.
[0048] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0049] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0050] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.
[0051] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.
[0052] Figure 3 Some embodiments of this application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0053] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .
[0054] In some embodiments, the control device 100 may be a smart device. For example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0055] In some embodiments, as Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0056] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and external and internal data processing functions.
[0057] Under the control of the controller 110, the communication interface 130 communicates control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, or other near field communication modules.
[0058] The user input / output interface 140 includes at least one of a microphone 141 , a touch panel 142 , a sensor 143 , a button 144 and other input interfaces.
[0059] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, to encode user input commands via the WiFi protocol, Bluetooth protocol, or NFC protocol and transmit them to the display device 200.
[0060] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0061] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0062] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system may provide a user interface (to control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0063] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0064] The operating system can be divided into different modules or layers according to the functions implemented.
[0065] For example, Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the application layer (abbreviated as "application layer"), the application framework layer (abbreviated as "framework layer"), the system library layer and the kernel layer.
[0066] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on these applications. The application layer can host at least one application, which can include built-in window programs, system settings programs, clock programs, and other applications provided by the operating system, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0067] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.
[0068] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to the system location service; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0069] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling application exit, opening, and back. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling changes in display windows, such as shrinking, shaking, or distorting the display window.
[0070] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0071] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4As shown, the kernel layer can be configured with hardware drivers, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0072] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0073] Most display devices currently offer subtitles to assist users in watching videos. However, traditional display devices typically use a fixed subtitle language mode. This means the subtitles are in a single, fixed language and require manual user switching. This can lead to users having difficulty understanding the subtitles in complex and changing language environments (e.g., in shopping malls, hotels, and other public places).
[0074] Some solutions have been proposed to address this issue, such as adjusting the subtitle language based on the surrounding language. However, complex language environments often contain more noise interference, which can easily affect the accuracy of language recognition, resulting in unreliable subtitle language.
[0075] In order to solve the above problem, in some embodiments, a display device is provided, the display device including:
[0076] The display is configured to display a user interface.
[0077] Among them, the user interface is the interface displayed on the display, and the user interface may include a subtitle display area for displaying subtitles in real time, and the subtitle display area may be fixed at the bottom of the screen or other locations on the screen. In addition to the subtitle display area, the user interface may also include a "subtitle display mode" control, which may be set in the upper right corner of the screen or other locations on the screen. The user can trigger the "subtitle display mode" control to set the subtitle display mode of the display device, such as a fixed subtitle display mode or an adaptive subtitle display mode. When the user selects the fixed subtitle display mode, a variety of candidate languages may be further displayed. After the user selects a language, it may be used as a subtitle language. When the user selects the adaptive subtitle display mode, it is necessary to determine the appropriate subtitle language based on the real-time environment language, which will be explained in detail in subsequent embodiments.
[0078] In some embodiments, the subtitle display mode can also be associated with the scene in which the display device is located. Specifically, the user interface can also display a "scene mode" control, and the user can trigger the "scene mode" control to select the scene mode of the display device, such as shopping mall mode, home mode, hotel mode, etc. The subtitle display mode can be bound to the scene mode. For example, when the user selects shopping mall mode or hotel mode, the display device can automatically set the subtitle display mode to an adaptive subtitle display mode; when the user selects home mode, the display device can automatically set the subtitle display mode to a fixed subtitle display mode, and the fixed display subtitle language can be a language selected by the user. Of course, in actual applications, other subtitle display modes and scene modes can be configured according to actual needs, and this embodiment does not limit this.
[0079] The microphone is configured to collect audio signals, wherein the audio signals may refer to audio signals of the environment in which the display device is located, for subsequent environmental language recognition.
[0080] The controller is configured to: when the display device turns on the adaptive subtitle display mode, perform language recognition based on the audio signal to obtain the environment language and the confidence level of the environment language.
[0081] Among them, the adaptive subtitle display mode is an intelligent subtitle display mode. In this mode, the subtitle language is no longer fixed, but can be automatically matched based on the real-time environment. Language recognition can refer to the process of identifying the language type of the input audio signal. The environmental language can refer to the dominant language type in the environment where the current display device is located, such as the language used most frequently or for the longest time in the environment. The environmental language type can include one or more language types. The confidence level can be used to indicate the credibility of the environmental language, usually quantified as a probability value or a score range (such as 0-100%).
[0082] For example, a user can activate the adaptive subtitle display mode of the display device through any method, such as a remote control, voice command, or touch control. At this time, the controller begins to receive the audio signal collected and transmitted by the microphone in real time, and performs language recognition based on the audio signal to obtain the ambient language of the environment in which the display device is located and the confidence score of the ambient language.
[0083] In some embodiments, the controller can capture audio data from an audio signal by calling the AudioRecord class, thereby performing language recognition based on the audio data. The AudioRecord class supports real-time capture of raw audio data from a microphone, reducing latency and preventing audio data loss, making it more suitable for real-time language recognition scenarios.
[0084] In some embodiments, after acquiring the audio data, the controller may further pre-process the audio data, such as noise reduction and normalization, before performing language recognition to improve the audio data quality and ensure the accuracy of language recognition. Furthermore, given that the microphone may capture private information such as the user's contact information and address, this private information may also be encrypted or desensitized to protect user privacy.
[0085] In some embodiments, the controller can use the display device's intelligent voice assistant to implement language recognition. It is understandable that mainstream intelligent voice assistants generally have integrated modules such as voice recognition and natural language understanding, thus having the ability to recognize languages. Therefore, the controller can call the voice assistant's audio stream interface to send audio data to the voice assistant. The voice assistant performs language recognition and feeds the recognition results back to the controller. By directly reusing the language recognition capabilities of the intelligent voice assistant, development costs can be reduced.
[0086] In other embodiments, given that intelligent voice assistants may not support recognition of certain specific languages (such as minority languages and dialects), a self-trained language recognition model can be used to implement language recognition. A language recognition model can refer to a mathematical model used to identify language types, and can be, but is not limited to, at least one of a convolutional neural network, a recurrent neural network, or a Transformer. Specifically, the controller can pre-collect a large amount of audio data, including samples of the target language to be recognized. It then uses feature extraction techniques such as MFCC (Mel Frequency Cepstral Coefficient) to extract features from the audio. Finally, a deep learning framework (TensorFlow or PyTorch) is used to train the language recognition model. Compared to intelligent voice assistants, self-trained language recognition models are not limited to widely used mainstream languages, offer greater adaptability to specific scenarios, and possess efficient inference speed, enabling timely language recognition.
[0087] In other embodiments, the intelligent voice assistant and the language recognition model can also be used in conjunction. Specifically, the intelligent voice assistant can send the language information generated during the interaction with the user in the past period of time to the language recognition model for auxiliary recognition. For example, the intelligent language assistant feedback that the language that the user mainly interacted with in the past 10 minutes was B, then when the language recognition model identifies the environmental language A and the environmental language B, the proportion of environmental language B will be greater, the confidence will be higher, and it will be given priority as the subtitle language. In this way, with the assistance of the intelligent voice assistant, the context perception ability of the language recognition model can be improved, thereby improving the recognition accuracy of the language recognition model.
[0088] In other embodiments, the controller can also update the language usage of each historical subtitle language in the historical subtitle language record in real time based on the language information generated by the intelligent language assistant during past interactions with the user, such as the number of interactions and the duration of interactions for each historical subtitle language. This improves the reliability and accuracy of the historical subtitle language record, and allows the optimal subtitle language to be selected from the historical subtitle language record when the confidence level of the currently recognized environment language is low.
[0089] When the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, the language usage of each historical subtitle language in the historical subtitle language record is obtained.
[0090] The confidence condition can refer to a pre-set threshold or standard used to assess the reliability of the recognized ambient language. Failure to meet the confidence condition means the confidence level of the current ambient language is less than or equal to the confidence threshold, indicating that the real-time language recognition results are insufficiently reliable and requiring alternative processing strategies. Conversely, the ambient language is considered reliable. The confidence condition can be customized based on application scenario requirements, such as 80% or 90%. The historical period is a preset time range, such as the past 10 minutes or the past hour, that limits the time window for analyzing historical subtitle language data. The specific time window can be determined based on actual needs. Historical subtitle language records can refer to the subtitle languages used by a display device during a historical period, along with their contextual information (such as usage timestamps, language types, usage scenario tags, and voice assistant interaction tags). Historical subtitle languages are the set of subtitle languages used by a display device during a historical period. Language usage can refer to metrics such as the frequency of use, duration (average or total duration of each use), and usage time of each historical subtitle language during the historical period. For example, English appeared 12 times, Chinese appeared 8 times, and Japanese appeared 3 times in the past hour.
[0091] In some embodiments, if the historical period is in days, the usage time distribution of the historical subtitle language can be further obtained, such as different hours of the day, or specific time periods (such as morning, afternoon, and evening). This indicator can be used to measure the concentration or regularity of the use of the historical subtitle language in a specific time period.
[0092] For example, when the confidence level of the identified environmental language does not meet the preset confidence conditions, the current environmental language is considered unreliable, and the controller may adopt other strategies, namely historical data analysis. Before performing historical subtitle language analysis, the controller needs to first determine whether there is a historical subtitle language record for the display device, or in other words, whether there is a historical subtitle language record for the display device within a historical period. In the case of determining that there is a historical subtitle language record for the display device, the frequency of use, duration, usage time and other indicators of each historical subtitle language in the historical subtitle language record within the historical period can be obtained for each historical subtitle language in the historical subtitle language record, so as to analyze the stability of use of each historical subtitle language.
[0093] If the display device has no historical subtitle language records, this means that the real-time language recognition is unreliable and there is no historical data for further analysis. In this case, the controller can implement further strategies to ensure proper subtitle display, such as using the display device's default language as the subtitle language, or using the local language of the display device's location as the subtitle language.
[0094] Based on the usage of each language, analyze the usage stability of each historical subtitle language.
[0095] Usage stability can be understood as a measure of the continued usability of a historical subtitle language over a historical period. This can be expressed as a stability score, a stability level (e.g., high stability, medium stability, low stability), or a percentage (e.g., 90% stability), though this embodiment does not impose any specific limitations.
[0096] For example, after obtaining the language usage of each historical subtitle language, the controller can analyze the continued use of each historical subtitle language over the historical period based on the language usage. For example, the higher the frequency of use of a historical subtitle language over the historical period, the higher its usage stability, and vice versa, the longer the usage duration, the higher its stability.
[0097] In some embodiments, the controller can combine multiple language usage indicators to comprehensively analyze the usage stability of historical subtitle languages. For example, a predefined formula for calculating the usage stability of historical subtitle languages is as follows: Stability (L) = a*frequency of use + b*duration, where a and b represent the weights of the frequency of use and duration, respectively, and Stability represents the usage stability of the historical subtitle language L. By collaboratively analyzing the usage stability of historical subtitle languages using multiple indicators, it is possible to avoid the problem that a single indicator is easily affected by short-term fluctuations or outliers, resulting in inaccurate analysis results. This improves the accuracy of the historical subtitle language stability analysis and facilitates matching of more accurate subtitle languages.
[0098] In other embodiments, if development costs are sufficient, a machine learning model can be introduced to analyze the usage stability of historical subtitle languages. For example, a random forest model can be trained and input into the model using metrics such as the frequency, duration, and duration of use of each historical subtitle language over a historical period. The model then performs a stability analysis and outputs the usage stability of each historical subtitle language. This allows for efficient and accurate analysis of the usage stability of historical subtitle languages.
[0099] From various historical subtitle languages, a language whose usage stability satisfies a usage stability condition is selected as the subtitle language of the display device.
[0100] Among them, the stability condition can be a preset threshold or logical rule, which is used to determine whether the historical subtitle language has sufficient stability. For example, setting the stability score threshold to 90 points means that the stability of the historical subtitle language needs to be greater than or equal to 90 points to be considered to meet the conditions. Or when the stability level is high, it can also be considered to meet the conditions. It should be noted that the stability condition can be flexibly set according to different scenarios. For example, the stability score threshold can be set to 90 points in scenario A and 80 points in scenario B. This is conducive to matching a subtitle language that is more suitable for the actual scenario.
[0101] For example, after analyzing the stability of each historical subtitle language, the controller can select historical subtitle languages that meet preset stability requirements. To ensure the accuracy of the final subtitle language displayed, the controller can perform further verification on the initially matched historical subtitle languages, such as uniqueness verification and consistency verification with the environment language. After the historical subtitle language passes these verifications, it can be used as the final subtitle language.
[0102] In this embodiment, when the display device turns on the adaptive subtitle display mode, it will first perform language recognition based on the audio signal of the environment in which the display device is located to obtain the environment language and the confidence of the environment language. When the confidence of the environment language does not meet the confidence condition, it means that the currently identified environment language has a low credibility. At this time, it can be further determined whether the display device has a historical subtitle language record in the historical period. If so, the language usage of each historical subtitle language in the historical subtitle language record is obtained to analyze the usage stability of each historical subtitle language, and the language that meets the stability condition is used as the current subtitle language of the display device. It can be understood that the language with the most stable performance can directly reflect the recent subtitle language display requirements of the display device in the current scenario. In this way, when the real-time language recognition confidence is insufficient, this solution can effectively improve the reliability of the subtitle language by analyzing the usage stability of the historical subtitle language and selecting the most stable language as the subtitle language, thereby ensuring the subtitle display effect and improving the user experience.
[0103] In actual applications, there may be multiple historical subtitle languages that meet the stability conditions. For example, the stability scores of historical subtitle language A and historical subtitle language B are both over 90 points, indicating that historical subtitle language A and historical subtitle language B have performed relatively stably in the historical period. In this case, the controller can further screen multiple historical subtitle languages that meet the stability conditions to select the optimal subtitle language.
[0104] Therefore, in some embodiments, when the controller is executing the process of screening out a language whose usage stability meets the usage stability condition from various historical subtitle languages as the subtitle language of the display device, the controller is further configured to: screen out candidate subtitle languages whose usage stability meets the usage stability condition from various historical subtitle languages; if the candidate subtitle language is unique, use the candidate subtitle language as the subtitle language of the display device; if there are multiple candidate subtitle languages, score each candidate subtitle language separately to obtain a scoring result for each candidate subtitle language; and screen out a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language as the subtitle language of the display device.
[0105] Among them, candidate subtitle languages refer to historical subtitle languages that have been screened from historical subtitle language records and that initially meet the stability conditions. A unique candidate subtitle language means that only one language meets the preset stability conditions and becomes the only option for the candidate subtitle language, so the final subtitle language can be locked in without further analysis. The existence of multiple candidate subtitle languages means that in the historical subtitle language records, two or more languages simultaneously meet the preset stability conditions and become a set of candidate subtitle languages that require further refined screening. In this case, a scoring mechanism is required. Scoring here refers to a multi-dimensional quantitative evaluation of multiple candidate subtitle languages to generate a basis for priority sorting. The scoring results can include the score and priority of each candidate subtitle language. The higher the score, the higher the priority of being selected as the subtitle language. The first scoring requirement can refer to the scoring rule or sorting rule for screening the final subtitle language, such as the highest score, first priority, etc. The selected subtitle language is the final subtitle language.
[0106] Exemplarily, the controller first selects candidate subtitle languages that preliminarily meet the usage stability condition from each historical subtitle language, and performs a uniqueness check on the candidate subtitle language, that is, verifies whether the candidate subtitle language is unique. If it is unique, it can be directly locked as the final subtitle language. If it is not unique, that is, the candidate subtitle language includes at least two languages, each candidate subtitle language can be scored, so as to select the one that meets the first scoring requirement, such as the highest score or the highest priority, as the final subtitle language.
[0107] In some embodiments, since all candidate subtitle languages meet the usage stability condition at this point, other evaluation indicators can be introduced for scoring, such as the specific usage time of the candidate subtitle language. It can be understood that the more recent the candidate subtitle language, that is, the closer it is to the current time, the more it can reflect the user's subtitle language usage preference in the current environment. Therefore, the controller can define a time decay weight to dynamically adjust the contribution of candidate subtitle languages in different time periods in the scoring calculation.
[0108] For example, the scoring formula is predefined as: Score1(M)=stability score*c+usage period*d, where c and d represent the indicator weights of the stability score and usage period respectively, and Score1 represents the score of the candidate subtitle language M. Taking the historical period of the past hour as an example, the candidate subtitle language with a usage period within the past half hour can be given a higher weight, while the candidate subtitle language with a usage period within the first half hour can be given a lower weight. In this way, a subtitle language that is more in line with the actual scenario can be matched. Of course, the indicator of language usage time is not limited to use in this embodiment, and can also be used in language stability analysis, which can further improve the accuracy and reliability of language usage stability analysis, and can also avoid the situation of multiple candidate subtitle languages to a certain extent. If there are still multiple candidate subtitle languages, other mechanisms such as consistency verification with the environment language can be selected for further matching.
[0109] In this embodiment, in view of the fact that there may be multiple historical subtitle languages that meet the usage stability condition, a multi-dimensional scoring mechanism is introduced to match the optimal subtitle language from multiple historical subtitle languages, thereby ensuring that the subtitle language is displayed accurately.
[0110] In addition to the situation where there are multiple historical subtitle languages, the environmental languages identified in real time may also be single-category languages or multiple-category languages. Therefore, in order to further ensure that the optimal subtitle language is matched, in some embodiments, if the candidate subtitle language is unique, the controller is further configured to use the candidate subtitle language as the subtitle language of the display device: if the candidate subtitle language is unique, if the environmental language is a single-category language, the candidate subtitle language is used as the subtitle language of the display device; if the environmental language includes multiple target languages, the candidate subtitle language is matched with each target language separately; when the candidate subtitle language matches any target language, the candidate subtitle language is used as the subtitle language of the display device.
[0111] Among them, for the case where the candidate subtitle language is unique, if the current real-time identified environmental language is a single-category language, that is, the current environmental language is single and the confidence level is not high, then the unique candidate subtitle language can be directly used as the subtitle language of the display device. If the current real-time identified environmental language is a multi-category language, that is, the environmental language contains multiple target languages, such as Chinese, English and French. It should be noted that in this case, the confidence level of the environmental language does not meet the confidence condition, which means that the confidence level of each target language does not meet the confidence condition, that is to say, each target language currently identified is unreliable, and a better subtitle language needs to be matched from the historical subtitle language records. In order to avoid the subtitle language being out of touch with the real-time environment, in this embodiment, the candidate subtitle language will also be selected for consistency matching with each target language. Consistency matching refers to comparing the candidate subtitle language with the target language one by one to determine whether there is a language-consistent item. If the language is consistent, the candidate subtitle language can be used as the subtitle language of the display device. The language consistency may be the consistency of the language nouns or the consistency of the language codes. This embodiment does not impose any limitation on the basis for determining the language consistency.
[0112] It should be noted that the uniqueness of the candidate subtitle language has passed the historical stability screening condition, indicating that the language has dominated in long-term use and has higher scene adaptability. At this point, it can be considered that the credibility of the candidate subtitle language is higher than the identified single environment language. Regardless of whether the candidate subtitle language is consistent with the language, the candidate subtitle language can be directly used as the final subtitle language. Therefore, in this case, there is no need to compare the candidate subtitle language with the single environment language for consistency, which can reduce computing resources to a certain extent. Of course, in order to further ensure that the subtitle language is accurate and reliable, you can also choose to re-score the candidate subtitle language and the single environment language, and select the one with the highest score as the final subtitle language.
[0113] In a multilingual environment, where the display device is operating in a variety of languages, assuming the language range includes both Chinese and English, to avoid the final subtitle language being completely unrelated to the real-time environment, the candidate subtitle languages can be compared against the target language one by one. If this comparison reveals that the candidate subtitle language is consistent with English, the display device's subtitle language will be set to English.
[0114] In this embodiment, analyses are performed for both single-language and mixed-language scenarios. In the single-language scenario, when the confidence level of the identified language is low, the only historical subtitle language that meets the stability requirements is deemed to have higher credibility and scenario adaptability, and therefore can be directly used as the final subtitle language. In the mixed-language scenario, when the language confidence level is low, to prevent the subtitle language from being out of sync with the real-time environment, the only historical subtitle language that meets the stability requirements is matched with multiple environment languages for consistency, and the language with the most consistent match is prioritized as the final subtitle language to improve the adaptability of the subtitle language to the environment.
[0115] Of course, there may be situations where the candidate subtitle language is not the same as any of the target languages. Therefore, in some embodiments, the controller is further configured to: when the candidate subtitle language is not the same as any of the target languages, score the candidate subtitle language and each of the target languages separately to obtain a score result for each of the candidate subtitle language and each of the target languages; and use the language whose score meets the second scoring requirement as the subtitle language for the display device.
[0116] Among them, the scoring process in this embodiment is similar to the scoring process between the aforementioned multiple candidate subtitle languages, both of which belong to multi-index weighted analysis. However, in this embodiment, the language confidence indicator can be introduced. For example, the scoring formula in this embodiment can be expressed as: Score2(N)=stability score*e+language confidence*f, where e and f represent the indicator weights of the stability score and language confidence respectively, and Score2 represents the score of language N, which is any one of the candidate subtitle languages and the target languages. The indicator weights can be set according to actual needs. The specific weight values and the specific indicators and number of indicators involved in the scoring are not limited in this embodiment. The second scoring requirement can refer to the scoring rules or sorting rules for screening the final subtitle language, such as the highest score, the first priority, etc. It can be understood that the first scoring requirement and the second scoring requirement can be the same, such as both need to be the highest score, or they can be different, such as the first scoring requirement can be the highest score, and the second scoring requirement can be not less than the scoring threshold. The specific requirements can be determined according to actual needs.
[0117] For example, if the only candidate subtitle language does not match each target language, directly selecting the candidate subtitle language may result in the final subtitle language being completely irrelevant to the current environment. Therefore, this embodiment selects to re-score the candidate subtitles and each target language, and then selects the language that meets the second scoring requirement, such as the highest score, as the final subtitle language for the display device.
[0118] This embodiment uses a cross-dimensional scoring system to identify historical subtitle languages and real-time language recognition results, finding the optimal balance between historical stability and real-time performance. Furthermore, this cross-dimensional scoring system improves the flexibility and adaptability of subtitle language matching, especially in multilingual environments. Therefore, this embodiment ensures optimal subtitle language matching.
[0119] The above embodiments describe subtitle language matching methods for a single-language scenario and a multi-language scenario, respectively, when there is only one candidate subtitle language. However, when there are multiple candidate subtitle languages, corresponding subtitle language matching methods can also be configured for a single-language scenario and a multi-language scenario.
[0120] Therefore, in some embodiments, in the process of selecting a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language as the subtitle language of the display device, the controller is further configured to: select a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language; if the environment language includes multiple target languages, match the selected subtitle language with each target language respectively; when the selected subtitle language matches any target language, use the selected subtitle language as the subtitle language of the display device.
[0121] Among them, it is mentioned in the aforementioned embodiment that when there are multiple candidate subtitle languages, each candidate subtitle language can be scored to screen out the selected subtitle language that meets the first scoring requirement, such as the highest score. It can be understood that the number of selected subtitle languages at this time should be one, which is similar to the subtitle language matching method for the single-category environmental language scenario and the multi-category environmental language scenario under the aforementioned only candidate subtitle language. In this embodiment, if the environmental language is a single-category language, the selected subtitle language will be used as the subtitle language of the display device. If the environmental language is a multi-category language, the selected subtitle language will be matched with each target language for consistency. When the selected subtitle language matches any target language, the selected subtitle language can be used as the subtitle language of the display device. If the selected subtitle language does not match any target language, the selected subtitle language and each target language will be scored again, so that the one with the highest score will be used as the subtitle language of the display device.
[0122] In this embodiment, in scenarios where there are multiple candidate subtitle languages that meet the usage stability condition, optimal subtitle language matching is also performed for single-language scenarios and multi-language mixed scenarios, to ensure that the most appropriate subtitle language can be matched in various complex scenarios.
[0123] In actual applications, for scenes with frequent flow of people, such as shopping malls, the subtitle language used at the previous moment may not be applicable at the next moment. Therefore, in some embodiments, the controller is further configured to: obtain the subtitle display requirements of the scene in which the display device is located; when the subtitle display requirements meet the subtitle update conditions, perform language recognition again based on the current audio signal collected by the microphone to obtain the current environment language and the confidence of the current environment language; when the confidence of the current environment language meets the confidence conditions, update the subtitle language of the display device to the current environment language.
[0124] The term "subtitle display requirements" can be understood as information related to subtitle display requirements in the current environment, such as the subtitle display scenario. Alternatively, it can refer to requirements for font styles, such as font size, font color, and display position on the user interface. It can also refer to requirements for simultaneous display of multiple languages, such as bilingual display. It is understood that different scenarios may have different subtitle display requirements. Subtitle update conditions can be prerequisites for triggering subtitle updates. Subtitle updates may include, but are not limited to, real-time updates to the subtitle language, font size, font color, subtitle display position, and multilingual subtitle updates. For example, if the subtitle display requirement is to display subtitles in shopping mall mode, it is understandable that the high turnover of people in a shopping mall environment may require real-time updates to the subtitle language. In this case, real-time subtitle language updates may be triggered. If the subtitle display requirement is for bilingual display, a multilingual subtitle update may be triggered. Of course, users can also trigger subtitle updates manually or through voice commands, and the display device can update subtitles based on the user's instructions. The current audio signal refers to the audio signal currently being captured by the microphone in real time. The current environment language refers to the currently recognized environment language in real time.
[0125] For example, in order to make the display of the subtitle language more compatible with the actual scene, the controller will further obtain the subtitle display requirements of the scene in which the display device is located. It can be understood that the timing of obtaining the subtitle display requirements is not restricted. It can be obtained synchronously when the display device turns on the adaptive subtitle display mode, or it can be obtained after completing a round of subtitle language updates. The controller will determine whether the subtitle display requirement triggers the real-time update condition of the subtitles. If triggered, the controller will loop the language recognition, that is, obtain the audio signal collected by the microphone in real time again, and then perform language recognition again. When the confidence of the identified environmental language meets the confidence condition, it will be displayed as the latest subtitle language. If the confidence condition is not met, it is necessary to screen the most suitable subtitle language in the historical subtitle language record. If the subtitle display requirement does not trigger the real-time update condition of the subtitles, the process can be ended after completing a round of subtitle language updates.
[0126] In some embodiments, the interval duration and stop conditions of the real-time update of the subtitle language can be determined according to actual needs. For example, the subtitle language is updated every 1 minute until the display device is turned off or the user manually or voice instructs to turn off the real-time update function of the subtitles and stops updating.
[0127] In this embodiment, when the subtitle display demand in a real-time scenario triggers a subtitle update condition, the display device can cyclically identify the ambient language to achieve real-time updates of the subtitle language. This allows the subtitle language to respond to environmental changes in real time, thereby improving the adaptability of the subtitle language to the real-time scenario and making the subtitle language display more intelligent.
[0128] In some embodiments, the controller is further configured to: if the confidence level does not meet the confidence condition and there is no historical subtitle language record for the display device within the historical period, obtain the default language of the display device; and use the default language as the subtitle language of the display device.
[0129] If the display device has no historical subtitle language records within a certain period of time, this means the device may be in its initial use and no relevant subtitle language usage records have been generated, or the device's historical subtitle language records have been lost. In this case, to avoid interruption of subtitle display service, the display device's default language can be used as the subtitle language. The default language can be understood as the system language preset at the factory for the display device, or a user-configured language.
[0130] Generally speaking, when a display device is first started or restored to factory settings, it will not have any subtitle language history. In this case, a default language fallback mechanism can be triggered, using the display device's default language as the subtitle language. This ensures basic usability of the subtitle feature while preventing issues with no subtitles on screen due to a lack of basis for decision making.
[0131] In some embodiments, the default language may be automatically set based on the geographic location of the display device in addition to being factory preset or user configured. For example, if the device address of the display device is in France, the default language may be automatically set to French.
[0132] In other embodiments, the default language fallback mechanism can also incorporate a priority fallback strategy. Specifically, multiple levels of default languages are preset and fall back according to priority. If the primary default language fails, a secondary priority language is used. For example, if the primary default language is Chinese and the Chinese font file is damaged, a secondary priority language such as English can be used.
[0133] In this embodiment, when the confidence of the real-time environment language does not meet the confidence condition and the display device does not have any historical subtitle language records in the historical period, the default language of the display device is used as the subtitle language. This can avoid interruption of subtitle display, thereby ensuring the basic availability of the subtitle function and improving user experience.
[0134] In some embodiments, the controller is further configured to perform volume analysis based on the audio signal to obtain the ambient volume.
[0135] Volume analysis involves processing audio signals in the time or frequency domain to calculate their sound pressure or energy intensity. Ambient volume refers to the ambient sound intensity measured in decibels.
[0136] During the process of performing language recognition based on the audio signal, the controller is further configured to perform language recognition based on the audio signal when the ambient volume is greater than the media volume of the display device.
[0137] The media volume of the display device refers to the volume of the audio content (such as advertisements, videos, etc.) currently being played by the display device, or the preset volume value. When the ambient volume is greater than the media volume of the display device, for example, the ambient volume is 70dB and the media volume is 60dB, it means that the surrounding environment is relatively noisy at this time, and the ambient language recognition function can be activated. That is, the display device begins to identify the ambient language through an intelligent language assistant or a self-trained language recognition model, and dynamically updates the subtitle language based on the identified ambient language. When the ambient volume is less than or equal to the media volume of the display device, it means that the surrounding environment is not noisy at this time, and the user can clearly hear the content currently played by the display device, and the ambient language recognition function can be disabled.
[0138] In some embodiments, whether to compare the ambient volume with the media volume can be further determined based on the subtitle display requirements of the display device. For example, if the display device is used in a shopping mall, the display device can be in shopping mall mode. Due to the high traffic and noisy nature of shopping malls, there is no need to compare the ambient volume with the media volume. Instead, the language recognition loop can be directly executed to update the subtitle language in real time. This can effectively reduce computing resources and energy consumption.
[0139] In this embodiment, by comparing the ambient volume with the media volume, it can be used as a basis for judging whether the environment is noisy, so as to start the language recognition function in a targeted manner, realize dynamic updating of the subtitle language, and reduce the energy consumption of the language recognition function to a certain extent.
[0140] In some embodiments, the controller is further configured to: if the confidence level of the environment language satisfies a confidence condition, use the environment language as the subtitle language of the user interface.
[0141] Among them, the confidence of the environmental language meets the confidence condition, which means that the environmental language identified in real time is reliable and can be directly displayed as a subtitle language. It should be noted that the confidence of the environmental language meeting the confidence condition can refer to the confidence of a single language meeting the confidence condition, or at least one language in multiple target languages meeting the confidence condition. Further analysis can be done for the case where at least one language in multiple target languages meets the confidence condition. Specifically, if only one language in multiple target languages meets the confidence condition, then the language can be directly used as the subtitle language. If there are multiple languages in multiple target languages that meet the confidence condition, for example, if the confidence of two languages is more than 90%, then one of them can be selected as the subtitle language.
[0142] In a specific embodiment, Figure 5 A schematic diagram shows the interaction between various modules during the subtitle language matching process. When the display device is in adaptive subtitle display mode, the controller acquires the ambient audio signal collected and transmitted by the microphone in real time. Using the display device's intelligent voice assistant or self-trained language recognition model, the controller then performs language recognition on the audio signal, obtaining the ambient language and its confidence level. If the ambient language confidence level meets the confidence criteria, it is considered highly reliable and can be directly displayed as the subtitle language. If the ambient language confidence level does not meet the confidence criteria, further determination is made as to whether the display device has historical subtitle language records. If so, the language usage of each historical subtitle language in these records is obtained. Based on this usage, the stability of each historical subtitle language is analyzed. The language that meets the stability criteria is selected as the final subtitle language. If the ambient language confidence level does not meet the confidence criteria and no historical subtitle language records exist for the display device, the display device's default language is directly displayed as the subtitle language.
[0143] Furthermore, the subtitle language can be updated in real time based on the display device's subtitle display requirements or the display device's environment. For example, in a shopping mall environment, when the display device is displaying subtitles in mall mode, the subtitle language can be updated in real time. The controller can then loop through language recognition and confidence assessment based on the ambient audio signals captured in real time by the microphone, and update the final matching language as the subtitle language.
[0144] In this embodiment, when the confidence level of real-time language recognition is insufficient, the most stable language is selected as the subtitle language by analyzing the historical stability of subtitle language usage. It is understood that the most stable language directly reflects the recent subtitle language display requirements of the display device in the current scenario, and selecting it as the subtitle language ensures the accuracy and reliability of the subtitle language display. Furthermore, this embodiment uses looping language recognition to update the subtitle language in real time, making it more suitable for scenarios with high traffic and noisy environments such as shopping malls, effectively improving the environmental adaptability of the subtitle language.
[0145] In some embodiments, the present application also provides a subtitle display method, which is applied to the above-mentioned display device. Figure 6 As shown, the subtitle display method includes the following steps:
[0146] Step S602: When the display device is in an adaptive subtitle display mode, language recognition is performed based on an audio signal collected by a microphone of the display device to obtain an ambient language and a confidence level of the ambient language.
[0147] Step S604: when the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, then obtaining the language usage of each historical subtitle language in the historical subtitle language record;
[0148] Step S606: Analyze the usage stability of each historical subtitle language based on the usage of each language;
[0149] Step S608 : Filter out, from the various historical subtitle languages, a language whose usage stability satisfies a usage stability condition as the subtitle language of the display device.
[0150] In some embodiments, as Figure 7 As shown, from each historical subtitle language, a language whose usage stability meets the usage stability condition is screened out as the subtitle language of the display device, including: screening out candidate subtitle languages whose usage stability meets the usage stability condition from each historical subtitle language; if the candidate subtitle language is unique, the candidate subtitle language is used as the subtitle language of the display device; if there are multiple candidate subtitle languages, each candidate subtitle language is scored separately to obtain a scoring result for each candidate subtitle language; and screening out a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language as the subtitle language of the display device.
[0151] In some embodiments, as Figure 8As shown, if the candidate subtitle language is unique, the candidate subtitle language is used as the subtitle language of the display device, and also includes: when the candidate subtitle language is unique, if the environmental language is a single language, the candidate subtitle language is used as the subtitle language of the display device; if the environmental language includes multiple target languages, the candidate subtitle language is matched with each target language respectively; when the candidate subtitle language is consistent with any target language, the candidate subtitle language is used as the subtitle language of the display device.
[0152] In some embodiments, the subtitle display method further includes: when the candidate subtitle language is inconsistent with each target language, scoring the candidate subtitle language and each target language respectively to obtain scoring results for the candidate subtitle language and each target language; and using the language whose scoring results meet the second scoring requirements as the subtitle language of the display device.
[0153] In some embodiments, a selected subtitle language whose scoring result meets a first scoring requirement is screened out from each candidate subtitle language as the subtitle language of the display device, including: screening out a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language; if the environment language includes multiple target languages, matching the selected subtitle language with each target language for consistency; when the selected subtitle language is consistent with any target language, using the selected subtitle language as the subtitle language of the display device.
[0154] In some embodiments, the subtitle display method also includes: obtaining the subtitle display requirements of the scene in which the display device is located; when the subtitle display requirements meet the subtitle update conditions, performing language recognition again based on the current audio signal collected by the microphone to obtain the current environment language and the confidence of the current environment language; when the confidence of the current environment language meets the confidence conditions, updating the subtitle language of the display device to the current environment language.
[0155] In some embodiments, the subtitle display method further includes: if the confidence level does not meet the confidence condition and there is no historical subtitle language record for the display device within the historical period, obtaining the default language of the display device; and using the default language as the subtitle language of the display device.
[0156] In some embodiments, the subtitle display method further includes: performing volume analysis based on the audio signal to obtain the ambient volume.
[0157] In some embodiments, performing language recognition based on the audio signal further includes: performing language recognition based on the audio signal when the ambient volume is greater than the media volume of the display device.
[0158] In some embodiments, the subtitle display method further includes: if the confidence level of the environment language satisfies a confidence condition, using the environment language as the subtitle language of the user interface.
[0159] In a specific embodiment, Figure 9 A schematic diagram of the execution flow of the subtitle display method is shown. First, the method determines whether the display device has enabled adaptive subtitle display mode. If not, the traditional subtitle display mode (i.e., fixed subtitle display) is used. If adaptive subtitle display mode is enabled, the method proceeds to the dynamic subtitle language update phase. The first step in this dynamic subtitle language update phase can be initialization of subtitle-related modules (e.g., a language database storing historical subtitle language records). This initialization can be achieved by collecting language information fed back by other modules, such as an intelligent language assistant. It should be noted that if this is the first time subtitles are displayed, the second step can monitor the subtitles' on-screen status (not shown). If subtitles are already on, indicating that the user has selected their preferred language, the traditional subtitle display mode is used. If subtitles are not on, the method proceeds to the language identification phase. In this language identification phase, it is necessary to ensure that the language identification module is ready, such as by ensuring that the intelligent voice assistant or self-trained language identification model is in normal operation. Next, the ambient volume is compared with the display device's media volume. If the condition is met (the ambient volume is greater than the media volume), the language of the ambient audio signal is identified, thereby obtaining the ambient language and the corresponding confidence level. If the condition is not met (the ambient volume is not greater than the media volume), the process returns to the language identification module preparation step and continues to compare the ambient volume with the media volume of the display device. During language identification, if the recognition is successful, that is, the confidence level of the ambient language meets the confidence condition, the recognized ambient language is directly used as the subtitle language and the subtitles are enabled. Furthermore, the ambient language can be updated to the language database. If the recognition fails, that is, the confidence level of the ambient language does not meet the confidence condition, a suitable language is selected from the language database as the subtitle language and the subtitles are enabled. It is understood that if the confidence level of the ambient language meets the confidence condition, the recognition can be considered successful, otherwise it is considered a failure. Furthermore, the language recognition duration can be further considered to determine whether the recognition is successful. For example, if the language is not recognized within a specified time (e.g., 5 seconds), it can also be considered a recognition failure. Figure 10 A schematic diagram showing the process of real-time updating of subtitle language is shown. Figure 9 The difference between the processes shown is that Figure 10 The environment language will be identified in a loop, and the subtitle language will be updated in real time based on the language identification results.
[0160] In this embodiment, when the display device turns on the adaptive subtitle display mode, it will first perform language recognition based on the audio signal of the environment in which the display device is located to obtain the environment language and the confidence of the environment language. When the confidence of the environment language does not meet the confidence condition, it means that the currently identified environment language has a low credibility. At this time, it can be further determined whether the display device has a historical subtitle language record in the historical period. If so, the language usage of each historical subtitle language in the historical subtitle language record is obtained to analyze the usage stability of each historical subtitle language, and the language that meets the stability condition is used as the current subtitle language of the display device. It can be understood that the language with the most stable performance can directly reflect the recent subtitle language display requirements of the display device in the current scenario. In this way, when the real-time language recognition confidence is insufficient, this solution can effectively improve the reliability of the subtitle language by analyzing the usage stability of the historical subtitle language and selecting the most stable language as the subtitle language, thereby ensuring the subtitle display effect and improving the user experience.
[0161] In some embodiments, the present application further provides a subtitle display device, which is applied to the above-mentioned display device. In this embodiment, the subtitle display device includes:
[0162] A language recognition module is used to perform language recognition based on the audio signal collected by the microphone of the display device when the display device is in adaptive subtitle display mode, and obtain the ambient language and the confidence level of the ambient language;
[0163] A historical data acquisition module, configured to, when the confidence level does not meet the confidence condition, acquire the language usage of each historical subtitle language in the historical subtitle language record if the display device has a historical subtitle language record within a historical period;
[0164] The language stability analysis module is used to analyze the usage stability of each historical subtitle language based on the usage of each language;
[0165] The subtitle language matching module is used to select a language whose usage stability meets the usage stability condition from various historical subtitle languages as the subtitle language of the display device.
[0166] In some embodiments, the subtitle language matching module is further used to: screen out candidate subtitle languages whose usage stability meets the usage stability condition from various historical subtitle languages; if the candidate subtitle language is unique, use the candidate subtitle language as the subtitle language of the display device; if there are multiple candidate subtitle languages, score each candidate subtitle language separately to obtain a scoring result for each candidate subtitle language; and screen out a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language as the subtitle language of the display device.
[0167] In some embodiments, the subtitle language matching module is further used to: when the candidate subtitle language is unique, if the ambient language is a single language, use the candidate subtitle language as the subtitle language of the display device; if the ambient language includes multiple target languages, match the candidate subtitle language with each target language separately; when the candidate subtitle language matches any target language, use the candidate subtitle language as the subtitle language of the display device.
[0168] In some embodiments, the subtitle display device is also used to: when the candidate subtitle language is inconsistent with each target language, score the candidate subtitle language and each target language respectively to obtain the scoring results of the candidate subtitle language and each target language; and use the language whose scoring results meet the second scoring requirements as the subtitle language of the display device.
[0169] In some embodiments, the subtitle language matching module is further used to: screen out a selected subtitle language whose scoring result meets the first scoring requirement from each candidate subtitle language; if the environment language includes multiple target languages, the selected subtitle language is matched with each target language respectively; when the selected subtitle language is consistent with any target language, the selected subtitle language is used as the subtitle language of the display device.
[0170] In some embodiments, the subtitle display device is also used to: obtain the subtitle display requirements of the scene in which the display device is located; when the subtitle display requirements meet the subtitle update conditions, perform language recognition again based on the current audio signal collected by the microphone to obtain the current environment language and the confidence of the current environment language; when the confidence of the current environment language meets the confidence conditions, update the subtitle language of the display device to the current environment language.
[0171] In some embodiments, the subtitle display device is further used to: if the confidence does not meet the confidence condition and there is no historical subtitle language record for the display device within the historical period, then obtain the default language of the display device; and use the default language as the subtitle language of the display device.
[0172] In some embodiments, the subtitle display device is further used to: perform volume analysis based on the audio signal to obtain the ambient volume.
[0173] In some embodiments, the language identification module is further configured to perform language identification based on the audio signal when the ambient volume is greater than the media volume of the display device.
[0174] In some embodiments, the subtitle display device is further configured to: if the confidence level of the environment language satisfies a confidence condition, use the environment language as the subtitle language of the user interface.
[0175] Each module in the above-mentioned subtitle display device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0176] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method steps are implemented.
[0177] In some embodiments, a computer program product is provided, comprising a computer program, which implements the above method steps when executed by a processor.
[0178] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0179] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0180] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A display device, characterized in that: include: a display configured to display a user interface; a microphone configured to collect audio signals; The controller is configured as: When the display device is in an adaptive subtitle display mode, performing language recognition based on the audio signal to obtain an ambient language and a confidence level of the ambient language; When the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, obtaining the language usage of each historical subtitle language in the historical subtitle language record; Analyzing the usage stability of each of the historical subtitle languages based on the usage of each of the languages; From the historical subtitle languages, the language whose usage stability satisfies the usage stability condition is selected as the subtitle language of the display device.
2. The display device according to claim 1, wherein The controller is further configured to, in the process of screening out the language whose usage stability satisfies the usage stability condition from the historical subtitle languages as the subtitle language of the display device: Filtering out candidate subtitle languages whose usage stability satisfies a usage stability condition from the historical subtitle languages; If the candidate subtitle language is unique, using the candidate subtitle language as the subtitle language of the display device; If there are multiple candidate subtitle languages, score each candidate subtitle language separately to obtain a score result for each candidate subtitle language; A selected subtitle language whose scoring result meets a first scoring requirement is screened out from the candidate subtitle languages as the subtitle language of the display device.
3. The display device according to claim 2, wherein In the process of using the candidate subtitle language as the subtitle language of the display device if the candidate subtitle language is unique, the controller is further configured to: In the case where the candidate subtitle language is unique, if the environment language is a single language, the candidate subtitle language is used as the subtitle language of the display device; If the environment language includes multiple target languages, the candidate subtitle language is matched with each target language respectively; When the candidate subtitle language matches any of the target languages, the candidate subtitle language is used as the subtitle language of the display device.
4. The display device according to claim 3, wherein The controller is further configured to: When the candidate subtitle language is inconsistent with each of the target languages, scoring the candidate subtitle language and each of the target languages to obtain scoring results for the candidate subtitle language and each of the target languages; The language of the scoring result that meets the second scoring requirement is used as the subtitle language of the display device.
5. The display device according to claim 2, wherein: In the process of selecting the selected subtitle language whose scoring result meets the first scoring requirement from the candidate subtitle languages as the subtitle language of the display device, the controller is further configured to: Filtering out a selected subtitle language whose scoring result meets the first scoring requirement from the candidate subtitle languages; If the environment language includes multiple target languages, the selected subtitle language is matched with each target language respectively; When the selected subtitle language matches any of the target languages, the selected subtitle language is used as the subtitle language of the display device.
6. The display device according to claim 1, wherein The controller is further configured to: Obtaining subtitle display requirements for the scene in which the display device is located; When the subtitle display requirement satisfies the subtitle update condition, language recognition is performed again based on the current audio signal collected by the microphone to obtain the current environment language and the confidence level of the current environment language; When the confidence level of the current environment language satisfies a confidence condition, the subtitle language of the display device is updated to the current environment language.
7. The display device according to claim 1, wherein The controller is further configured to: If the confidence level does not meet the confidence condition, and there is no historical subtitle language record for the display device within a historical period, obtaining a default language for the display device; The default language is used as the subtitle language of the display device.
8. The display device according to claim 1, wherein The controller is further configured to: Performing volume analysis based on the audio signal to obtain an ambient volume; During the process of performing language recognition based on the audio signal, the controller is further configured to: When the ambient volume is greater than the media volume of the display device, language recognition is performed based on the audio signal.
9. The display device according to claim 1, wherein The controller is further configured to: If the confidence level of the environment language satisfies the confidence condition, the environment language is used as the subtitle language of the user interface.
10. A subtitle display method, characterized in that: The method comprises: When the display device is in an adaptive subtitle display mode, performing language recognition based on an audio signal collected by a microphone of the display device to obtain an ambient language and a confidence level of the ambient language; When the confidence level does not meet the confidence condition, if the display device has a historical subtitle language record within a historical period, obtaining the language usage of each historical subtitle language in the historical subtitle language record; Analyzing the usage stability of each of the historical subtitle languages based on the usage of each of the languages; From the historical subtitle languages, the language whose usage stability satisfies the usage stability condition is selected as the subtitle language of the display device.