Display device and identification method

By dynamically recording audio and image recognition based on the display interface status in the display device, the problem of invalid track recognition in the display device is solved, saving system resources and improving user experience.

CN120358380APending Publication Date: 2025-07-22JUHAOKAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409980.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When existing display devices identify tracks in media resources, regardless of whether the media resources contain audio data, they need to wait for the complete track recognition process to end, resulting in wasted system resources and user waiting time too long.

Method used

By in response to the target trigger signal, the content status displayed on the current display interface is determined, and audio data is recorded in the playback state for track recognition, and audio is not recorded in the unplayed state, combined with image recognition, it reduces the invalid recognition process and saves system resources.

Benefits of technology

Dynamic track recognition is realized, the invalid recognition process is reduced, the system resources are saved, the user experience is improved, and resource waste and waiting time are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358380A_ABST
    Figure CN120358380A_ABST
Patent Text Reader

Abstract

The invention discloses a display device and an identification method, and relates to the field of display, the display device comprises a display and a controller, the controller is configured to execute the following steps: receiving a target trigger signal sent by a remote controller, the target trigger signal representing that a target key on the remote controller is pressed; in response to the target trigger signal, determining the state of the content displayed by the current display interface; under the condition that the state of the content displayed on the current display interface is a playing state, recording an audio corresponding to the current display interface, and sending a song recognition request to a server according to target audio data generated by recording so as to perform song recognition; when the state of the content displayed by the current display interface is a non-playing state, not recording the audio corresponding to the current display interface; under the condition that the states of the contents displayed on the current display interface are different, different operations are executed on the audio corresponding to the current display interface, dynamic track identification is achieved, invalid identification processes are reduced, and system resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of displays, and in particular, to a display device and an identification method. Background Art

[0002] With the development of technology, display devices (such as televisions, mobile phones, computers, etc.) can provide users with diverse functions.

[0003] Currently, when a user uses a display device to watch media resources such as movies and TV dramas, if they want to know relevant information about the media resources, such as the program name, cast information, etc., and also want to know what the background music played in the media resource is, the display device can respond to the user's key operation on the remote control, record audio data of a fixed duration for the content currently displayed on the display interface of the display device, and perform track identification on the audio data to obtain an identification result. Generally, for the identification of audio data of media resources, the entire track identification process needs to be completed.

[0004] However, in practical applications, whether the media resource includes audio data or not, or whether it includes identifiable audio data or not, the user needs to wait for the track identification process to end, resulting in a waste of system resources. Summary of the Invention

[0005] This application provides a display device and an identification method, which can dynamically identify the audio corresponding to the current display interface according to the state of the content displayed on the current display interface.

[0006] In a first aspect, a display device is provided, including:

[0007] A display; a controller connected to the display, the controller being configured to: receive a target trigger signal sent by a remote control, where the target trigger signal indicates that a target key on the remote control is pressed, and the target key is the key on the remote control used to trigger track identification; in response to the target trigger signal, determine the state of the content displayed on the current display interface; in the case where the state of the content displayed on the current display interface is the playing state, record the audio corresponding to the current display interface, and send a track identification request to a server based on the generated target audio data for track identification; in the case where the state of the content displayed on the current display interface is the non-playing state, do not record the audio corresponding to the current display interface.

[0008] In an embodiment of the present application, the state of the content displayed on the current display interface is the playing state, which also indicates that the state of the current media asset is the playing state, indicating that the display device will output audio and video signals, and the track recognition process can be executed. While the state of the content displayed on the current display interface is the non-playing state, it indicates that the display device does not output audio and video signals, and the audio corresponding to the current display interface is not recorded to avoid executing the track recognition process, thereby enabling dynamic track recognition, reducing invalid track recognition processes, saving system resources, and avoiding resource waste.

[0009] In a possible implementation manner, after receiving the target trigger signal sent by the remote controller, the controller is further configured to: in response to the target trigger signal, take a screenshot of the content displayed on the current display interface, and send an image recognition request to the server based on the screenshot image obtained from the screenshot for image recognition.

[0010] In an embodiment of the present application, by taking a screenshot of the content displayed on the current display interface in response to the target trigger signal to obtain a screenshot image, and then performing image recognition based on the screenshot image, an image recognition result can be obtained. Moreover, by performing image recognition through the server, the recognition accuracy can be improved, the image recognition effect can be enhanced, and at the same time, the power consumption of the display device can be reduced.

[0011] In a possible implementation manner, the controller determines the state of the content displayed on the current display interface in response to the target trigger signal, including: the controller determines the state of the content displayed on the current display interface in response to the target trigger signal, and determines the media asset type corresponding to the content displayed on the current display interface; the controller executes recording the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state, including: the controller records the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state and the media asset type corresponding to the content displayed on the current display interface does not belong to the preset media asset type; the controller does not record the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state and the media asset type corresponding to the content displayed on the current display interface belongs to the preset media asset type.

[0012] Among them, the preset media asset type is a media asset type that does not include music, and the preset media asset type may include media assets such as picture book media assets, educational media assets, and sports media assets.

[0013] In the embodiments of the present application, when it is determined that the state of the content displayed on the current display interface is the playing state, it is further possible to determine whether the media type of the content displayed on the current display interface is a preset media type. When the media type of the content displayed on the current display interface is a preset media type, the audio corresponding to the current display interface is not recorded, and the voice recognition process is not executed. In this way, the voice recognition process can be triggered more precisely, avoiding an invalid voice recognition process. Thus, only the image recognition result is required, which can shorten the response time and improve the user experience. Moreover, not performing the invalid voice recognition process can save system resources and avoid resource waste.

[0014] In addition, when it is determined that the state of the content displayed on the current display interface is the playing state, if the media type of the content displayed on the current display interface is not a preset media type, the audio corresponding to the current display interface is recorded to execute the voice recognition process, which can improve the success rate of voice recognition and enhance the user experience.

[0015] In a possible implementation, the application corresponding to the content displayed on the current display interface is the desktop application of the display device.

[0016] In the embodiments of the present application, when the application corresponding to the content displayed on the current display interface is the desktop application of the display device, different operations can be performed according to the state of the content displayed on the current display interface, so as to realize dynamic track recognition, reduce the invalid track recognition process, save system resources, and avoid resource waste.

[0017] In a possible implementation, before determining the state of the content displayed on the current display interface in response to the target trigger signal, the controller is further configured to: determine whether the application corresponding to the content of the current display interface is a desktop application; when the application corresponding to the content of the current display interface is a desktop application, determine the state of the content displayed on the current display interface in response to the target trigger signal; when the application corresponding to the content of the current display interface is a non-desktop application, send a recognition request to the server, including: sending a recognition request including a first parameter to the server, where the first parameter indicates that the application corresponding to the content of the current display interface is a non-desktop application, and the recognition request is further used to enable the server to determine the interface type corresponding to the screenshot image according to the screenshot image; receive the interface type corresponding to the screenshot image feedback by the server, where the interface type corresponding to the screenshot image is determined by the server according to the screenshot image; when the interface type corresponding to the screenshot image belongs to the preset interface type, record the audio corresponding to the current display interface; when the interface type corresponding to the screenshot image does not belong to the preset interface type, do not record the audio corresponding to the current display interface.

[0018] In the embodiments of the present application, the application programs installed in the display device include a desktop application program and a non-desktop application program. In this regard, it is possible to first determine whether the application program corresponding to the content displayed on the current display interface is a desktop application program. In this way, for different types of application programs, different methods are adopted to determine whether to record audio and execute the voice recognition process, which can trigger the voice recognition process more accurately.

[0019] Moreover, when the application program corresponding to the content displayed on the current display interface is a non-desktop application program, it is possible to take a screenshot of the content displayed on the current display interface, and determine the interface type corresponding to the current display interface according to the screenshot image obtained by the screenshot, so as to determine whether to record the audio corresponding to the current display interface according to the interface type corresponding to the current display interface. When the interface type corresponding to the current display interface belongs to the preset interface type, record the audio corresponding to the current display interface. When the interface type corresponding to the current display interface does not belong to the preset interface type, do not record the audio corresponding to the current display interface.

[0020] In addition, after obtaining the screenshot image, send a recognition request to the server according to the screenshot image obtained by the screenshot. The recognition request includes a first parameter indicating that the application program corresponding to the content of the current display interface is a non-desktop application program. The recognition request is used to request the server to determine the interface type corresponding to the current display interface according to the screenshot image, which can improve the accuracy and the processing efficiency.

[0021] In a possible implementation manner, the controller executes sending a song recognition request to the server according to the generated target audio data for song recognition, and is specifically configured to: when the target audio data meets the first preset condition, send a song recognition request to the server according to the first audio segment in the target audio data for song recognition, and the target audio data includes multiple audio frames; wherein, the first preset condition is that the number of valid audio frames in the first audio segment is greater than or equal to the first threshold, and the first audio segment is the first preset number of consecutive audio frames in the target audio data. The controller is further configured to: when the target audio data does not meet the first preset condition, do not send a song recognition request to the server.

[0022] In the embodiments of the present application, when the number of valid audio frames in the first audio segment in the target audio data is greater than or equal to the first threshold, indicating that there is an identifiable audio segment in the target audio data, the identifiable audio segment is recognized, which improves the recognition accuracy and avoids invalid recognition.

[0023] In a possible implementation, the controller is further configured to: for a first audio frame in the target audio data, use an audio signal classification model to process the first audio frame to obtain a first probability, where the first probability is the probability that the first audio frame is a valid audio frame, and the first audio frame is any one of a plurality of audio frames; and determine that the first audio frame is a valid audio frame when the first probability is greater than or equal to a preset probability.

[0024] In the embodiments of the present application, an audio frame with a probability greater than or equal to a preset probability is determined as a valid audio frame, which improves the accuracy of determining a valid audio frame.

[0025] In a second aspect, an identification method is provided, which is applied to the display device in the first aspect. The method includes: receiving a target trigger signal, where the target trigger signal indicates that a target button on the remote control is pressed, and the target button is a button on the remote control for triggering track identification; in response to the target trigger signal, determining the state of the content displayed on the current display interface; when the state of the content displayed on the current display interface is the playing state, recording the audio corresponding to the current display interface, and sending a track identification request to the server for track identification according to the generated target audio data; and when the state of the content displayed on the current display interface is the non-playing state, not recording the audio corresponding to the current display interface.

[0026] In a third aspect, an identification device is provided, including units for executing the identification method in the second aspect. The device may be a terminal device or a chip in the terminal device.

[0027] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is run by the identification device, the identification device is caused to execute the identification method in the second aspect.

[0028] In a fifth aspect, a computer program product is provided. The computer program product includes: a computer program, and when the computer program is run by the identification device, the identification device is caused to execute the identification method in the second aspect.

[0029] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram showing an operation scenario between a display device and a control device provided in some embodiments of the present application is shown;

[0031] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown;

[0032] Figure 3 Exemplarily shows a configuration block diagram of a display device 200 according to an exemplary embodiment;

[0033] Figure 4 Exemplarily shows a system block diagram in a display device 200 according to an exemplary embodiment;

[0034] Figure 5 Shows a schematic diagram of an identification process and an identification result in the prior art provided by an embodiment of the present application;

[0035] Figure 6 Shows a schematic flowchart of an identification method provided by an embodiment of the present application;

[0036] Figure 7 Shows a schematic flowchart of another identification method provided by an embodiment of the present application;

[0037] Figure 8 Shows a schematic flowchart of yet another identification method provided by an embodiment of the present application;

[0038] Figure 9 Shows a schematic diagram of steps for determining whether to record audio corresponding to the current display interface provided by an embodiment of the present application;

[0039] Figure 10 Shows a user interface provided by an embodiment of the present application;

[0040] Figure 11 Shows a schematic structural diagram of an identification device provided by an embodiment of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.

[0042] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0043] For purposes of illustration and not limitation, specific details such as specific system architectures, technologies, etc. are provided to facilitate a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0044] As used in this application, the term "remote control" refers to a component of a display device (such as the display device disclosed in this application), which can generally wirelessly control the display device within a relatively short distance range. It is generally connected to the display device through infrared and / or radio frequency (RF) signals and / or Bluetooth, and can also include functional modules such as WiFi, wireless USB, Bluetooth, and motion sensors. For example: A handheld touch remote control replaces most of the physical built-in hard keys in a general remote control device with a user interface on a touch screen.

[0045] Embodiments of the present application provide a display device, and the display device provided by the embodiments of the present application can have various implementation forms. For example, it can be a television, a smart TV, a laser projection device, a monitor, an electronic bulletin board, etc. Figure 1 and Figure 2 is a specific implementation manner of the display device of the present application.

[0046] Figure 1 is a schematic diagram of the operation scenario between the display device and the control device according to the embodiment. As Figure 1 shown, this operation scenario includes: a control device 100, a display device 200, and a smart device 300. The user can operate the display device 200 through the smart device 300 or the control device 100.

[0047] In some embodiments, the user operates the display device 200 through the control device 100. The control device 100 can be a remote control. The communication between the remote control and the display device includes infrared protocol communication or Bluetooth protocol communication, and other short-distance communication methods, and controls the display device 200 wirelessly or wiredly. The user can input user instructions through buttons on the remote control, voice input, control panel input, etc. to control the display device 200.

[0048] In some embodiments, the user can also control the display device 200 through the smart device 300 (such as: a mobile terminal, a tablet computer, a computer, a laptop, etc.). For example, use an application program running on the smart device to control the display device 200.

[0049] In some embodiments, the display device 200 may also be controlled in ways other than the control device 100 and the smart device 300. For example, it may directly receive the user's voice commands for control through a module for obtaining voice commands configured inside the display device 200, or may receive the user's voice commands for control through a voice control device provided outside the display device 200.

[0050] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 is allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 may provide various contents and interactions to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers. For example, the server 400 may be a data processing server in any form such as a cloud server or a distributed server.

[0051] Figure 2 Exemplarily shown is a configuration block diagram of the control device 100 according to an exemplary embodiment. As Figure 2 shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive the operation instructions input by the user and convert the operation instructions into instructions recognizable and responsive by the display device 200, serving as an interaction intermediary between the user and the display device 200.

[0052] Next, through Figure 3 embodiments, the structure of the display device provided by the present application will be described in detail. Figure 3 This is a schematic structural diagram of the display device provided by an embodiment of the present application. As Figure 3 shown, the display device 200 includes at least one of a tuner demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.

[0053] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, a RAM, a ROM, and a first interface to an nth interface for input / output.

[0054] The display 260 includes a display screen component for presenting a picture, and a driving component for driving image display, which is a component for receiving an image signal output from the controller and displaying video content, image content, and a menu manipulation interface, as well as a user manipulation UI interface.

[0055] The display 260 may be a liquid crystal display, an OLED display, and a projection display, and may also be a projection device and a projection screen.

[0056] The communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips such as a near-field communication protocol chip, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the control device 100 or the server 400 through the communicator 220.

[0057] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control, etc.).

[0058] The detector 230 is used to collect signals from the external environment or for external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0059] The external device interface 240 includes at least one High Definition Multimedia Interface (HDMI), and may also include, but is not limited to, any one or more of the following: analog or digital high-definition component input interfaces (component), composite video input interfaces (CVBS), RGB ports, etc., a Display Port (DP), a USB interface (such as USB Type-C). It can also be a composite input / output interface formed by the above multiple interfaces.

[0060] Among them, HDMI is a digital video / audio interface that can transmit audio and video signals simultaneously, eliminating the need for additional audio connection cables, simplifying the installation difficulty of the lines, and ensuring high-quality audio and video signal transmission. And with the development of HDMI technology, the width of HDMI is continuously increasing, having a higher transmission bandwidth, which can support high-definition, ultra-high-definition, and even higher-resolution video transmission, and can meet the growing demand for high-definition video.

[0061] It should be understood that various external devices such as game consoles, Blu-ray players, set-top boxes, etc. can be connected to the display device 200 through HDMI, enabling users to obtain a richer and more immersive audio-visual entertainment experience.

[0062] It should also be understood that, for the convenience of distinguishing and managing HDMI interfaces, the display device and the system of the display device will label the HDMI interfaces, and different HDMI interfaces can be represented by numbers, letters, or specific symbols. For example, HDMI 1, HDMI2, or HDMI A, HDMI B, etc.

[0063] In the system of the display device, each HDMI interface has its own independent interface code, which can be identified by a digital code and has a corresponding relationship with the label of the HDMI interface. For example, HDMI 1 corresponds to the number 1281, and HDMI 2 corresponds to the number 1282.

[0064] The tuner demodulator 210 receives broadcast television signals through wired or wireless reception methods, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.

[0065] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different separate devices, that is, the tuner demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0066] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example: in response to receiving a user command for selecting a UI object to be displayed on the display 260, the controller 250 can perform operations related to the object selected by the user command.

[0067] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), a random access memory (RAM), a read-only memory (ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0068] The user can input a user command on the graphical user interface (GUI) that can be displayed on the display 260, and then the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and then the user input interface receives the user input command by recognizing the sound or gesture through a sensor.

[0069] "User interface" is a media interface for interaction and information exchange between an application or an operating system and a user. It realizes the conversion between the internal form of information and the form acceptable to the user. The common manifestation form of the user interface is the Graphic User Interface (GUI), which refers to the user interface related to computer operation displayed in a graphical manner. It can be an interface element such as an icon, a window, a control, etc. displayed on the display screen of an electronic device, where the control can include visible interface elements such as an icon, a button, a menu, a tab, a text box, a dialog box, a status bar, a navigation bar, a Widget, etc.

[0070] Such as Figure 4 For some embodiments of the present application Figure 1 is a schematic diagram of the system configuration of the display device. In some embodiments, the system is divided into four layers, from top to bottom are the Applications layer (abbreviated as "application layer"), the Application Framework layer (abbreviated as "framework layer"), the Android runtime and the System Libraries layer (abbreviated as "system runtime library layer"), and the kernel layer.

[0071] In some embodiments, at least one application program runs in the application layer. These application programs can be window programs, system setting programs, or clock programs, etc. that come with the operating system; they can also be application programs developed by third-party developers. In specific implementation, the application program packages in the application layer are not limited to the above examples.

[0072] The application framework layer provides application programming interfaces (APIs) and programming frameworks for application programs. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center, and this center determines the actions of the application programs in the application layer. Application programs can access the resources in the system and obtain system services through the API interface during execution.

[0073] Such as Figure 4As shown in the figure, in the embodiment of the present application, the application framework layer includes managers (Managers), content providers (Content Provider), etc. Among them, the managers include at least one of the following modules: The ActivityManager is used to interact with all the activities running in the system; the Location Manager is used to provide access to the system location service for system services or applications; the Package Manager is used to retrieve various information related to the application packages currently installed on the device; the NotificationManager is used to control the display and clearing of notification messages; the Window Manager is used to manage icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0074] In some embodiments, the ActivityManager is used to manage the life cycles of various applications and the usual navigation back functions, such as controlling the exit, opening, and backward of applications. The Window Manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the changes of the display window (such as shrinking the display window, jittering the display, distorting the display, etc.).

[0075] In some embodiments, the system runtime layer provides support for the upper layer, i.e., the framework layer. When the framework layer is used, the Android operating system will run the C / C++ libraries included in the system runtime layer to implement the functions to be achieved by the framework layer.

[0076] In some embodiments, the kernel layer is the layer between hardware and software. As Figure 4 shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.

[0077] The related technologies of the embodiments of the present application will be described below.

[0078] When a user uses a display device to watch media resources such as movies and TV series, if they want to know the relevant information of the currently playing media resources, such as: program name, cast information, etc., and at the same time want to know what the background music in the media resource is, the display device receives a user instruction through the remote control, takes a screenshot of the current display interface of the display device and performs image recognition. At the same time, it records audio data for a fixed duration and performs track recognition on the audio data to obtain a track recognition result. Generally, to recognize the audio data of a media resource, the entire track recognition process needs to be completed.

[0079] However, in practical applications, since the display device needs to follow the complete track recognition process (including steps such as audio sampling, feature extraction, data matching, and result return), regardless of whether the current media resource contains a valid audio signal, for example: there is no valid audio signal in a silent scene or a dialogue segment without background music, or whether the audio data has recognizable features, for example: low-quality sound sources or audio data with large environmental noise interference may not have recognizable features, the user needs to wait until the entire track recognition process ends to obtain the final feedback. As Figure 5 shown, whether it is music recognition in a scene with or without movie soundtracks, it starts from clicking "Identify Song by Listening" on the music recognition card and waiting for the music recognition time (e.g., 15s) to obtain the track recognition result. In some scenes, the track recognition result is obtained, while in some scenes, "No track recognition result found" is obtained. This indiscriminate processing of each media resource for track recognition leads to a large amount of ineffective waiting time in the track recognition process. Especially in scenes where the track recognition success rate is relatively low, the display device may require a longer time in the processes of feature extraction and data matching of the audio data, making the entire track recognition process time-consuming, requiring the user to wait for a long time, and continuously occupying system resources, resulting in a waste of system resources.

[0080] In view of this, embodiments of the present application provide a display device and an identification method. By responding to a target trigger signal, the state of the content displayed on the current display interface can be determined; when the state of the content displayed on the current display interface is the playing state, it indicates that the current media asset is in the playing state, characterizing that the display device will output audio and video signals. Record the audio corresponding to the current display interface, and send a song recognition request to the server for track recognition based on the generated target audio data; when the state of the content displayed on the current display interface is the non-playing state, it characterizes that the display device does not output audio and video signals and does not record the audio corresponding to the current display interface to avoid executing the track recognition process, so as to dynamically perform track recognition on the current display interface, reduce ineffective track recognition processes, save system resources, and avoid resource waste.

[0081] To facilitate a further understanding of the technical solutions in some embodiments of the present application, the technical solutions of the display device and the identification method, and how this technical solution solves the above technical problems will be described in detail below with reference to some specific embodiments and drawings. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.

[0082] The display device includes a display and a controller, and the controller is communicatively connected to the display. The display device may further include at least one HDMI interface for connecting the display device to an external device.

[0083] Figure 6 The flowchart of the recognition method in the display device in some embodiments of the present application is shown. As Figure 6 shown, taking a smart TV as an example of the display device, wherein the controller is configured to perform the following steps:

[0084] S610. Receive a target trigger signal sent by the remote control.

[0085] Wherein, the target trigger signal indicates that a target button on the remote control is pressed, and the target button is the button on the remote control for triggering track recognition.

[0086] In some embodiments, the user presses one or more target buttons on the remote control, so that the remote control can generate a corresponding key code according to the key information of the one or more target buttons, modulate to generate an infrared or radio frequency signal, and then send the infrared or radio frequency signal to the display device, so that the display device receives and decodes the infrared or radio frequency signal (i.e., the target trigger signal), so that the display device can perform functions mapped according to the target button corresponding to the target trigger signal.

[0087] In some embodiments, the target trigger signal is mapped to a track recognition instruction.

[0088] S620. In response to the target trigger signal, determine the state of the content displayed on the current display interface.

[0089] Exemplarily, the present application embodiment provides the specific process for the controller to execute S620. Specifically, the framework layer in the operating system converts the target trigger signal into a target event corresponding to the target button, and the framework layer sends the target event to the application layer. The application layer receives the target event and triggers the first application program to execute the target event according to the predefined logic. It can be understood that before triggering the first application program to execute the target event, the system can first query whether the first application program has been started. If the first application program has not been started, then in response to the target trigger signal, first start the first application program and then execute the target event. If the first application program has already been started, then in response to the target trigger signal, directly execute the target event, and the target event is to determine the state of the content displayed on the current display interface.

[0090] In some embodiments, the target trigger signal is mapped to a track recognition instruction. In response to the target trigger signal, the first application obtains the attributes of the content displayed on the current display interface from the second application, and the first application determines the state of the content displayed on the current display interface according to the attributes of the content displayed on the current display interface returned by the second application.

[0091] It can be understood that the attributes of the content displayed on the current display interface include: the state of the content displayed on the current display interface. It should be noted that in this example, the application corresponding to the content displayed on the current display interface is the second application.

[0092] In some embodiments, the state of the content displayed on the current display interface includes a playing state and a non-playing state. The non-playing state may specifically further include: a paused state, a stopped state, a background buffering state, etc. The above states can be characterized by different parameters.

[0093] S6301. When the state of the content displayed on the current display interface is the playing state, record the audio corresponding to the current display interface, and send a track recognition request to the server based on the generated target audio data for track recognition.

[0094] In some embodiments, the audio corresponding to the current display interface is audio information synchronized or related to the content displayed on the current display interface. The audio output by the microphone (the current sound stream) can be recorded by calling the system audio interface to obtain the target audio data. Exemplarily, the target audio data can be sampled in Pulse Code Modulation (PCM) format, with a sampling rate of 16 kHz and mono-channel. Of course, the target audio data can also be sampled in other forms, which is not limited in this application.

[0095] In the embodiments of the present application, the state of the content displayed on the current display interface being the playing state means that the state of the media resource displayed on the display device is the playing state, and the display device is outputting audio and video signals. For effective track recognition of the content displayed on the current display interface, at this time, the system port can be called to record the audio corresponding to the current display interface to obtain the target audio data, and the track recognition request including the target audio data is sent to the server, so that the server performs track recognition on the target audio data according to the track recognition request.

[0096] In some embodiments, track recognition refers to matching and identifying the collected target audio data in a preset database to feedback the identified information. The information identified by track recognition can be the information of any one of various types of media resources such as operas, music, variety shows, etc. For example, the information identified for media resources of the music type may include: song name, song writer, lyrics, play address, download address, singer, etc.; the information identified for opera media resources may include: opera name, opera singer, brief introduction of opera-related background knowledge, etc.

[0097] It should also be noted that when the server receives the target audio data, it performs track recognition on the target audio data according to the track recognition request to obtain the information corresponding to the target audio. After that, the server can compare the target audio data with the preset database to find the audio data matching the target audio data in the preset database. Then, the information corresponding to the matched audio data is sent to the display device, and the display device receives and displays the information corresponding to the matched audio data on the display.

[0098] S6302. When the content displayed on the current display interface is in the unplayed state, do not record the audio corresponding to the current display interface.

[0099] In the embodiments of the present application, the state of the content displayed on the current display interface being in the unplayed state means that the state of the media resource displayed on the display device is in the unplayed state, and the display device does not output audio and video signals at this time. If track recognition is performed on the content displayed on the current display interface, it is an invalid recognition. Therefore, do not record and recognize the audio corresponding to the current display interface, thereby reducing the invalid track recognition process, saving system resources, and avoiding waste of system resources. At the same time, the user does not need to wait for the track recognition process, saving the user's time and improving the user experience.

[0100] In some embodiments, when the content displayed on the current display interface is in the unplayed state, the display device also displays a prompt message to prompt the user that track recognition is not currently being performed.

[0101] In the embodiments of the present application, executing S6302 means that the track recognition process ends. After S6302 does not record the audio corresponding to the current display interface, the display device can display a prompt message of "Tracks cannot be recognized in the current scene" on the display interface to prevent the user from continuing to wait for the track recognition result.

[0102] The embodiments of the present application provide a display device and an identification method, which can determine the state of the content displayed on the current display interface in response to a target trigger signal; when the state of the content displayed on the current display interface is the playing state, it indicates that the state of the current media asset is the playing state, characterizing that the display device is outputting audio and video signals, recording the audio corresponding to the current display interface, and sending a song recognition request to the server according to the generated target audio data for song recognition; when the state of the content displayed on the current display interface is the non-playing state, it characterizes that the display device does not output audio and video signals and does not record the audio corresponding to the current display interface, so that the song recognition of the current display interface can be dynamically realized, reducing the invalid song recognition process, saving system resources and avoiding resource waste.

[0103] In some embodiments, the image recognition operation and the song recognition operation are usually executed synchronously. Therefore, in the embodiments of the present application, after S610, the controller is further configured to: in response to a target trigger signal, take a screenshot of the content displayed on the current display interface, and send an image recognition request to the server according to the captured screenshot image for image recognition.

[0104] Obviously, in this embodiment, the target trigger signal can be mapped to a song recognition instruction and an image recognition instruction, or the target trigger signal is mapped to a target instruction, and the target instruction includes a song recognition instruction and an image recognition instruction.

[0105] In some embodiments, the controller is configured to: in response to the image recognition instruction mapped by the target trigger signal, take a screenshot of the content displayed on the current display interface to obtain a screenshot image, and send the image recognition request including the screenshot image to the server, so that the server performs image recognition on the screenshot image according to the image recognition request to obtain an image recognition result.

[0106] In order to further reduce the invalid song recognition and save system resources, when it is determined that the content displayed on the current display interface is in the playing state, the media type corresponding to the content displayed on the current display interface can be further judged to exclude the media types that do not require song recognition, thereby reducing the invalid song recognition, saving system resources and avoiding the waste of system resources.

[0107] Exemplarily, as Figure 7 The flowchart of the identification method in the display device in some embodiments of the present application shown, the controller executing S620 includes:

[0108] S620A. In response to a target trigger signal, determine the state of the content displayed on the current display interface and determine the media type corresponding to the content displayed on the current display interface.

[0109] In some embodiments, the media type corresponding to the content displayed on the current display interface may be included in the attributes of the content displayed on the current display interface and obtained together with the state of the content displayed on the current display interface.

[0110] In some embodiments, the preset media type is a media type that does not include music. The preset media type may include picture book media, educational media, sports media, etc. Among them, picture book media is that while the display device flips the pages like a book (picture book) as the picture is displayed, it outputs human voices such as storytelling or book explanation that are the same as those of the picture book. Educational media is related to learning, such as media for explaining topics. Sports media is related videos of various sports events. The preset media type can also be understood as media that has the value of music recognition, or media in which the probability of including music exceeds the probability threshold.

[0111] The controller executes S6301 and is specifically configured as follows:

[0112] S6301A: When the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface does not belong to the preset media type, record the audio corresponding to the current display interface.

[0113] When the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface does not belong to the preset media type, it indicates that the display device has audio and video signal output, and the content displayed on the current display interface has the value of music recognition, and there is a high probability that the track can be recognized. Therefore, record the audio corresponding to the current display interface.

[0114] In some embodiments, the controller is further configured as follows:

[0115] S6301B: When the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface belongs to the preset media type, do not record the audio corresponding to the current display interface.

[0116] When the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface belongs to the preset media type, it indicates that the display device has audio and video signal output, but the content displayed on the current display interface has no value of music recognition, and there is a high probability that the track cannot be recognized. Therefore, do not record the audio corresponding to the current display interface, thereby reducing the ineffective recognition process and saving system resources.

[0117] In the embodiments of the present application, the application program corresponding to the content displayed on the current display interface is the desktop application program of the display device.

[0118] The desktop application of the display device is installed by default in the operating system of the display device at the time of factory, and can be used as an application for the signal source of the display device.

[0119] The "second application" in the foregoing text is the desktop application of the display device.

[0120] Since the source of the content displayed on the current display interface can be the desktop application of the display device or a non-desktop application. For example, it can be a third-party application installed in the display device, or it can be the content transmitted to the display device by an external device connected to the display device through a high-definition multimedia interface.

[0121] In the case where the sources of the content displayed on the current display interface are different, the embodiments of the present application provide different methods to determine whether to perform track recognition on the content displayed on the current display interface, which will be specifically introduced below.

[0122] In some embodiments, before S620, the controller is further configured to:

[0123] Determine whether the application corresponding to the content of the current display interface is a desktop application.

[0124] Exemplarily, the application corresponding to the content of the current display interface is the current top-level application. The first application can send a request to the operating system for the list of currently running applications, and determine the identifier of the current top-level application from the list of currently running applications. Alternatively, the first application can directly request the identifier of the current top-level application from the operating system. In the case of obtaining the identifier of the current top-level application, match the identifier of the current top-level application with the identifier of the desktop application of the operating system. If the match is successful, it is determined that the application corresponding to the content of the current display interface is a desktop application. If the match fails, it is determined that the application corresponding to the content displayed on the current display interface is a third-party application, not a desktop application, that is, a non-desktop application.

[0125] Exemplarily, if the source of the content displayed on the current display interface is transmitted by an external device connected to the display device through a high-definition multimedia interface, and the data packet name transmitted by the external transmission device is obtained from the operating system instead of the identifier of the application, then match the data packet name with the desktop application name of the operating system. If the match fails, it is determined that the application corresponding to the content displayed on the current display interface is a non-desktop application.

[0126] In some embodiments, in the case where the application corresponding to the content of the current display interface is a desktop application, execute S620, S6301 or S6302.

[0127] When the application corresponding to the content displayed on the current display interface is a non-desktop application, the status corresponding to the content displayed on the current display interface cannot be obtained. In this case, it is possible to further filter the display interfaces in the content displayed on the current display interface that do not require the execution of the track recognition process by whether the interface type corresponding to the current display interface belongs to the preset interface type.

[0128] Exemplarily, as Figure 8 shown in the schematic flowchart of the recognition method in the display device in some other embodiments of the present application, the controller is further configured to:

[0129] S810. When the application corresponding to the content displayed on the current display interface is a non-desktop application, send a recognition request including a first parameter to the server.

[0130] Exemplarily, when it is determined that the application corresponding to the content of the current display interface is not a desktop application, that is, a non-desktop application, a first parameter indicating that the application corresponding to the content displayed on the current display interface is a non-desktop application can be carried in the view request, so that after receiving the recognition request, the server can determine the interface type corresponding to the screenshot image according to the screenshot image.

[0131] S820. Receive the interface type corresponding to the screenshot image feedback by the server.

[0132] Wherein, the interface type corresponding to the screenshot image is determined by the server according to the screenshot image.

[0133] Specifically, the server inputs the screenshot image into the interface type recognition model, and the interface type recognition model outputs the interface type corresponding to the screenshot image. Among them, the interface type recognition model is a pre-trained model.

[0134] S8301. When the interface type corresponding to the screenshot image belongs to the preset interface type, record the audio corresponding to the current display interface.

[0135] In some embodiments, the preset interface type is an interface type including a music scene. As an example rather than a limitation, the preset interface types include: video playback interface, music playback interface, concert interface, MV playback interface. The interface types included in the preset interface types are interface types that very likely contain music and have the value of song recognition.

[0136] Exemplarily, when the interface type corresponding to the screenshot image belongs to the preset interface type, it means that the content played on the current display interface very likely contains music and has the value of song recognition. Therefore, record the audio corresponding to the current display interface and execute the track recognition process.

[0137] S8302. When the interface type corresponding to the screenshot image does not belong to the preset interface type, do not record the audio corresponding to the currently displayed interface.

[0138] Exemplarily, when the interface type corresponding to the screenshot image does not belong to the preset interface type, it indicates that the content played on the currently displayed interface probably does not contain music and has no value for song recognition. Therefore, do not record the audio corresponding to the currently displayed interface.

[0139] By filtering the displayed interfaces that do not need to execute the song recognition process through the "preset interface type", the invalid recognition process is effectively reduced, thereby saving system resources and avoiding waste of system resources.

[0140] In order to improve the accuracy of song recognition and avoid invalid recognition, for this reason, the embodiments of the present application provide that it is possible to first determine whether the target audio data contains recognizable valid audio segments. If the target audio data contains recognizable valid audio segments, recognize the valid audio segments. The following is a specific introduction to this process:

[0141] The controller executes the song recognition request to the server according to the target audio data generated by recording in S6301, and is specifically configured as:

[0142] When the target audio data meets the first preset condition, send a song recognition request to the server according to the first audio segment in the target audio data for song recognition. The target audio data includes multiple audio frames.

[0143] Among them, the first preset condition is that the number of valid audio frames in the first audio segment is greater than or equal to the first threshold, and the first audio segment is the first preset number of consecutive audio frames in the target audio data.

[0144] In some embodiments, the first preset number is set to 10 frames, and the first threshold is set to 8 frames. In this way, the first preset condition is that the number of valid audio frames in 10 consecutive frames of the target audio data is greater than or equal to 8 frames. Among them, taking 8 valid music frames as an example, it can be 8 consecutive frames in 10 consecutive frames, for example: the 1st frame to the 8th frame, the 2nd frame to the 9th frame, the 3rd frame to the 10th frame, the Nth frame to the N + 8th frame, or it can also be 8 non - consecutive frames in 10 consecutive frames, for example: the 1st frame and the 3rd frame to the 9th frame, the 9th, 10th frames, the 12th frame, and the 14th - 18th frames.

[0145] The following explains how to determine the valid audio frames in the target audio data.

[0146] Specifically, for the first audio frame in the target audio data, an audio signal classification model is used to process the first audio frame to obtain a first probability. When the first probability is greater than or equal to a preset probability, the first audio frame is determined to be a valid audio frame.

[0147] Wherein, the first probability is the probability that the first audio frame is a valid audio frame, and the first audio frame is any one of multiple audio frames.

[0148] In some embodiments, when the first probability is less than the preset probability, the first audio frame is determined to be an invalid audio frame.

[0149] In some embodiments, the preset probability is set to 90%. When the first probability is greater than or equal to 90%, the first audio frame is determined to be a valid audio frame. When the first probability is less than 90%, the first audio frame is determined to be an invalid audio frame.

[0150] First, an explanation is given on how to obtain the first audio frame in the target audio data:

[0151] After obtaining the target audio data, it can be segmented according to a preset duration and a preset frame shift amount to generate a continuous audio frame sequence. The continuous audio frame sequence is the above-mentioned multiple audio frames, and the first audio frame is one frame in the continuous audio frame sequence.

[0152] In some embodiments, the preset duration can be set to 20 ms and the preset frame shift amount can be set to 10 ms to obtain multiple audio frames. It can be understood that there is frame overlap between audio frames.

[0153] It can be understood that the target audio data can also be preprocessed to obtain the preprocessed target audio data, and then the preprocessed target audio data is segmented according to the preset duration and the preset frame shift amount to generate a continuous audio frame sequence. Among them, the preprocessing includes but is not limited to: noise reduction processing. Exemplarily, a digital filter (such as: FIR filter) is used to perform noise reduction processing on the target audio data, so as to filter out the environmental noise in the target audio data and obtain the noise-reduced target audio data, that is, the preprocessed target audio data.

[0154] Secondly, an explanation is given on how to obtain the first probability corresponding to the first audio frame:

[0155] After obtaining the first audio frame, feature extraction is performed on the first audio frame to obtain a feature vector corresponding to the first audio frame, and an audio signal classification model is used to determine the first probability corresponding to the first audio frame.

[0156] Among them, the feature extraction of the first audio frame specifically includes: performing a fast Fourier transform (FFT) on the first audio frame, extracting the spectral features of the first audio frame, and limiting the frequency range to 20 Hz - 5 kHz; then, based on the spectral features of the first audio frame, calculating the Mel-frequency cepstral coefficients (MFCC), and extracting the first 13-dimensional coefficients as feature vectors through a discrete cosine transform (DCT). The feature vector corresponding to the first audio frame is input into the audio signal classification model, and the first probability corresponding to the first audio frame, that is, the probability that the first audio frame is a valid audio frame, is output by the audio signal classification model.

[0157] In some embodiments, the audio signal classification model can be a pre-trained support vector machine (SVM) classification model. The pre-trained support vector machine (SVM) classification model is used to determine whether the sound is a structured music type (such as songs, soundtracks, extended segments, etc.). The audio signal classification model can also be a random forest model (RF), a lightweight neural network model (such as MOBILENET), etc. The audio signal classification model can also be a fusion model obtained by fusing at least two of the above-mentioned multiple models. It can be understood that when the audio signal classification model is a fusion model, the first probability corresponding to the first audio frame obtained through the audio signal classification model can be the probability output by each model in the multiple models weighted and summed according to different weights. For example, if the audio signal classification model is a fusion model of a pre-trained support vector machine classification model, a random forest model, and a lightweight neural network model, the weight corresponding to the pre-trained support vector machine classification model can be taken as 0.5, the weight corresponding to the random forest model can be 0.3, and the weight corresponding to the random forest model can be 0.2, so as to obtain the first probability corresponding to the first audio frame.

[0158] Finally, the first audio frame can be determined whether it is a valid audio frame by comparing the first probability corresponding to the first audio frame with a preset probability.

[0159] In some embodiments, the controller is further configured to: not send a song recognition request to the server when the target audio data does not meet the first preset condition.

[0160] When the target audio data does not meet the first preset condition, the target audio data does not have the condition for song recognition, or in other words, the target audio data does not include a valid audio segment (the first audio segment). Do not send a song recognition request to the server, prompt the user that "the current scene cannot recognize the song", and terminate the subsequent song recognition process to avoid wasting resources.

[0161] In a possible implementation, the controller sends a song recognition request to the server according to the first audio segment in the target audio data for song recognition, which specifically includes: generating an audio fingerprint corresponding to the first audio segment; and sending a song recognition request to the server according to the music fingerprint corresponding to the first audio segment for song recognition.

[0162] In some examples, the controller calculates the short-time energy of the first audio segment, extracts the significant peak points (Peak Points) corresponding to the first audio segment, and generates multiple hash fingerprints (Hash Fingerprint) (i.e., audio fingerprints) corresponding to the first audio segment according to the time and frequency pairs of the significant peak points corresponding to the first audio segment. For example: (t1, f1, t2, f2) → hash_code = 0x3A5B. Encapsulate multiple hash fingerprints into a JSON array. Example: {"fingerprints": ["0x3A5B", "0x7C2D",...]} and upload the JSON array corresponding to the first audio segment to the server, so that the server performs song recognition according to the JSON data.

[0163] It should be noted that when the server receives the song recognition request, it determines whether the multiple audio fingerprints in the JSON data meet the second preset condition. When the multiple audio fingerprints meet the second preset condition, it obtains the song recognition result corresponding to the audio fingerprint with the number of consecutive first audio fingerprints in the multiple audio fingerprints as the song recognition result corresponding to the first audio segment; where the second preset condition is that the number of consecutive first audio fingerprints in the multiple audio fingerprints is greater than or equal to the second threshold, and the first audio fingerprint is an audio fingerprint with a matching degree greater than the preset matching degree with the pre-stored feature fingerprint.

[0164] If the multiple audio fingerprints meet the second preset condition, it is determined that the audio fingerprint corresponding to the first audio segment is successfully matched, and the song information matched by the audio fingerprint corresponding to the first audio segment can be returned. The song information includes the song name.

[0165] If the multiple audio fingerprints do not meet the second preset condition, it is determined that the audio fingerprint corresponding to the first audio segment is failed to be matched, and the user is prompted that "the current scene cannot be recognized".

[0166] Exemplarily, the second threshold is set to 5, and the preset matching degree is set to 98%. Thus, the second preset condition is that the number of audio fingerprints with a matching degree greater than 98% with the pre-stored feature fingerprint in the multiple audio fingerprints is greater than 5.

[0167] In some examples, when the server receives the song recognition request, if the recognition of the first audio segment stops after the preset time expires. Exemplarily, the preset time can be set to 15S.

[0168] The following will be described in detail by Figure 9 the embodiments shown below. Figure 9 FIG. shows a schematic diagram of steps for determining whether to record audio corresponding to the current display interface in some embodiments of the present application. The method is applied to a first application and includes:

[0169] S901. In response to a target instruction, take a screenshot of the content displayed on the current display interface to obtain a target image;

[0170] S902. Obtain the identifier of the top-level APP in the currently running APP;

[0171] S903. Determine whether the identifier of the top-level APP in the currently running APP matches the identifier of the desktop APP;

[0172] If the identifier of the top-level APP in the currently running APP matches the identifier of the desktop APP successfully, execute S904;

[0173] If the identifier of the top-level APP in the currently running APP does not match the identifier of the desktop APP, execute S909;

[0174] S904. Obtain the attributes of the content displayed on the current display interface;

[0175] The attributes of the content displayed on the current display interface include: the state of the content displayed on the current display interface and the media type corresponding to the content displayed on the current display interface.

[0176] S905. Determine whether the state of the content displayed on the current display interface is a playing state;

[0177] If the state of the content on the current display interface is a playing state, execute S906;

[0178] If the state of the content on the current display interface is not a playing state, execute S907;

[0179] S906. Determine whether the media type corresponding to the content displayed on the current display interface belongs to a preset media type;

[0180] If the media type corresponding to the content displayed on the current display interface belongs to the preset media type, execute S907;

[0181] If the media type corresponding to the content displayed on the current display interface does not belong to the preset media type, execute S908;

[0182] S907. Do not record the audio corresponding to the current display interface;

[0183] S908. Record the audio corresponding to the current display interface;

[0184] S909. Obtain the scene corresponding to the target image;

[0185] S910. Determine whether the interface type corresponding to the target image belongs to a preset interface type;

[0186] If the interface type corresponding to the target image belongs to the preset interface type, execute S908;

[0187] If the interface type corresponding to the target image does not belong to the preset interface type, execute S907.

[0188] It should be noted that for the implementation processes of S901 to S9010, reference can be made to the detailed descriptions in the foregoing embodiments of this application, which will not be elaborated herein again. Moreover, after recording the audio corresponding to the current display interface in this embodiment, the steps of the picture recognition process and the track recognition process in the foregoing embodiments can also be executed, which will not be specifically elaborated herein.

[0189] It should be noted that the embodiments of this application further include: displaying the image recognition process and the track recognition process on the display device. Exemplarily, the image recognition result and the track recognition result can be displayed on the display device, or, when no image recognition result or track recognition result is obtained, user prompt information can be displayed.

[0190] Exemplarily, as Figure 10 shown, Figure 10 shows a user interface provided by the embodiments of this application. The information in the image recognition process and the information in the track recognition process are displayed in the interface control list. The interface control list includes a first display area and a second display area. The first display area is used to display the information in the image recognition process, and the second display area is used to display the information in the track recognition process. The first display area and the second display area do not overlap.

[0191] Specifically, displaying the information in the image recognition process in the first display area may include: during the image recognition process, displaying the prompt message "Recognizing the screenshot image, please wait patiently"; after obtaining the image recognition result, displaying the image recognition result, or if no recognition result is obtained after image recognition, displaying the prompt message "No image recognition result obtained".

[0192] Displaying information during track recognition in the second display area may include: during track recognition, displaying a prompt message "Recognizing, please wait" in the second display area; in the case of obtaining a track recognition result, the track recognition result may be displayed in the second display area; if the audio corresponding to the current interface content is not recorded, that is, the track recognition step is not executed and no track recognition result is obtained, the second display area may be hidden, that is, the second display area is not displayed in the interface control, or a prompt message "The current scene cannot be recognized" may be displayed in the second display area, or, in the case where the audio corresponding to the current display interface includes a first audio segment, that is, the audio corresponding to the current display interface does not include recognizable valid music, at this time, a prompt message "No recognizable music detected" may also be displayed in the second display area.

[0193] In some embodiments, the interface control is displayed as a floating layer on top of the current display interface, covering part or all of the display content of the current display interface. Generally, the interface control is arranged on the left side of the interface, and can also be at any position of the interface (such as: right side, top, bottom, etc.). The playback window of the current display interface displays the picture of the current playback content, which can be to continue playing the current playback content or to display a paused picture after pausing the current playback content. It can be understood that the interface control may also include a display area for displaying other content.

[0194] It should be understood that the magnitudes of the sequence numbers of the processes in the above embodiments do not mean the order of execution is prior or posterior. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. Each of the embodiments described herein can be an independent solution or can be combined according to internal logic, and these solutions all fall within the protection scope of this application.

[0195] It should also be understood that there is no strict order limit for the execution of each step in the flowcharts in the above embodiments, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0196] In the embodiments of the present application, the display device can be divided into functional modules according to the above method examples. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and other feasible division methods can be adopted in actual implementation. The following will take the example of dividing each functional module corresponding to each function for illustration.

[0197] Next, the device embodiments of the present application will be described in detail in conjunction with Figure 11 , It should be understood that the recognition device in the embodiments of the present application can execute various recognition methods of the foregoing embodiments of the present application, that is, the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.

[0198] Figure 11 is a schematic structural diagram of the recognition device provided by the embodiments of the present application. It should be understood that the recognition device can execute Figures 6 to 9 the recognition method shown; the recognition device 1000 includes: a receiving unit 1010, a determining unit 1020, and a recording unit 1030, where:

[0199] The receiving unit 1010 is configured to receive a target trigger signal sent by a remote control. The target trigger signal indicates that a target button on the remote control is pressed, and the target button is a button on the remote control for triggering track recognition;

[0200] The determining unit 1020 is configured to determine the state of the content displayed on the current display interface in response to the target trigger signal;

[0201] The recording unit 1030 is configured to record the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state, and send a track recognition request to the server according to the generated target audio data for track recognition; when the state of the content displayed on the current display interface is the non-playing state, the audio corresponding to the current display interface is not recorded.

[0202] Each unit module of the recognition device can execute the corresponding steps in the above method embodiments respectively, so the unit modules will not be elaborated here. For details, please refer to the description of the above corresponding steps.

[0203] It should be noted that the above recognition device is embodied in the form of functional units. The term "unit" here can be implemented in the form of software and / or hardware, and no specific limitation is made thereto.

[0204] For example, a "unit" may be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a combined logic circuit, and / or other suitable components that support the described functions.

[0205] Therefore, the units of the various examples described in the embodiments of the present application can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0206] As Figure 1 、 Figure 2 shown, the display device 200 can be used to implement the recognition method described in the above method embodiments.

[0207] The display device 200 may include one or more memories on which there is a program that can be run by the controller 250 to generate instructions, so that the controller 250 executes the recognition method described in the above method embodiments according to the instructions.

[0208] Optionally, data may also be stored in the memory. Optionally, the controller 250 may also read the data stored in the memory. The data may be stored at the same storage address as the program, or the data may be stored at a different storage address from the program.

[0209] The controller 250 and the memory may be provided separately or integrated together; for example, integrated on a system on chip (SOC) of a terminal device.

[0210] The present application also provides a computer program product that implements the recognition method of any method embodiment in the present application when executed by the controller 250.

[0211] The computer program product may be stored in the memory, for example, it is a program that is finally converted into an executable target file that can be executed by the controller 250 after processes such as preprocessing, compilation, assembly, and linking.

[0212] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the recognition method in any method embodiment of the present application is implemented. The computer program can be a high-level language program or an executable target program.

[0213] The computer-readable storage medium is, for example, a memory. The memory can be a volatile memory or a non-volatile memory, or the memory can include both a volatile memory and a non-volatile memory at the same time. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0214] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items" or its similar expression means any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0215] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0216] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0217] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0218] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0219] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0220] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0221] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application and should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A display device, characterized in that, Comprising: A display; A controller connected to the display, the controller being configured to: Receive a target trigger signal sent by a remote control, wherein the target trigger signal indicates that a target button on the remote control is pressed, and the target button is a button on the remote control for triggering track recognition; In response to the target trigger signal, determine the state of the content displayed on the current display interface; When the state of the content displayed on the current display interface is the playing state, record the audio corresponding to the current display interface, and send a track recognition request to the server based on the generated target audio data for track recognition; When the state of the content displayed on the current display interface is the non-playing state, do not record the audio corresponding to the current display interface.

2. The display device according to claim 1, characterized in that, After receiving the target trigger signal sent by the remote control, the controller is further configured to: In response to the target trigger signal, take a screenshot of the content displayed on the current display interface, and send an image recognition request to the server based on the obtained screenshot image for image recognition.

3. The display device according to claim 1, characterized in that The controller determining the state of the content displayed on the current display interface in response to the target trigger signal includes: The controller determines the state of the content displayed on the current display interface in response to the target trigger signal, and determines the media type corresponding to the content displayed on the current display interface; The controller performing the recording of the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state includes: The controller records the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface does not belong to a preset media type; The controller does not record the audio corresponding to the current display interface when the state of the content displayed on the current display interface is the playing state and the media type corresponding to the content displayed on the current display interface belongs to a preset media type.

4. The display device according to any one of claims 1 to 3, characterized in that: The application corresponding to the content displayed on the current display interface is the desktop application of the display device.

5. The display device according to claim 2, characterized in that, Before determining the state of the content displayed on the current display interface in response to the target trigger signal, the controller is further configured to: Determine whether the application corresponding to the content of the current display interface is a desktop application; When the application corresponding to the content of the current display interface is a desktop application, perform determining the state of the content displayed on the current display interface in response to the target trigger signal; When the application corresponding to the content of the current display interface is a non-desktop application, the sending of the image recognition request to the server includes: sending an image recognition request including a first parameter to the server, wherein the first parameter indicates that the application corresponding to the content of the current display interface is a non-desktop application; Receive the interface type corresponding to the captured screen image fed back by the server, where the interface type corresponding to the captured screen image is determined by the server based on the captured screen image; When the interface type corresponding to the captured screen image belongs to a preset interface type, record the audio corresponding to the currently displayed interface; When the interface type corresponding to the captured screen image does not belong to the preset interface type, do not record the audio corresponding to the currently displayed interface.

6. The display device according to claim 1, wherein The controller is configured to send a song recognition request to the server based on the target audio data generated by recording for song recognition, specifically configured as: When the target audio data meets a first preset condition, send a song recognition request to the server for song recognition based on a first audio segment in the target audio data, where the target audio data includes multiple audio frames; Wherein, the first preset condition is that the number of valid audio frames in the first audio segment is greater than or equal to a first threshold, and the first audio segment is a first preset number of consecutive audio frames in the target audio data; The controller is further configured to: When the target audio data does not meet the first preset condition, do not send a song recognition request to the server.

7. The display device according to claim 6, wherein The controller is further configured to: For the first audio frame in the target audio data, use an audio signal classification model to process the first audio frame to obtain a first probability, where the first probability is the probability that the first audio frame is a valid audio frame, and the first audio frame is any one of the multiple audio frames; When the first probability is greater than or equal to a preset probability, determine that the first audio frame is a valid audio frame.

8. A recognition method, characterized in that, Applied to the display device according to any one of claims 1 to 7, the method includes: Receive a target trigger signal, where the target trigger signal indicates that a target button on the remote control is pressed, and the target button is a button on the remote control for triggering song recognition; In response to the target trigger signal, determine the state of the content displayed on the current display interface; When the state of the content displayed on the current display interface is a playing state, record the audio corresponding to the current display interface, and send a song recognition request to the server based on the target audio data generated by recording for song recognition; When the state of the content displayed on the current display interface is a non-playing state, do not record the audio corresponding to the current display interface.

9. The method according to claim 8, characterized in that, The determining the state of the content displayed on the current display interface in response to the target trigger signal includes: In response to the target trigger signal, determine the state of the content displayed on the current display interface and determine the media type corresponding to the content displayed on the current display interface; The recording the audio corresponding to the current display interface when the state of the content displayed on the current display interface is a playing state specifically includes: When the state of the content displayed on the current display interface is a playing state and the media type corresponding to the content displayed on the current display interface does not belong to a preset media type, record the audio corresponding to the current display interface; When the state of the content displayed on the current display interface is the playing state, and the media type corresponding to the content displayed on the current display interface belongs to a preset media type, the audio corresponding to the current display interface is not recorded.

10. The method according to any one of claims 8 or 9, characterized in that: The application corresponding to the content displayed on the current display interface is the desktop application of the display device.