Display apparatus and audio recognition method
Patent Information
- Application Number
- CN202211635428.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2042-12-19
AI Technical Summary
因此,往往录制较长时间的音频才能得到较准确的识别结果,这导致需要用户长时间等待,导致用户体验感差
[0013]本申请一些实施例提供一种显示设备和音频识别方法,可响应于用户输入的音频识别请求,与正在播放音频的第一应用程序建立通信,从第一应用程序对应的媒资缓存中获取当前播放时刻之前已播放的历史音频数据,再通过对历史音频数据对应的待识别音频数据执行识别操作,确定待识别音频数据的识别结果,通过待识别的音频对应的历史音频数据进行识别,减少了用户等待录制的时间,提高了识别音频的效率,提高用户体验。
Smart Images

Figure CN117292681B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of display technology, and more particularly to a display device and an audio recognition method. Background Technology
[0002] With the development of technology, display devices are becoming increasingly diversified, offering users a wider range of functions. Display devices include smart TVs, smartphones, laser projectors, and other products with integrated display screens. They provide users with various entertainment functions such as video, audio, games, and karaoke, satisfying diverse entertainment needs.
[0003] In some display device applications, if a user is interested in the audio played by the display device, they can learn about the relevant information of the audio through audio recognition controls or other means. In audio recognition methods, after receiving a request from the user to recognize the audio, the display device begins recording audio data and performs audio recognition using that audio data.
[0004] In audio recognition methods, the longer the recorded audio data, the greater the likelihood of it being recognized. Therefore, recording longer audio sessions is often necessary to obtain more accurate recognition results, which leads to long waiting times for users and a poor user experience. Summary of the Invention
[0005] This application provides a display device and an audio recognition method, which can be used to solve the technical problem of long waiting times for users during the audio recognition process.
[0006] This application provides a method for recording status display and a display device, which can improve the user experience of operating the display device.
[0007] In a first aspect, some embodiments of this application provide a display device, including a display and a controller; the controller is communicatively connected to the display, and the controller is configured to:
[0008] In response to the user's input audio recognition request, communication is established with the first application that is playing audio, and historical audio data that has been played before the current playback time is obtained from the media asset cache corresponding to the first application;
[0009] By performing a recognition operation on the audio data to be recognized corresponding to the historical audio data, the recognition result of the audio data to be recognized is determined and displayed in the audio recognition display area of the monitor.
[0010] Some embodiments of this application provide an audio recognition method, characterized in that it includes:
[0011] In response to the user's input audio recognition request, communication is established with the first application that is playing audio, and historical audio data that has been played before the current playback time is obtained from the media asset cache corresponding to the first application;
[0012] By performing a recognition operation on the audio data to be recognized corresponding to the historical audio data, the recognition result of the audio data to be recognized is determined and displayed in the audio recognition display area of the monitor.
[0013] Some embodiments of this application provide a display device and an audio recognition method that, in response to an audio recognition request input by a user, establishes communication with a first application that is currently playing audio, obtains historical audio data that has been played before the current playback time from the media asset cache corresponding to the first application, and then determines the recognition result of the audio data to be recognized by performing a recognition operation on the audio data to be recognized corresponding to the historical audio data. By recognizing the audio data to be recognized through the historical audio data corresponding to the audio data to be recognized, the user's waiting time for recording is reduced, the efficiency of audio recognition is improved, and the user experience is enhanced. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This application illustrates operational scenarios between a display device and a control device according to some embodiments;
[0016] Figure 2 A hardware configuration block diagram of a control device 100 according to some embodiments of this application is shown;
[0017] Figure 3 A hardware configuration block diagram of a display device 200 according to some embodiments of this application is shown;
[0018] Figure 4 The present application illustrates a software configuration diagram in a display device according to some embodiments;
[0019] Figure 5 A timing diagram of audio recognition in a display device according to some embodiments of this application is shown;
[0020] Figure 6 The illustration shows a schematic diagram of a user's request for the display device to recognize audio in some embodiments of this application;
[0021] Figure 7aThis application shows a schematic diagram of data in a media asset cache according to some embodiments;
[0022] Figure 7b This illustration shows a schematic diagram of data in another media asset cache in some embodiments of this application;
[0023] Figure 7c This illustration shows a schematic diagram of data in another media asset cache in some embodiments of this application;
[0024] Figure 7d This illustration shows a schematic diagram of data in another media asset cache in some embodiments of this application;
[0025] Figure 7e This illustration shows a schematic diagram of data in another media asset cache in some embodiments of this application;
[0026] Figure 7f This illustration shows a schematic diagram of data in another media asset cache in some embodiments of this application;
[0027] Figure 8 A timing diagram of audio recognition in another display device according to some embodiments of this application is shown;
[0028] Figure 9 A timing diagram of audio recognition in another display device according to some embodiments of this application is shown;
[0029] Figure 10 The following is a timing diagram illustrating the audio recognition operation performed by the display device through an audio recognition server in some embodiments of this application. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the exemplary embodiments of this application clearer, the technical solutions in the exemplary embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0031] Based on the exemplary embodiments shown in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can constitute a complete technical solution on its own.
[0032] It should be understood that the terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate, for example, to allow implementation in orders other than those given in the embodiments illustrated or described in this application.
[0033] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.
[0034] The display device provided in this application can have various implementation forms, such as a television, smart television, mobile terminal, tablet computer, computer, laptop computer, laser projection device, monitor, electronic bulletin board, electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.
[0035] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.
[0036] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.
[0037] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.
[0038] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.
[0039] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.
[0040] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.
[0041] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0042] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0043] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.
[0044] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0045] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.
[0046] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the control device 100 or the server 400 through the communicator 220.
[0047] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).
[0048] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0049] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.
[0050] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.
[0051] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0052] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can execute operations related to the object selected by the user command.
[0053] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.
[0054] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0055] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0056] like Figure 4 In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android Runtime and System Library layer (referred to as the "System Runtime Layer"), and the kernel layer.
[0057] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs, system settings programs, or clock programs that come with the operating system; they may also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0058] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0059] like Figure 4 As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: ActivityManager, which interacts with all activities running in the system; LocationManager, which provides access to system location services for system services or applications; PackageManager, which retrieves various information related to application packages currently installed on the device; NotificationManager, which controls the display and clearing of notification messages; and WindowManager, which manages user interface elements including icons, windows, toolbars, wallpapers, and desktop widgets.
[0060] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling display window changes (e.g., shrinking the display window, shaking the display, distorting the display, etc.).
[0061] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0062] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0063] During the use of the display device, if the user is interested in the audio played by the display device, they can send a recognition request through an audio recognition control or other means. After receiving the user's request for audio recognition, the display device starts recording audio data and performs audio recognition using that audio data. In order to ensure the probability of recognition, a certain duration of audio data needs to be recorded, and the longer the recorded audio data, the greater the probability of it being recognized. Therefore, the user needs to wait for the corresponding duration of audio recording before obtaining a more accurate recognition result, which results in the user having to wait for a long time.
[0064] To reduce the recording time for users while waiting for recognition, some embodiments of this application provide a display device and an audio recognition method. In response to a user-inputted audio recognition request, the method establishes communication with a first application that is currently playing audio, retrieves historical audio data played before the current playback time from the media asset cache corresponding to the first application, and then determines the information of the audio data to be recognized by performing a recognition operation on the audio data to be recognized corresponding to the historical audio data. Recognition is then performed using the historical audio data corresponding to the audio data to be recognized, thereby reducing the user's recording time, improving the efficiency of audio recognition, and enhancing the user experience.
[0065] Figure 5 The following is a timing diagram illustrating audio recognition in a display device according to some embodiments of this application, such as... Figure 5 As shown, the display device includes a display and a controller, the controller being communicatively connected to the display, wherein the controller is configured to perform the following steps:
[0066] Based on the user's playback request, play media assets containing audio to the server;
[0067] Download and cache the audio and video data corresponding to the media assets parsed by the wall mount for playback.
[0068] S310, in response to an audio recognition request input by the user, establishes communication with the first application that is playing audio.
[0069] The user's request to recognize audio can be made by clicking the recognition control on the display device 200, or by pressing buttons or using voice input on the control device 100.
[0070] Figure 6 The following diagram illustrates a user's request for the display device to recognize audio in some embodiments of this application, such as... Figure 6 As shown, the display device has a recognition control 261. The user can operate the recognition control 261 through a control device or the like to trigger a request for audio recognition.
[0071] In response to a user's request to identify audio, communication is first established with the first application that is playing the audio, which facilitates communication with the media asset cache corresponding to the first application.
[0072] like Figure 5 As shown, the controller is configured to perform the following steps: S320, retrieve historical audio data that has been played before the current playback time from the media asset cache corresponding to the first application.
[0073] It should be understood that media asset caches can include both audio and video data.
[0074] Among them, historical audio data refers to audio data that the first application has played before the current playback time. When a user is interested in the audio being played while using the display device, the audio is in a playing state. When the user realizes that they want to learn about the relevant information of the audio through audio recognition controls or other means, the audio has already been playing for a period of time. Therefore, when the user issues an audio recognition request through audio recognition controls or other means, the corresponding historical audio data is already in the media asset cache of the first application.
[0075] Figure 7a This application illustrates a schematic diagram of data in a media asset cache in some embodiments, such as... Figure 7a As shown, the audio data in the media asset cache that is before the current playback time is the historical audio data that has already been played.
[0076] The duration of historical audio data retrieved from the media asset cache each time may be different.
[0077] For example, if the first application stores 8 seconds of played audio data in its media asset cache at the current playback moment, that is, it can obtain 8 seconds of historical audio data from the media asset cache immediately after the user clicks the audio recognition control; if it stores 5 seconds of played audio data in its media asset cache, that is, it can obtain 5 seconds of historical audio data from the media asset cache immediately after the user clicks the audio recognition control; and if it stores 2 seconds of played audio data in its media asset cache, that is, it can obtain 2 seconds of historical audio data from the media asset cache immediately after the user clicks the audio recognition control.
[0078] It should be understood that audio recognition can be performed on the audio data to be recognized, which includes the historical audio data obtained in step 320. For example... Figure 5 As shown, the controller is configured to perform the following steps:
[0079] S330. By performing a recognition operation on the audio data to be recognized corresponding to the historical audio data, the recognition result of the audio data to be recognized is determined.
[0080] The recognition operation is performed on the audio data to be recognized corresponding to the historical audio data. The recognition operation can be implemented on the display device or by sending the audio data to be recognized to the server and then relying on the return result from the audio recognition server.
[0081] It should be understood that there are many types of audio data to be identified. Generally, the types of data to be identified are more comprehensive and the identification efficiency is faster when the identification is performed on the server side. In some embodiments, the identification of audio data to be identified can also be achieved through data on the local display device.
[0082] The capacity of the media asset cache storing historical audio data is limited, and the length of historical audio data may be long or short. Therefore, the length of the audio data to be recognized corresponding to the historical audio data is also limited. When performing audio recognition through the audio data to be recognized, the recognition result of the recognized audio data may be recognized information or it may not be recognized. In the case of not being recognized, the length of the audio data can be increased before further recognition.
[0083] In some embodiments, since the duration of historical audio data obtained from the media asset cache is uncertain, the historical audio data can be further judged before the recognition operation is performed to determine whether the duration and other information of the historical audio data meet the preset recognition conditions.
[0084] For example, by using methods such as statistical analysis of audio data, the minimum duration of the audio data to be identified in audio recognition is determined to be 5 seconds. In this case, if the obtained historical audio data is greater than or equal to 5 seconds, the audio data to be identified corresponding to the historical audio data can be recognized. If the obtained historical audio data is less than 5 seconds, the corresponding audio data to be identified cannot obtain the corresponding recognition result. In this case, the duration of the audio data to be identified can be increased by recording audio data after the current playback time on the display device, or by obtaining the last audio data downloaded and played after the current playback time from the media asset cache, etc., to further realize the recognition of historical audio data less than 5 seconds.
[0085] By adding a function to determine whether historical audio data meets preset recognition conditions, the efficiency of audio recognition can be improved, and the probability of audio data being recognized can also be increased.
[0086] S340: Display the recognition result of the audio data to be recognized in the audio recognition display area of the monitor.
[0087] For the display of the recognition results of the identified audio data, if the recognition result is information of the audio data to be recognized, it can be displayed directly in the result display area, or indirectly in the result display area through a QR code or other means; in some embodiments, it can also be displayed through voice information or other means.
[0088] If the recognition result is unrecognized, information such as recognition failure can be displayed in the audio recognition display area, or the user can be prompted to perform further recognition.
[0089] This application embodiment reduces the user's waiting time for recording, improves the efficiency of audio recognition, and enhances the user experience by recognizing historical audio data corresponding to the audio data to be recognized.
[0090] It should be understood that the duration of historical audio data is uncertain. After retrieving historical audio data played before the current playback moment from the media asset cache corresponding to the first application in step 320, it may also include recording the audio data currently being played by the first application. Figure 8 The following is a timing diagram illustrating audio recognition in a display device according to some embodiments of this application, such as... Figure 8 As shown, the controller in the display device is also configured to perform the following steps:
[0091] S410, in response to an audio recognition request input by the user, establishes communication with the first application that is playing audio.
[0092] S420: Obtain historical audio data that has been played before the current playback time from the media asset cache corresponding to the first application.
[0093] Historical audio data refers to audio data that the first application has played before the current playback time. When a user is interested in the audio being played while using the display device, the audio is in a playing state. When the user realizes that they want to learn about the relevant information of the audio through audio recognition controls or other means, the audio has already been playing for a period of time. Therefore, when the user issues an audio recognition request through audio recognition controls or other means, the historical audio data corresponding to the audio is already in the media asset cache of the first application.
[0094] In some embodiments, the audio data in the media asset cache is stored in segments and is continuously updated in segments. The media asset cache may contain the first segment data that has been played and the second segment data that is currently being played, or it may only contain the second segment data that is currently being played.
[0095] The lengths of the first data segment and the second data segment can be the same or different.
[0096] For example, media asset data in a display device is stored in a media asset cache as TS (Transport Stream, an audio and video encapsulation format) fragments, and is stored in the form of an audio stream.
[0097] Figure 7b This application illustrates a schematic diagram of data in another media asset cache, as shown in some embodiments. Figure 7b As shown, when the media asset cache contains the first segment data that has already been played and the second segment data that is currently being played, the media asset cache contains historical audio data that has been played before the current playback time, and this historical audio data includes the first sub-segment data that was played before the current playback time in the second segment data that is currently being played, and the first segment data that was played before the second segment data.
[0098] The second segment data may include the first sub-segment data that has been played before the current playback time and the second sub-segment data that has not been played after the current playback time.
[0099] In some embodiments, the second segment data may also include first sub-segment data that has been played before the current playback time, or the second segment data may include second sub-segment data that has not been played after the current playback time.
[0100] Figure 7c This application illustrates a schematic diagram of data in another media asset cache, as shown in some embodiments. Figure 7c As shown, when the media asset cache contains only the currently playing second segment data, and the second segment data includes the first sub-segment data that was played before the current playback time, the media asset cache contains historical audio data that was played before the current playback time, and this historical audio data includes the first sub-segment data that was played before the current playback time in the currently playing second segment data.
[0101] Data in media asset cache, such as Figure 7a As shown, step 430 can be executed simultaneously with or after step 420:
[0102] like Figure 8 As shown, the controller is also configured to perform the following steps: S430, recording audio data played by the first application within a first preset time period after receiving the audio recognition request.
[0103] It should be understood that during the audio recognition process, historical audio data can be obtained through media asset caching. At this time, the audio data played by the first application is recorded. This audio data is the data played by the first application within a first preset time period after the audio recognition request. The recorded audio data is combined with the historical audio data to increase the length of the audio data to be recognized and improve the recognition efficiency.
[0104] In existing audio recognition, the display device only starts recording audio data after receiving a user's request to input audio for recognition. This application embodiment increases the length of the audio data to be recognized by recording audio data based on historical audio data (i.e., a portion of audio data is already available). The duration of the recorded audio data is controlled to be less than the duration of audio data recording in existing audio recognition, thereby reducing the waiting time required for the user to record audio data during the audio recognition process.
[0105] For example, in existing audio recognition, the display device only starts recording audio data after receiving a user's request to input the audio for recognition. If the recording time is 15 seconds, the user needs to wait 15 seconds before the audio recognition function can be performed. However, in the embodiment of this application, after receiving a user's request to input the audio for recognition, the display device retrieves historical audio data from the media asset cache. If the duration of the historical audio data is 10 seconds, the corresponding first preset time may be 5 seconds. In other words, the user only needs to wait 5 seconds for the recording to be performed before the audio recognition function can be performed, reducing the waiting time required for the user to record audio data during the audio recognition process.
[0106] The first preset time is the minimum duration that is less than the duration required for audio recognition and recording.
[0107] In some embodiments, a determination of historical audio data can be added before step 430. Since the duration of historical audio data obtained from the media asset cache is uncertain, it is possible to further determine whether information such as the duration of historical audio data meets preset recognition conditions.
[0108] For example, by statistical analysis of audio data, the minimum duration of the audio data to be identified in audio recognition is determined to be 10 seconds. If the obtained historical audio data is greater than or equal to 10 seconds, the recognition operation can be performed on the audio data to be identified corresponding to the historical audio data in step 330. If the obtained historical audio data is less than 10 seconds, the corresponding audio data to be identified cannot obtain the corresponding recognition result. In this case, the duration of the audio data to be identified can be increased by recording the audio data after the current playback time on the display device in step 430, thereby further realizing the recognition of the corresponding audio data (i.e., recording data and historical audio data).
[0109] By adding a function to determine whether historical audio data meets preset recognition conditions, the efficiency of audio recognition can be improved, and the probability of audio data being recognized can also be increased.
[0110] It should be understood that when the duration of the audio data to be identified is insufficient for audio recognition, step 430 can be added to further expand the audio data to be identified. That is, step 430 records the audio data played by the first application within a first preset time period after receiving the audio recognition request. When the duration of the audio data to be identified meets the requirements of audio recognition or meets the preset recognition conditions, step 430 does not need to be executed.
[0111] like Figure 7b As shown, if the historical audio data includes the first sub-segment data that was played before the current playback time in the second segment data that is currently playing, step 430 can be configured to obtain the first alternative segment data that is played by the first application within a second preset time after the current playback time of the historical audio data, and the recording data includes the first alternative segment data.
[0112] In step 430, the first preset time period can be a preset fixed time length or a dynamically adjustable time length.
[0113] In some embodiments, the duration of recording audio data for the first application can be determined by comparing the duration of the first sub-segment data with the preset recognition duration, which is the minimum time for audio data to be recognized.
[0114] If the duration of the first sub-segment data is less than the preset recognition duration, the first playback duration (i.e., the first preset time period in step 430) is determined. The first playback duration is the difference between the preset recognition duration and the duration of the first sub-segment data.
[0115] After obtaining the current playback time of the historical audio data, the first alternative segment data is played by the first application within the first playback duration, and the recording data includes the first alternative segment data.
[0116] If the duration of the first sub-segment data is not less than the preset recognition duration, recognition can be performed using the audio data to be recognized corresponding to the first sub-segment data.
[0117] like Figure 7c As shown, if the historical audio data includes the first sub-segment data that was played before the current playback time and the first segment data that was played before the second segment data, step 430 can be configured to obtain the second alternative segment data that is played by the first application within a third preset time after the current playback time of the historical audio data, and the recording data includes the second alternative segment data.
[0118] In some embodiments, a preset recognition duration can be set, which is the minimum time for audio data to be recognized. The playback data duration is compared with the preset recognition duration, wherein the playback data duration is the sum of the duration of the first sub-segment data and the duration of the first segment data.
[0119] If the duration of the played data is less than the preset recognition duration, the second playback duration (i.e., the first preset time period in step 430) is determined. The second playback duration is the difference between the preset recognition duration and the duration of the played data.
[0120] After acquiring the current playback time of the historical audio data, the first application plays the second alternative segment data within the second playback duration, and the recording data includes the second alternative segment data.
[0121] If the duration of the played data is not less than the preset recognition duration, recognition can be performed using the first sub-segment data and the corresponding audio data to be recognized.
[0122] S440. By performing a recognition operation on the audio data to be recognized, the recognition result of the audio data to be recognized is determined, wherein the audio data to be recognized includes historical audio data and recording data.
[0123] Among them, if historical audio data is like Figure 7b As shown, the audio data to be identified includes the first segment data, the first sub-segment data, and the recording data; if the historical audio data is as follows... Figure 7c As shown, the audio data to be identified includes the first sub-segment data and the recording data.
[0124] Regardless of whether the audio data to be identified includes historical audio data, or both historical audio data and recorded audio data, the identification operation can be performed on the display device or by sending the audio data to be identified to the server and then relying on the results returned by the audio server.
[0125] It should be understood that there are many types of audio data to be identified. Generally, identification is performed on the server side. In some embodiments, the identification of audio data to be identified can also be performed through data on the local display device.
[0126] S450: Display the recognition results of the audio data to be recognized in the audio recognition display area of the monitor.
[0127] For the display of the recognition results of the identified audio data, if the recognition result is information of the audio data to be recognized, it can be displayed directly in the result display area, or indirectly in the result display area through a QR code or other means; in some embodiments, it can also be displayed through voice information or other means.
[0128] This application embodiment reduces user waiting time for recording, improves audio recognition efficiency, and enhances user experience by recognizing historical audio data or historical audio data and recording data corresponding to the audio data to be recognized.
[0129] It should be understood that the media asset cache may include not only played historical audio data, but also downloaded but unplayed end audio data. The historical audio data played before the current playback moment and the end audio data downloaded but not yet played after the current playback moment can be obtained from the media asset cache corresponding to the first application. Figure 9 The following is a timing diagram illustrating audio recognition in a display device according to some embodiments of this application, such as... Figure 9 As shown, the controller in the display device is also configured to perform the following steps:
[0130] S510, in response to an audio recognition request input by the user, establishes communication with the first application that is playing audio.
[0131] S520: Determine whether there is end audio data in the media asset cache corresponding to the first application.
[0132] It should be understood that during the use of the first application, its media asset cache may include historical audio data that has been played, as well as downloaded but unplayed audio data.
[0133] In some embodiments, the end audio data may be determined by unplayed data parsed from the media asset cache, or by unplayed data that has already been parsed from the media asset cache.
[0134] In some embodiments, the final audio data may be determined by converting the parsed digital audio data.
[0135] If historical audio data and end audio data exist in the media asset cache, the historical audio data and end audio data can be obtained, the corresponding audio data to be identified can be determined, the length of the audio data to be identified can be increased, and the probability of the audio data to be identified can be increased.
[0136] S530: Retrieve historical audio data played before the current playback time and the last audio data played after the current playback time from the media asset cache.
[0137] The end audio data consists of data that has been downloaded but not yet played since the current playback time.
[0138] Figure 7d This application illustrates a schematic diagram of data in another media asset cache, as shown in some embodiments. Figure 7dAs shown, when the media asset cache contains historical audio data that has already been played and downloaded but not yet played audio data, the last audio data is the audio data that has been downloaded but not yet played after the current playback time.
[0139] Figure 7e This application illustrates a schematic diagram of data in another media asset cache, as shown in some embodiments. Figure 7e As shown, when the media asset cache contains historical audio data that has been played and downloaded but not played end audio data, the end audio data is the audio data that has been downloaded but not played after the current playback time, and the end audio data includes the second sub-segment data that has been downloaded but not played after the current playback time in the second segment data that is currently being played.
[0140] Figure 7f This application illustrates a schematic diagram of data in another media asset cache, as shown in some embodiments. Figure 7f As shown, when the media asset cache contains historical audio data that has been played and downloaded but not played audio data at the end, the audio data at the end is the audio data that has been downloaded but not played after the current playback time, and the audio data at the end includes the second sub-segment data that has been downloaded but not played after the current playback time in the second segment data that is currently being played, and the third segment data that has been downloaded but not played after the second segment data.
[0141] It should be understood that, for Figure 7d , Figure 7e and Figure 7f The historical audio data can include the first sub-segment data that was played before the current playback time in the second segment data that is currently playing, or the historical audio data can also include the first segment data and the first sub-segment data that were played before the second segment data.
[0142] In other words, the historical audio data and the last audio data in the media asset cache can be... Figure 7b , Figure 7c Historical audio data can be compared with Figure 7e and Figure 7f The audio data at the end of the file can be combined in any two pairs.
[0143] S540. By performing a recognition operation on the audio data to be recognized, the recognition result of the audio data to be recognized is determined, wherein the audio data to be recognized includes historical audio data and end audio data.
[0144] Among them, if the audio data at the end is like Figure 7e As shown, the audio data to be identified includes historical audio data and second sub-segment data; if the historical audio data is as follows... Figure 7c As shown, the audio data to be identified includes historical audio data, second sub-segment data, and third segment data.
[0145] For example, in the media asset cache, the historical audio data is 8 seconds long and the last audio data is 8 seconds long. At this time, the corresponding audio data to be recognized is 16 seconds long. In response to the user's input audio recognition request, the historical audio data and the last audio data in the media asset cache are obtained, totaling 16 seconds. Audio recognition can be performed using the audio data to be recognized corresponding to the historical audio data and the last audio data, so that the user does not need to wait for the audio recording time, thus improving the speed of audio recognition and enhancing the user experience.
[0146] Regardless of whether the audio data to be identified includes historical audio data, or whether the audio data to be identified includes both historical audio data and the last audio data, the recognition operation can be implemented on the display device or by sending the audio data to be identified to the server and then relying on the return result from the audio recognition server.
[0147] It should be understood that there are many types of audio data to be identified. Generally, identification is performed on the server side. In some embodiments, the identification of audio data to be identified can also be performed through data on the local display device.
[0148] S550 displays the recognition results of the audio data to be recognized in the audio recognition display area of the monitor.
[0149] For the display of the recognition results of the identified audio data, if the recognition result is information of the audio data to be recognized, it can be displayed directly in the result display area, or indirectly in the result display area through a QR code or other means; in some embodiments, it can also be displayed through voice information or other means.
[0150] Figure 10 The following are timing diagrams illustrating the audio recognition operation performed by the display device through an audio recognition server in some embodiments of this application, such as... Figure 10 As shown, the audio data to be recognized can be determined based on different audio data, specifically including the following steps:
[0151] S610. Determine the audio data to be recognized.
[0152] based on Figures 5 to 8 The timing diagram should be understood to include the audio data to be identified, which includes: first audio data, second audio data, third audio data, or first audio data and third audio data.
[0153] S620: Transmit the audio data to be recognized to the audio recognition server.
[0154] On the audio recognition server side, the audio data to be recognized is recognized.
[0155] In some embodiments, the type of the audio data to be identified can be reduced to a preset data type on the audio recognition server. The reduced type results in a smaller data volume of the audio data to be identified, a faster transmission speed to the audio recognition server, and an improved recognition efficiency by the audio recognition server. Therefore, before step 520, the type of the audio data to be identified can be reduced to a preset data type on the audio recognition server.
[0156] S630 receives the recognition result sent by the audio recognition server and displays the recognition result on the display device.
[0157] The information (recognition result) of the identified audio data can be displayed directly in the result display area, or indirectly in the result display area through a QR code or other means; in some embodiments, it can also be displayed through voice information or other means.
[0158] It should be understood that the types of audio data to be identified are numerous, and generally, through... Figure 10 The recognition shown is implemented on the server side. In some embodiments, the recognition of the audio data to be recognized can also be achieved through data on the local display device.
[0159] This application embodiment reduces user waiting time for recording, improves audio recognition efficiency, and enhances user experience by recognizing historical audio data corresponding to the audio data to be recognized, or historical audio data and the last audio data.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0161] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, include: monitor; The controller, which is communicatively connected to the display, is configured to: In response to an audio recognition request input by the user, communication is established with the first application that is playing audio, and historical audio data that has been played before the current playback time is obtained from the media asset cache corresponding to the first application; Determine whether the duration of the historical audio data meets the preset recognition conditions, wherein the preset recognition conditions are the shortest duration required to perform the recognition operation; If the duration of the historical audio data is greater than or equal to the minimum duration required to perform the recognition operation, then the recognition result of the audio data to be recognized is determined by performing the recognition operation on the audio data to be recognized corresponding to the historical audio data, and the recognition result is displayed in the audio recognition display area of the display. If the duration of the historical audio data is less than the minimum duration required to perform the recognition operation, then the recording data played by the first application within a first preset time period after receiving the audio recognition request is obtained, or the last audio data downloaded but not played after the current playback time is obtained from the media asset cache; the recognition operation is performed on the audio data to be recognized; wherein, the audio data to be recognized includes the historical audio data and the recording data, or the audio data to be recognized includes the historical audio data and the last audio data.
2. The display device according to claim 1, characterized in that, If the historical audio data includes a first sub-segment of data that was played before the current playback time in the second segment of data that is currently being played, then in the step of obtaining the recording data played by the first application within a first preset time period after receiving the audio recognition request, the controller is configured to: After the current playback time of the historical audio data is obtained, the first alternative segment data is played by the first application within a second preset time period, and the recording data includes the first alternative segment data; or, Compare the duration of the first sub-segment data with the preset recognition duration; If the duration of the first sub-segment data is less than the preset recognition duration, a first playback duration is determined, and the first playback duration is the difference between the preset recognition duration and the duration of the first sub-segment data. After the current playback time of the historical audio data is obtained, the first alternative segment data is played by the first application within the first playback duration, and the recording data includes the first alternative segment data.
3. The display device according to claim 1, characterized in that, If the historical audio data includes the first sub-segment data played before the current playback time in the currently playing second segment data, and the first segment data played before the second segment data, in the step of obtaining the recording data played by the first application within a first preset time period after receiving the audio recognition request, the controller is configured to: After the current playback time of the historical audio data is obtained, the first application plays the second alternative segment data within a third preset time period, and the recording data includes the second alternative segment data; or, Compare the duration of the played data with the preset identification duration, wherein the duration of the played data is the sum of the duration of the first sub-segment data and the duration of the first segment data; If the duration of the played data is less than the preset identification duration, a second playback duration is determined, and the second playback duration is the difference between the preset identification duration and the duration of the played data. After the current playback time of the historical audio data, the first application plays a second alternative segment data within the second playback duration, and the recording data includes the second alternative segment data.
4. The display device according to claim 1, characterized in that, If the historical audio data includes the first sub-segment data that was played before the current playback time in the second segment data that is currently being played, or the historical audio data includes the first segment data and the first sub-segment data that were played before the second segment data; and the media asset cache also includes the last audio data that was downloaded but not played after the current playback time, the controller is further configured to: Obtain the end audio data; A recognition operation is performed on the audio data to be recognized, wherein the audio data to be recognized includes the historical audio data and the last audio data.
5. The display device according to claim 4, characterized in that, If the final audio data includes the second sub-segment data that has been downloaded but not played after the current playback time in the second segment data, then in the step of performing the recognition operation on the audio data to be recognized, the audio data to be recognized includes the historical audio data and the second sub-segment data.
6. The display device according to claim 4, characterized in that, If the final audio data includes the second sub-segment data that has been downloaded but not played after the current playback time and the third segment data that has been downloaded but not played after the second segment data, then in the step of performing the identification operation on the audio data to be identified, the audio data to be identified includes the historical audio data, the second sub-segment data and the third segment data.
7. The display device according to claim 1, characterized in that, In the step of determining the recognition result of the audio data to be recognized by performing a recognition operation on the audio data to be recognized corresponding to the historical audio data, the controller is configured as follows: The audio data to be identified is sent to the audio recognition server; Receive the recognition result of the audio data to be recognized returned by the audio recognition server.
8. An audio recognition method, characterized in that, include: In response to an audio recognition request input by the user, communication is established with the first application that is playing audio, and historical audio data that has been played before the current playback time is obtained from the media asset cache corresponding to the first application; Determine whether the duration of the historical audio data meets the preset recognition conditions, wherein the preset recognition conditions are the shortest duration required to perform the recognition operation; If the duration of the historical audio data is greater than or equal to the minimum duration required to perform the recognition operation, then the recognition result of the audio data to be recognized is determined by performing the recognition operation on the audio data to be recognized corresponding to the historical audio data, and the recognition result is displayed in the audio recognition display area of the display. If the duration of the historical audio data is less than the minimum duration required to perform the recognition operation, then the recording data played by the first application within a first preset time period after receiving the audio recognition request is obtained, or the last audio data downloaded but not played after the current playback time is obtained from the media asset cache; the recognition operation is performed on the audio data to be recognized; wherein, the audio data to be recognized includes the historical audio data and the recording data, or the audio data to be recognized includes the historical audio data and the last audio data.
Citation Information
Patent Citations
Audio recognition method and device
CN105657535A
Audio recognition method and device, terminal, earphone and readable storage medium
CN108922537A