Systems and methods for speech-based triggers for supplementing content

CN122580882APending Publication Date: 2026-08-14VIZIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480075951.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-25
Filing Date
2024-10-23
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因此,当呈现非交互式内容时,智能电视可能无法利用智能电视可用的交互式功能

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122580882A_ABST
    Figure CN122580882A_ABST
Patent Text Reader

Abstract

Media devices can perform contextual processing based on the presented media segments. A media device can receive identification of a video segment from an automatic content identification service, which is being displayed by a display device. The media device can send a notification including information associated with the video segment and a request for audio input. Upon detecting one or more audio segments associated with the notification, the media device can facilitate the presentation of objects associated with the video segment.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This patent application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 593,132, filed October 25, 2023, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0002] This disclosure generally relates to audio processing, and more particularly to context-controlled audio feedback processing performed by a computing device. Background Technology

[0003] Smart TVs can present interactive media (e.g., interactive videos, games, etc.) using various input devices (e.g., remote controls, mobile devices, etc.). Interactive media may include one or more functions configured to execute upon detecting an input event. The smart TV can detect the event (e.g., via an input / output interface, communication interface, etc.) and pass a notification of the input event to the interactive media being presented, causing the interactive media to execute the one or more functions. Non-interactive media (e.g., broadcast television, video, etc.) may not include instructions for performing specific or particular functions of the smart TV. Therefore, when presenting non-interactive content, the smart TV may not be able to utilize the interactive functions available to it. Summary of the Invention

[0004] This document describes methods and systems for contextual audio processing. The methods may include: receiving an identification of a video segment from an automatic content identification service, wherein the video segment is being displayed by a display device; sending a notification to the display device based on the identification of the video segment, the notification including information associated with the video segment and a request for audio input; detecting one or more audio segments associated with the notification; and facilitating the presentation of objects associated with the video segment in response to the detection of the one or more audio segments.

[0005] This document describes a system for contextual audio processing. The system may include one or more processors and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods previously described.

[0006] The non-transitory computer-readable medium described herein can store instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods previously described.

[0007] These illustrative examples are mentioned not to limit or define this disclosure, but to aid in its understanding. Further embodiments are discussed in the detailed description, and are further described therein. Attached Figure Description

[0008] The features, implementation methods and advantages of this disclosure will be better understood when the following detailed description is read in conjunction with the accompanying drawings.

[0009] Figure 1 A block diagram of an example computing device configured to perform contextual audio processing according to various aspects of this disclosure is shown.

[0010] Figure 2 This is a block diagram of an example system for contextual audio processing according to various aspects of this disclosure.

[0011] Figure 3 A diagram illustrating an example process for performing contextual audio processing in a computing device according to various aspects of this disclosure is shown.

[0012] Figure 4 A flowchart illustrating an example process for performing contextual audio processing in a computing device, according to various aspects of this disclosure, is shown.

[0013] Figure 5 An example computing device architecture is shown that can implement various technologies described herein according to various aspects of this disclosure. Detailed Implementation

[0014] This paper describes systems and methods for voice-based triggers for supplementing content. Media devices can be configured to present various types of media from diverse sources, such as over-the-air (OTA), cable, satellite, and the internet. Most media types include non-interactive media, which includes audiovisual data but may not include instructions for performing other functions of the media device. Therefore, many functions of the media device may be inaccessible, impacting the user experience. The methods and systems described herein provide instruction transmission in parallel with the media presented by the media device, which, when executed, causes the media device to perform functions in association with the presented media. Any non-interactive media can be transformed into interactive media without modifying the media being received by the media device. Furthermore, the functions of the media device can be conditionally performed in association with either non-interactive or interactive media.

[0015] For example, some content may include contact information (e.g., phone number, quick response (QR) code, etc.) that provides a user with a way to access additional information associated with the content. Because the contact information may be presented for a limited amount of time, it may be difficult for the user to perform actions related to it (e.g., remembering the phone number, taking a picture of the QR code, etc.). Instead, the media device may receive instructions to be executed during the presentation of the content, starting when the content begins to be presented and potentially extending to the end of the presentation of the specialized content or a period of time after the presentation of the specialized content, monitoring the media device's input source. Upon receiving input, the media device may execute additional instructions associated with the content. These instructions may cause the media device to present additional information associated with the specialized content, replace the specialized content with alternative specialized content, send user-customized communications associated with the specialized content, etc. For example, the content may provide information about a product, and when input is received within the time interval, a notification may be sent to the user (e.g., to a media device, the user's mobile device, or other devices associated with the user), the notification having additional information about the product, coupons for the product, additional media associated with the content, combinations thereof, etc.

[0016] Media devices can be configured to activate triggers that execute instructions when a specific event is detected. These triggers can be associated with specific media or media segments to enable the detection of events associated with said specific media or media segments and the execution of context-dependent processes. For example, a trigger can be associated with a television program or advertisement, such that when the television program or advertisement is presented by the media device, the trigger causes the media device to begin monitoring for a corresponding event. Events can be actions performed by a user, specific input received by the media device, execution of a specific function of the media device, execution of instructions, presentation of specific media or media segments, or combinations thereof. For example, an event could correspond to detecting audio input from the media device's microphone interface, where the audio input includes specific words or phrases spoken by the user of the media device.

[0017] In some cases, a media device may define one or more triggers (e.g., based on a user's historical viewing behavior, user input, web browsing activity, etc.). In other cases, a media device may receive communications defining one or more triggers. Triggers may include identification of a new trigger (e.g., an identifier, etc.), identification of an event (e.g., specific input from a specific input source, etc.), identification of a set of instructions to be executed in response to detecting an event, identification of a set of instructions to be executed for detecting specific media and / or detecting an event, identification of the time interval within which detecting the occurrence of an event will cause the instructions to be executed, identification of trigger media (e.g., specific media or specific media segments that cause the media device to begin monitoring for the occurrence of an event, etc.), identification of supplementary media segments to replace specific media or specific media segments, combinations thereof, etc. If trigger information is missing, the media device may request the missing information from a remote device and / or from the user.

[0018] In some cases, triggers may include instructions that instruct a media device to request additional information, supplementary media segments, and / or instructions from a remote device (e.g., a server, a content delivery network (CDN), other devices, etc.). Triggers may identify instructions, media or media segments, data, etc., that the media device retrieves to enable the trigger to be executed. For example, a trigger may identify one or more functions to be performed in response to the detection of audio input from a microphone interface and instruct the media device to retrieve instructions for performing said one or more functions. The media device may request said instructions, and when audio input is detected, execute said instructions to cause said one or more functions of the media device to be performed. For another example, a trigger may identify supplementary media segments to replace a media segment when it is displayed by the media device. The media device may request supplementary media segments from a remote device. When the media device detects the presentation of a media segment (i.e., an event), the media device may replace the presentation of that media segment with the supplementary media segments. Because instructions can be received from a different source than the trigger, triggers can be defined for different categories of media devices (e.g., different device types, etc.). Media devices can be configured to retrieve instructions (e.g., executable instructions, application programming interfaces, etc.), media, etc., to achieve the function of triggers.

[0019] Once a trigger is defined, a media device can begin monitoring the media segment being rendered by the media device for the specific media segment associated with the trigger (i.e., the trigger media). In some cases, the media device can use metadata embedded in the media to detect the rendering of the trigger media. Metadata embedded in the media segment may include the identification of the media segment, the identification of one or more media segments to be rendered in the future, the schedule of the media segment, etc.

[0020] In other cases, the media may not include metadata identifying the media segment being presented. Instead, the media device can use an Automatic Content Identification (ACR) service to identify the media segment being presented. For example, the media device can generate one or more unknown cues by sampling the video and / or audio channels of the media. The media device can generate cues from pixel values ​​of one or more consecutive sets of pixels in a video frame. Alternatively or additionally, unknown cues may include representations of audio segments extracted from the media segment (e.g., analog or digital signals, word sets from a speech-to-text model, etc.). The media device can then compare the unknown cues with known cues associated with known media segments stored in a reference database to identify the currently presented media segment, channel, media source, and / or the like.

[0021] In some examples, the reference database may be stored in the media device's local memory, enabling the media device to locally identify unknown cues. The media device may receive updates to the reference database from an ACR server (or other remote device) to enable it to identify new media segments. If the media device cannot identify a matching known cue in the reference database, it may send the cue to the ACR server for identification. The ACR server may identify the known cues corresponding to the unknown cues and send the identification of the media segments associated with the known cues to the media device. In other examples, the reference database may be stored in remote memory (e.g., the memory of an ACR server or other remote device). In these examples, the media device may send cues to the ACR server for identification. The ACR server may identify the known cues corresponding to the unknown cues in the reference database. The ACR server may then send the identification of the media segments associated with the known cues to the media device.

[0022] Once a media segment is identified, the media device can determine whether the identified media segment corresponds to the trigger media of a new trigger. If the identified media segment does not correspond to the trigger media of a new trigger, the media device can wait until the media segment terminates and the new media segment begins. The media device can then identify the new media segment. If the identified media segment does correspond to the trigger media of a new trigger, the media device can instantiate one or more event listeners associated with that trigger to monitor events.

[0023] An event can correspond to a specific input received through a specific interface. A specific interface can include, but is not limited to, a microphone, an optical sensor, an input device (e.g., a mouse, keyboard, remote control, or other device configured to remotely control the operation of a media device, a game controller, and / or the like), a network interface of the media device (e.g., an interface using Bluetooth, Ethernet, Wi-Fi, Zigbee, Z-Wave, etc.), or combinations thereof. One or more event listeners can generate an event upon detecting a specific input from a specific interface. The input can be an audio segment (e.g., a spoken word or phrase), an alphanumeric string (e.g., a text message, email, input from a remote control, etc.), a graphic (e.g., a QR code, image, symbol, etc.), light or a sequence of light (e.g., a camera flash, etc.), and / or the like. The event generated by the event listener can include identification of the interface receiving the input, the received input, a timestamp corresponding to the time the event was received by the media device, or a combination thereof. An event listener can be terminated once the time interval associated with a new trigger expires.

[0024] Media devices can perform different processes based on the generated events. For example, if the event corresponds to audio input from a microphone, the media device can send communications to a device associated with the user of the media device, send communications to a user profile, send communications to a device associated with the identified media segment corresponding to the event, etc. If the event corresponds to input from a network interface, the media device can be configured to present supplementary media segments, restart the currently presented media segment, present other media segments, present information associated with the identified media segment, send instructions that can be executed by the device sending the input (e.g., to present media, navigate to a webpage, subscribe to push notifications associated with the media segment, download an application or service, send communications to a remote device, etc.), or combinations thereof.

[0025] Media devices can also perform different processes based on received input corresponding to an event. For example, a media device can present a media clip associated with a product. The media clip can present an indication that additional information or a product object (e.g., a coupon) can be received by providing specific input. If the user provides specific input by, for example, speaking a specific word or phrase identified in the media clip, the media device can present additional information and / or send a product object to a device or user profile associated with the user (e.g., an email address). If the user provides different input, such as different words or phrases, the media device can present additional information without sending a product object, send a product object without presenting additional information, re-present a supplementary media clip, restart the trigger media from the beginning, present an alternative media clip associated with the same product, present an alternative media clip associated with a different product or service, terminate the event listener, or a combination thereof.

[0026] In an illustrative example, a media device (e.g., a mobile device, tablet, computing device, television, etc.) can receive an identification of a video segment from an Automatic Content Identification Service. The identification of the video segment may correspond to a video segment currently being presented by the media device. For example, the media device may receive a request to activate a trigger from a remote device. This trigger may cause the media device to perform one or more processes in response to detecting that a specific media segment is being displayed by the media device. Upon activating the trigger, the media device may begin sending one or more cues derived from the video and / or audio channels of the video segment to the Automatic Content Identification Service to determine whether the currently presented media segment corresponds to the specific media segment associated with the trigger.

[0027] A clue can be a data structure comprising a representation of pixel values ​​of one or more consecutive pixels from a video frame extracted from a video channel. Alternatively or additionally, the data structure may include a representation of one or more audio segments extracted from an audio channel. Alternatively or additionally, the data structure may also include metadata derived by the media device from the video segment, such as, but not limited to, identification of the channel or media source, a timestamp corresponding to the generation of the clue, a time interval indicating the time since the start of the video segment, information associated with the media device (e.g., media device identifier, media device type, hardware and / or software installed on the media device, Internet Protocol address, etc.), and combinations thereof. The media device can then send the one or more clues to an automatic content identification service.

[0028] The automatic content identification service may be a component of a media device and / or a component of a remote device (e.g., a server, content provider, etc.). The automatic content identification service may include or have access to a database of known cues associated with known media segments. Upon receiving an unknown cue, the automatic content identification service may identify the closest matching known cue in the database (e.g., determined by distance algorithms, pattern matching, machine learning models, and / or the like). The remote device may then assign an identifier of the known media segment of the closest matching known cue to the unknown cue. The automatic content identification service may then send this identifier to the media device.

[0029] A media device can send a notification to its display device based on the identification of a video clip. The display device can be a component of the media device (e.g., a component that displays the video component of the clip). Alternatively, the display device can be a device connected to the media device (e.g., via an HDMI cable, DisplayPort cable, network connection, etc.). Alternatively, the media device can send the notification to a device associated with the display device or its user (e.g., a mobile device). The notification can include information associated with the video clip and a request for input. The notification can include alphanumeric text, images, video, audio clips, combinations thereof, etc. The notification can indicate one or more input options and one or more interfaces on which said one or more input options can be sent. For example, the notification could include alphanumeric text such as: “Say ‘pizza’ to receive a coupon for your next order.”

[0030] Media devices can detect one or more audio segments associated with a notification. For audio-based input, the media device may include a speech-to-text model configured to translate audio into alphanumeric strings. The media device can compare the alphanumeric string with the requested input to determine whether the requested input has been provided. Returning to the previous example, the media device can determine whether the alphanumeric string contains the word "pizza". In some cases, the media device can also determine whether the alphanumeric string corresponds to a common variant of the requested input (e.g., slang, different languages, synonyms, etc.). For non-audio-based input (e.g., text, email, input from a remote control, etc.), the media device can directly compare the input with the requested input.

[0031] In some cases, the media device may monitor one or more audio segments during the presentation of an identified media segment. If the one or more audio segments are received after the presentation of the identified media segment has ended (e.g., the media device identifies a new media segment being presented), the media device may ignore the incoming audio segments. In other cases, the media device may monitor one or more audio segments for a time interval longer than the presentation time of the identified media segment. This time interval may begin when a notification is presented and end at some point after the media segment has ended, giving the user more time to provide the requested input. For example, the time interval could be 30 minutes, 1 hour, 24 hours, etc. The time interval can be defined when the trigger is defined and communicated to the user via a notification. Continuing with the previous example, the notification could include “Say ‘pizza’ within the next 24 hours to receive a coupon for your next order.”

[0032] The media device can then facilitate the presentation of an object associated with a video segment in response to the detection of one or more audio segments. Facilitating the presentation of the object may include, but is not limited to, displaying a representation of the object (e.g., alphanumeric text, images, videos, etc., serial numbers, or product codes associated with the object), sending the object to a device associated with the media device (e.g., a mobile device, tablet, computing device, etc.), sending the object to a user profile associated with a user of the media device (e.g., a user identified using one or more audio segments) (e.g., an email address, a profile associated with a product or service in the identified video segment, an entity identified in the identified video segment, etc.), or combinations thereof. Returning to the previous example, after detecting that a user says "pizza," the media device can send a coupon to the user's mobile device.

[0033] Figure 1 A block diagram of an example computing device configured to perform contextual audio processing according to various aspects of this disclosure is shown. Media device 104 may include one or more processing components (e.g., system-on-a-chip, central processing unit, application-specific integrated circuit, field-programmable gate array, and / or the like), memory (e.g., volatile and non-volatile memory, database, etc.), network processor (e.g., including Wi-Fi transceivers, Bluetooth transceivers, and other transceivers, etc.), and one or more sensors.

[0034] Media device 104 can be configured to present media to one or more users using display 108 and / or one or more wireless devices (e.g., other display devices, mobile devices, tablets, and / or the like) connected via a network processor. Media device 104 can retrieve media from media database 152 (or alternatively receive media from one or more broadcast sources, remote sources via the network processor, external devices, etc.). Media can be loaded by media player 148, which can process media based on a video container (e.g., MPEG-4, QuickTime Movie, WavefileAudio File Format, Audio Video Interleave, etc.). Media player 148 can pass the media to video decoder 144, which decodes the video into a sequence of video frames that can be displayed by display 108. The video frame sequence can be passed to video frame processor 140 for preparation for display. Alternatively, media can be generated by an interactive service operating within application manager 136. Application manager 136 can pass the frame sequence generated by the interactive service to video frame processor 140.

[0035] A sequence of video frames can be passed to a system-on-a-chip (SOC) 112. The SOC 112 may include processing units configured to render sequences of video components and / or audio components. The SOC 112 may include a central processing unit (CPU) 124, a graphics processing unit (GPU) 120, memory 128 (e.g., volatile memory such as random access memory or read-only memory, non-volatile memory such as magnetic, flash memory, etc.), an input / output interface 132, and a video frame buffer 116.

[0036] The SOC 112 can generate cues from one or more video frames stored in the video frame buffer 116 before or during the rendering of one or more video frames on the display 108. Cues can be generated from one or more pixel arrays (also called pixel patches) of the video frames. A pixel patch can be any arbitrary shape or pattern, such as (but not limited to) a y×z pixel array comprising y pixels horizontally multiplied by z pixels vertically from the video frames. Pixels can include color values ​​(such as red, green, and blue values) and intensity values. The color values ​​of a pixel can be represented by eight-bit binary values ​​for each color. Other suitable color values ​​that can be used to represent pixel colors include luminance and chromaticity (Y, Cb, Cr, also known as YUV) values ​​or any other suitable color values.

[0037] SOC 112 can derive an average value for each thread. The average value can be a 4-bit data record representing the thread. The display device can generate a thread by aggregating the average value of each pixel patch and adding a timestamp corresponding to the frame from which the pixel patches were obtained. The timestamp can correspond to an epoch time (e.g., it can represent the total elapsed time in fractions of seconds since midnight on January 1, 1970), a scheduled start time, an offset time (e.g., from the start of the media being presented or when the display device is powered on), etc. The thread can also include metadata, which may include any information about the media being presented, such as a program identifier, program time, program length, or any other information (if known).

[0038] In some examples, cues can be derived from any number of pixel patches obtained from a single video frame. Increasing the number of pixel patches included in a cue increases the data size of the cue, which can increase the processing load on the display device and the processing load on one or more cloud networks that may be operating to identify content. For example, a cue derived from 5 pixel patches could correspond to 600 bits of data (24 bits per pixel patch multiplied by 5 pixel patches), excluding timestamps and any metadata. Increasing the number of video patches obtained from a video frame can improve the accuracy of boundary detection and content recognition at the cost of increased processing load. Reducing the number of video patches obtained from a video frame can reduce the accuracy of boundary detection and content recognition while reducing the processing load on the display device. The display device can dynamically determine whether to use more or fewer pixel patches to generate cues based on target accuracy and / or the processing load of the display device.

[0039] Unknown cues can be compared with known cues of known media stored in database 156 to identify the media segment corresponding to the unknown cues. The media device can use distance algorithms (e.g., Euclidean distance, cosine distance, semi-sine distance, Minkowski distance, etc.) or other matching algorithms to identify the known cues closest to the unknown cues. If the distance is less than a threshold distance, SOC 112 can assign the identifier of the known cues to the unknown cues, thereby identifying the media segment from which the unknown cues originate. The cue database 156 can be a component of media device 104 (e.g., stored in memory 128 or other memory of media device 104 (not shown)) or can be a remote component (as shown).

[0040] Upon identifying media being presented, media device 104 can determine whether a trigger is associated with that media. The trigger can be stored in memory 128 and defines a process for providing contextual presentation of the media based on input associated with the media being presented. The trigger can define: a notification to be presented when a segment of media is detected being presented; identification of input requested in response to the presentation notification; identification of a time interval within which input will be accepted; identification of one or more processes to be performed in response to the detection of input within that time interval; and / or the like.

[0041] If media device 104 identifies a trigger associated with the identified media, media device 104 may activate the trigger. Activating the trigger may include presenting a notification to the user of media device 104 (e.g., via display 108 and / or via the network interface of media device 104) requesting input from the user. The notification may include alphanumeric text, images, video clips, audio clips, etc. If the notification will be presented by display 108, the notification may be presented on top of the media clip being presented at predetermined time intervals (e.g., via a pop-up window, etc.). Alternatively, the notification may be included within the identified media clip, so that media device 104 does not need to present additional content separate from the media being presented.

[0042] Media device 104 may initiate an event listener when it determines that the identified media is associated with a trigger (and / or when a notification is presented, etc.). The event listener may be a process executed by CPU 124 to monitor I / O interface 132 for specific input. Upon detecting input from a specific input interface (e.g., as identified by the trigger) or from any input interface, SOC 112 may process the input to determine whether it corresponds to the input identified by the trigger. For example, a notification may request specific input (e.g., “Say ‘pizza’ within the next 24 hours to receive a coupon.”), and the event listener may monitor an audio interface (e.g., a microphone, etc.) for audio input. Upon detecting an audio segment, SOC 112 may process the audio segment (e.g., speech-to-text, etc.) to determine whether the audio segment corresponds to the input requested by the trigger (e.g., the word “pizza”, etc.).

[0043] SOC 112 may include or have access to one or more speech-to-text models configured to convert audio segments into alphanumeric strings. In some cases, the one or more speech-to-text models may include one or more machine learning models. The machine learning models may output alphanumeric strings corresponding to the audio segments. SOC 112 may then compare this alphanumeric string with the input requested by the trigger. Alternatively, SOC 112 may send the audio segments to a server for recognition. The server may return an indication of whether the alphanumeric string corresponding to the audio segment and / or whether the audio segment corresponds to the input requested by the trigger.

[0044] In some examples, the machine learning model may include an additional classification layer configured to identify the speaker of an audio segment. This additional classification layer may be trained based on ambient audio detected by the I / O interface 132. In some cases, the ambient audio may be filtered prior to training to prevent audio channels of the media presented by the media device 104 from being included. The additional classification layer may be configured to distinguish and / or identify speakers using the media device 104 (e.g., by name and / or by user identifier, etc.). Alternatively, a separate machine learning model may be used to distinguish and / or identify speakers.

[0045] One or more machine learning models for speech-to-text models can include any type of machine learning model. Examples of such machine learning models include, but are not limited to, neural networks such as recurrent neural networks (e.g., Long Short-Term Memory (LSTM), masked recurrent neural networks, etc.), You Only See Once (YOLO), EfficientDet, deep learning networks, transformers (Generative Pre-trained Transformer (GPT), Bidirectional Encoder Representation from Transformer (BERT), Text-to-Text Transformer (T5), etc.), Generative Adversarial Networks (GANs), Recurrent Gated Units (GRUs), combinations thereof, and / or the like. In implementations with an additional classification layer, the classification layer can be part of one of the aforementioned machine learning models or a separate machine learning model. Examples of such classification layers include, but are not limited to, one or more of the following: Naive Bayes, logistic regression models, perceptrons, support vector machines, random forest models, linear discriminant analysis models, k-nearest neighbors, gradient boosting, combinations thereof, and / or the like.

[0046] The one or more machine learning models can be trained using audio segments derived from media presented to and / or media that can be presented by media device 104 (e.g., broadcast media, streaming media, etc.) to customize the training of the one or more machine learning models to the type of audio segments that the one or more machine learning models will process at runtime (e.g., after training). If the one or more machine learning models are trained by media device 104, media device 104 can sample the media presented by media device 104 over time to collect training data that can be used to train the one or more machine learning models. Alternatively or additionally, media device 104 can receive audio segments and / or training data for training data from one or more remote devices. If the one or more machine learning models are trained by remote devices, the remote devices can sample multiple media sources based on the type of media that can be presented by media device 104. Alternatively, one or more machine learning models can be trained using audio clips derived from any source (e.g., media that can be presented by media device 104, other broadcast media such as radio or music, audiobooks, speeches, manually generated media, trained text-to-speech models, any other media source, and / or combinations thereof and / or the like). In some examples, the data can be augmented with additional training data (e.g., process-generated data, manually generated data, combinations thereof, and / or the like) and / or metadata (e.g., labels for supervised learning, features derived from the training data, combinations thereof, and / or the like).

[0047] The one or more machine learning models may be trained using supervised learning, unsupervised learning, semi-supervised learning, transfer learning, meta-learning, reinforcement learning, or combinations thereof. The one or more machine learning models may be trained for predetermined time intervals, predetermined number of iterations, and / or until one or more accuracy metrics are achieved (e.g., such as, but not limited to, accuracy, precision, area under the curve, log loss, F1 score, longest common subsequence (LCS) such as ROUGE-L, bilingual evaluation substitute (BLEU), mean absolute error, mean squared error, etc.).

[0048] In some other cases, the one or more speech-to-text models may not include machine learning, but instead use instructions to perform spectral pattern analysis. Spectral pattern analysis transforms the unknown audio segment into the frequency domain and compares portions of the audio segment in the frequency domain with spectral features corresponding to known sounds or words. If a portion of the unknown audio segment matches one or more spectral features of a known sound or word, the known sound or word is assigned to the unknown audio segment.

[0049] In some other cases, the one or more speech-to-text models may include a combination of one or more machine learning models and spectral pattern matching.

[0050] Media device 104 may perform one or more processes in response to processing an audio segment. For example, if an alphanumeric string corresponds to input requested by a trigger, media device 104 may facilitate the transmission of a product object to a user (e.g., to a user registered on media device 104, to a user identified as providing the audio segment, etc.). Media device 104 may present the product object via display 108, transmit a representation of the product object to a user-associated mobile device, user profile (e.g., such as an email address, an account held by the entity presenting the media segment, etc.), or a combination thereof via I / O interface 132 (e.g., via a Wi-Fi connection to the Internet or a mobile device, an Ethernet connection, a Bluetooth connection, etc.), or a combination thereof. The trigger may include instructions that can perform the one or more processes. The specific process performed may depend on the requested input, the received input (e.g., an alphanumeric string, etc.), and / or the interface on which the input is received (e.g., a remote control, microphone, camera, network interface, etc.).

[0051] The following are example procedures for one or more of the aforementioned processes: If the alphanumeric string corresponds to a termination condition (e.g., "Stop", "Cancel", "Terminate", etc.), media device 104 may prevent further processing associated with the trigger and / or terminate the presentation of the identified media segment (e.g., if it is still being presented). If the alphanumeric string corresponds to an information condition (e.g., "More Information", "Information", "Who is...", "What is...", etc.), media device 104 may present additional information associated with the identified media segment, such as, but not limited to, information associated with the product depicted in the identified media segment, the service depicted in the identified media segment, actors, directors, scenes, filming locations, etc. The additional information may be provided and / or requested along with the trigger when the alphanumeric string is detected. The additional information may be presented within a window on the portion of the media being presented on display 108. If the alphanumeric string is a presentation condition (e.g., "restart", "reboot", etc.), media device 104 can restart the identified media segment (if it is currently being presented) or play the identified media segment from the beginning (e.g., pause the currently presented media, present the identified media segment, and return to the paused media when the identified media segment terminates, etc.). If the alphanumeric string corresponds to an alternative presentation condition (e.g., "replace", "play alternative content", etc.), media device 104 can replace the identified media segment with an alternative media segment. The alternative media segment can be associated with the same product or service as the identified media segment (e.g., a different advertisement for the same product or service), or it can be associated with a different product or service (e.g., an advertisement for some other product or service).

[0052] Figure 2 This is a block diagram of an example system for contextual audio processing according to various aspects of this disclosure. A media device may perform contextual audio processing in part based on a media segment being presented by the media device. In some cases, the media device may receive a trigger that defines the context of context-dependent processing to be associated with a specific media segment. The trigger may identify a specific media segment (e.g., an identifier, one or more cues of a specific media segment, etc.), audio input configured to trigger a contextual process, and the identification of one or more contextual processes. Instructions for implementing the identification of the specific media segment, receiving and / or processing the audio input, one or more contextual processes, etc., may be included in the trigger, and these instructions may be obtained by the media device in response to receiving the trigger, and / or the instructions may be stored by the media device. In some cases, the trigger may be received via a network interface of the media device. In other cases, the trigger may be embedded in the media stream being presented by the media device (e.g., as metadata, watermark, etc.).

[0053] Media devices can monitor media streams and identify media segments being presented within those streams. In some cases, media devices may monitor media streams only if triggers are stored in memory to prevent consuming processing resources when no triggers are present. If the media device determines that triggers are stored in memory, it can perform a monitoring process. The monitoring process can cause a cue generator 204 to begin generating cuees that can be used to identify the media being presented by the media device. In some examples, the cue generator 204 may generate cuees at regular intervals (e.g., every 1 second, every 1 minute, 5 minutes, etc.). The regular intervals may vary based on the determination that the currently presented media segment is a new media segment (e.g., changes in average brightness and / or chroma, machine learning models, etc.). When a new media segment is detected being presented, the cue generator 204 temporarily increases the rate at which it generates cuees until that media segment is identified. The cue generator 204 can then revert to a reduced rate. For example, the cue generator 204 may generate cuees every 3 minutes until a new media segment (e.g., a new advertisement, television program, movie, etc.) is detected being presented. Then, the lead generator 204 can generate leads every second until a new media segment is identified. The lead generator 204 can then revert to generating leads every 3 minutes. The described time intervals (e.g., both normal and increased rates) can be individually selected based on the characteristics of the currently presented media segment, the characteristics of the media segment expected to be presented next, user input, previous iterations of the lead generator 204, machine learning models, combinations thereof, etc.

[0054] In other examples, the cue generator 204 can generate cues when a change in the media being presented by the media device is detected. For example, the cue generator 204 can detect changes in the average luminance and / or chroma of the media being presented across one or more video frames to detect a change in the media being presented (e.g., a new media segment). The cue generator 204 can determine that a new media segment is being presented when the difference in the average luminance and / or chroma of a video frame relative to a previous video frame is greater than a first threshold and / or less than a second threshold. By calculating the difference in average luminance and / or chroma across a set of multiple video frames, the media device can detect features that may be characteristics of a change in the media being presented, such as fade-in to black, fade-out from black, and scene changes. Additionally or alternatively, the media device can use a machine learning model to predict the probability that the currently presented media segment is different from a previously presented media segment. The machine learning model can be a classifier configured to process a set of video frames (and / or audio associated with the video frames) to determine whether a first video frame corresponds to the same media segment as a previous second video frame. Examples of machine learning models include, but are not limited to, deep learning networks, convolutional neural networks, recurrent neural networks, Naive Bayes, support vector machines, k-nearest neighbors, perceptrons, logistic regression, and / or the like. If the confidence level of the prediction is greater than a threshold, the media device can determine that new media is being presented. Upon detecting that a new media segment is being generated, the cue generator 204 can begin generating cue from the media segment at regular intervals as previously described.

[0055] Media devices may include register flags, namely Client TV ACR Enable / Disable 208, which can be used to enable and disable content identification. In some cases, register flags can be set to True (or '1' or 'On', etc.) to enable content identification performed by the media device. Register flags can be set to False (or '0' or 'Off', etc.) to enable content identification performed by the ACR server. In some cases, register flags can be toggled via a resource allocation process to allocate processing resources to the media device as needed. For example, if processing resources are low or if a process with a higher priority than the content identification process requests additional processing resources, register flags can be set to False (or '0' or 'Off', etc.) to offload content identification to the ACR server. Once sufficient processing resources are available, or if content identification has a higher priority than another process that has allocated processing resources (e.g., causing the media device to reallocate processing resources to other processes and assign those processes to the content identification process and set register flags to True, etc.), register flags can be set to True (or '1' or 'On', etc.).

[0056] If the register flag is set to true (or '1' or 'on', etc.), the generated thread can be processed by the client ACR thread processor 212, which can be a process executed by the media device or a device connected to the media device. If the register flag is set to false (or '0' or 'off', etc.), the generated thread can be processed by the server ACR thread processor 216, which can be a process executed by a remote device. The client ACR thread processor 212 and the server ACR thread processor 216 can process threads in a similar manner. For example, a thread can be processed by comparing a thread associated with an unknown media segment with known threads associated with known media segments stored in a known thread database. A distance algorithm can be used to identify the known thread that is closest to the unknown thread. If the distance is within a threshold distance, the identifier of the known media segment corresponding to the closest matching known thread can be assigned to the unknown thread.

[0057] In some cases, if the client ACR line processor 212 cannot identify the closest matching known line within a threshold distance, the media device may send the line to the server ACR line processor 216, as the server ACR line processor 216 may include a larger database of known lines (or have access to a larger database of known lines). Alternatively, the client ACR line processor 212 may return "unknown" (or null, etc.), and the process may be repeated with another line derived from the same media segment until a matching known line is identified. The identifier of the known media segment may be sent from the client ACR line processor 212 and / or the server ACR line processor 216 (depending on which processor is selected) to the context processor 220.

[0058] Context processor 220 can identify triggers associated with the identified media segment. Triggers can identify context instructions to be associated with the execution of the identified media segment. In some cases, context instructions can cause the media device to present a notification (e.g., a request for audio input) via the media device's display, text message, email, mobile device (e.g., via an app running on the mobile device, push notification, etc.), instant or direct message, or a combination thereof. Alternatively, the notification can be presented by the media segment. Context instructions can cause the media device to instantiate an event listener to detect events associated with the notification (e.g., receiving audio input via audio processor 224). The event listener can execute over a time interval that begins when the notification is presented and ends sometime after the identified media segment has terminated. For example, the identified media segment could be a 30-second advertisement. When the media segment is identified, the media device can present the notification and instantiate the event listener approximately simultaneously. The event listener can execute over a longer time interval than the media segment (e.g., 1 hour, 6 hours, 24 hours, etc.) to enable detection of input associated with the media segment long after its termination.

[0059] An event listener can detect audio input from audio processor 224. Audio processor 224 can process audio segments detected by the media device and output an alphanumeric string corresponding to the audio segment. For example, audio processor 224 may include a speech-to-text model configured to translate audio segments of speech into alphanumeric strings. The event listener can detect the audio input and generate an event including the alphanumeric string and a timestamp corresponding to the time the audio input was received. Context processor 220 can process the event to determine whether the alphanumeric string corresponds to the expected alphanumeric string of the trigger (e.g., to notify of the requested input). If the alphanumeric string does not correspond to the expected alphanumeric string of the trigger, the process can return to context processor 220 until further audio input is received.

[0060] If the alphanumeric string corresponds to the expected alphanumeric string of the trigger, the context processor 220 may facilitate a context output 228, which may include the execution of one or more procedures of the trigger. The one or more procedures may be identified in a notification and based on specific input received. Examples of the one or more procedures include, but are not limited to, presenting additional information associated with the identified media segment (e.g., featured products or services), sending product objects (e.g., coupons for the product and / or service) to a device or user profile (e.g., email address) associated with the media device, restarting the identified media segment, presenting alternative media segments associated with the same product and / or service, presenting alternative media segments associated with different products or services, terminating event listeners, combinations thereof, etc.

[0061] Figure 3 A diagram illustrating an example process for performing contextual audio processing in a computing device according to various aspects of this disclosure is shown. The media device may initiate one or more triggers configured to execute a process in response to detecting the presentation of a specific media segment. The media device may define one or more triggers (e.g., based on a user's viewing history, user input, web browsing activity, instructions embedded in the media segment such as metadata or watermarks, etc.). Alternatively or additionally, the media device may receive triggers from a remote device (e.g., such as a context server or other remote device). A trigger may be a data structure defining a process to be performed associated with a specific media segment. In some examples, a trigger may include, but is not limited to, an identifier, an identification of a specific media segment, an identification of an instruction to present a notification to a user, an identification of an event associated with the notification (e.g., a specific input from a specific input source), an identification of an instruction to be executed in response to the detection of an event, an identification of information associated with a specific media segment (e.g., information such as information associated with a product or service represented in a specific media segment, an identification of an actor, an identification of a director, production information, etc.), a time interval in which the detection of an event will trigger the execution of an instruction, an identification of a supplementary media segment to replace a specific media segment, or a combination thereof.

[0062] In some cases, a trigger may include instructions that can be executed by a media device to request additional information that can be used to execute the trigger. This additional information may include instructions, data, supplementary media segments, etc. For example, a trigger may recognize one or more application programming interfaces (APIs) that can be used to process the specific input requested by the trigger. The media device may request additional information to enable the trigger's execution.

[0063] A media device can begin identifying media segments presented by the media device to determine whether the media segment corresponds to a media segment triggered by a trigger. The media device can use metadata or a watermark embedded in the media segment to identify the media segment being presented. A watermark can include information (or instructions) embedded within the video and / or audio components of the media segment. For example, a watermark can be embedded by modifying the chroma and / or luminance values ​​of a consecutive set of pixels in a video frame. Luminance (as an example) can be increased by a predetermined value to indicate a first value (e.g., such as '1') and decreased by a predetermined value to indicate a second value (e.g., such as '0'). The consecutive set of pixels can be decoded into a sequence of 0s and 1s (e.g., binary code), which can convey information such as the identification of the media segment. Metadata or a watermark can include the identification of the media segment, the identification of one or more media segments to be presented in the future, a media segment schedule, a timestamp, a time offset (indicating the position of the media segment being presented relative to the start time), etc.

[0064] If a media segment does not include metadata or a watermark, the media device can generate clues that can be used to identify the media segment. The media device can generate clues by sampling pixel data and / or audio segments from one or more video frames. The media device can generate clues at regular intervals (e.g., 1 second, 1 minute, 3 minutes, etc.) until an unknown media segment is identified. The media device can reduce the rate at which new clues are generated until it is determined that the current media segment is different from the most recently identified media segment and that the current media segment is unknown.

[0065] In a 304 error, the media device can send an unknown clue to the ACR server for identification. The ACR server can compare the unknown clue with known clues associated with known media segments. The ACR server can use a fuzzy matching algorithm to identify the closest matching known clue to the unknown clue. The fuzzy matching algorithm can be a distance algorithm, a machine learning model, and / or the like. If the distance between the closest known clue and the unknown clue is less than a threshold, the ACR server can assign the identifier of the known media segment of the known clue to the unknown clue.

[0066] In step 308, the ACR server can send an identifier assigned to an unknown clue (e.g., an identifier for a known media segment of a known clue, or an identifier if the nearest matching known clue is no more than a threshold away from the known clue) to the media device. If the media device receives an "unknown" from the ACR server, it can send additional clues to the ACR server until an identifier is returned.

[0067] The media device can determine whether an identifier received from the ACR server corresponds to a specific media segment of the trigger. If the identifier received from the ACR server corresponds to a specific media segment of the trigger, the media device can present a notification (e.g., via the media device's display, a mobile device associated with the media device, a text message, email, etc.) to identify the trigger and request input associated with the identified media segment. The notification can be included in the trigger or retrieved in response to receiving the trigger. Alternatively, the notification can be included in the media segment.

[0068] For example, a media segment (such as an advertisement) can be associated with a trigger. When the presentation of a media segment is detected, the media device can present (if not included in the media segment) a notification requesting the user to provide audio input related to the context of the media segment (such as something related to the product or service represented in the media segment) to obtain a product object (such as a coupon for that product or service).

[0069] Alternatively, the media device may present a notification in response to detecting a watermark embedded in a media segment associated with the trigger. The media device may not need to identify the media segment (e.g., via metadata, watermark, or automatic content identification) to perform this action. Figure 3 The remaining blocks (e.g., blocks 312 through 328, etc.). In some cases, the media device may store instructions that are executed in response to the detection of a watermark. In other cases, the watermark may include instructions that, when executed by the media device, cause the media device to present a notification.

[0070] In section 312, a media device can instantiate an event listener configured to monitor an interface identified by a notification. Returning to the previous example where the notification requests audio input, the event listener can monitor an audio interface (e.g., a microphone or other audio-based input device). The media device can then wait for the event listener to generate an event. The media device can be configured to instantiate an event listener for any interface or device configured to monitor the media device, including but not limited to microphones, optical sensors, input devices (e.g., mice, keyboards, remote controls, or other devices configured to remotely control the operation of the media device, game controllers, and / or the like), network interfaces of the media device (e.g., interfaces using Bluetooth, Ethernet, Wi-Fi, Zigbee, Z-Wave, etc.).

[0071] Event listeners can execute at predetermined time intervals (e.g., defined by triggers, user input, default settings, etc.). In some cases, an event listener can begin execution when it recognizes that the media segment being rendered corresponds to a trigger and terminate after the media segment ends, giving the user more time to provide the requested input. Returning to the previous example, the media segment could be 15 seconds long, making it difficult for the user to provide the requested input before the media segment ends. The time interval can be any predetermined time interval, such as, but not limited to, 30 minutes, 1 hour, 6 hours, 24 hours, etc.

[0072] In section 316, an event listener can detect input from an input device corresponding to an interface and generate an event. The event may include identification of the interface receiving the input, the received input, a timestamp corresponding to the time the event was received by the media device, and / or the like. Input may be audio clips (e.g., spoken words or phrases), alphanumeric strings (e.g., text messages, emails, input from a remote control), graphics (e.g., QR codes, images, symbols), light or sequences of light (e.g., camera flashes), and / or the like. The event listener may continue execution until it is determined that the received input corresponds to the requested input.

[0073] In some cases, the media device can process the input to determine whether the received input corresponds to the requested input. For example, for audio-based input, the media device may include a speech-to-text model configured to process audio segments and output alphanumeric strings. The speech-to-text model can translate words spoken by the user into text. The media device can process the audio segment into an alphanumeric string and compare that alphanumeric string with the requested input. Alternatively or additionally, at 320, the media device can send the input to an audio server. The audio server can analyze the audio segment, and at 324, return the alphanumeric string representing the audio segment to the media server.

[0074] In some cases, at 328, the audio server may send a communication containing an alphanumeric string to the context server. The context server may determine whether the alphanumeric string corresponds to the requested input. Alternatively, if the alphanumeric string corresponds to the requested input, then at 332, the media device may send a communication to the context server indicating that the received input corresponds to the requested input. The context server may perform one or more procedures associated with the media segment and the trigger. Returning to the previous example, the context server may send the product object (at 336) to the media device's input device (e.g., a mobile device connected to the media device, a remote control, a speaker, and / or the like) or to the media device (at 340). In some cases, the one or more procedures may be performed by the media device.

[0075] In some cases, triggers can recognize inputs for multiple requests, each associated with one or more different processes. For example, the input for a request might ask a user to “say ‘pizza’ anytime within the next hour to receive a coupon.” If the audio input corresponds to the word “pizza,” the context server can send a product object to the media device. If the audio input corresponds to “restart,” the media device can restart or replay the media segment. The media device can re-insert a notification when restarting or replaying the media segment. Examples of processes that can be performed include, but are not limited to, facilitating the sending of a product object (e.g., via an input device, a media device as an image and / or as audio, a mobile device, push notification, text message, email, direct or instant message, combinations thereof, etc.), restarting a media segment, replaying a media segment, replacing a media segment with a supplementary media segment, providing additional information associated with the media segment (and / or the product or service represented by the media segment), presenting a webpage associated with the media segment (e.g., on the media segment, etc.), and / or the like.

[0076] While executing one or more of the processes, the media device may terminate the event listener to prevent the generation of duplicate events and the waste of processing resources to process duplicate events.

[0077] Figure 4A flowchart illustrating an example process for contextual audio processing in a media device according to various aspects of this disclosure is shown. At block 404, the media device (e.g., a mobile device, tablet, computing device, television, etc.) may receive an identification of a video segment from an automatic content identification service. The identification of the video segment may correspond to a video segment currently being presented by the media device. For example, the media device may receive a request to activate a trigger from a remote device. This trigger may cause the media device to perform one or more processes in response to detecting that a specific media segment is being displayed by the media device. Upon activating the trigger, the media device may begin sending one or more cues derived from the video channel and / or audio channel of the video segment to the automatic content identification service to determine whether the currently presented media segment corresponds to a specific media segment corresponding to the trigger.

[0078] A clue can be a data structure comprising a representation of pixel values ​​of one or more consecutive pixels from a video frame extracted from a video channel. Alternatively or additionally, the data structure may include a representation of one or more audio segments extracted from an audio channel. Alternatively or additionally, the data structure may also include metadata derived by the media device from the video segment, such as, but not limited to, identification of the channel or media source, a timestamp corresponding to the generation of the clue, a time interval indicating the time since the start of the video segment, information associated with the media device (e.g., media device identifier, media device type, hardware and / or software installed on the media device, Internet Protocol address, etc.), and combinations thereof. The media device can then send the one or more clues to an automatic content identification service.

[0079] The automatic content identification service may be a component of a media device and / or a component of a remote device (e.g., a server, content provider, etc.). The automatic content identification service may include or have access to a database of known cues associated with known media segments. Upon receiving an unknown cue, the automatic content identification service may identify the closest matching known cue in the database (e.g., determined by distance algorithms, pattern matching, machine learning models, and / or the like). The remote device may then assign an identifier of the known media segment of the closest matching known cue to the unknown cue. The automatic content identification service may then send this identifier to the media device.

[0080] In block 408, a media device may send a notification to its display device based on the identification of a video segment. The display device may be a component of the media device (e.g., a component displaying the video component of the segment). Alternatively, the display device may be a device connected to the media device (e.g., via an HDMI cable, DisplayPort cable, network connection, etc.). Alternatively, the notification may be included within the video segment (e.g., the media device may not need to do anything to present the notification). Alternatively, the media device may send the notification to a device associated with the display device or its user (e.g., a mobile device, etc.). The notification may include information associated with the video segment and a request for input. The notification may include alphanumeric text, images, video, audio segments, combinations thereof, etc. The notification may indicate one or more input options and one or more interfaces on which said one or more input options can be sent. For example, the notification may include alphanumeric text such as: “Say ‘pizza’ to receive a coupon for your next order.”

[0081] Media devices can instantiate event listeners in response to the recognition of video segments. An event listener can be a process that monitors one or more interfaces for input and generates an event in response to the detection of input. An event can include the input, the identification of the interface on which the input is received, a timestamp corresponding to the time the input is received, a combination thereof, etc.

[0082] In block 412, the media device may detect one or more audio segments associated with the notification. For audio-based input, the media device may include a speech-to-text model configured to translate audio into alphanumeric strings. The speech-to-text model may be implemented by one or more machine learning models and / or spectral pattern matching.

[0083] Media devices can compare the alphanumeric string with the requested input to determine if the requested input has been provided. Returning to the previous example, the media device can determine if the alphanumeric string contains the word "pizza". In some cases, the media device can also determine if the alphanumeric string corresponds to a common variant of the requested input (e.g., slang, different languages, synonyms, etc.). For non-audio-based input (e.g., text, email, input from a remote control, etc.), the media device can directly compare the input with the requested input.

[0084] In some cases, the media device may monitor one or more audio segments during the presentation of an identified media segment. If the one or more audio segments are received after the presentation of the identified media segment has ended (e.g., the media device identifies a new media segment being presented), the media device may ignore the incoming audio segments. In other cases, the media device may monitor one or more audio segments for a time interval longer than the presentation time of the identified media segment. This time interval may begin when a notification is presented and end at some point after the media segment has ended, giving the user more time to provide the requested input. For example, the time interval could be 30 minutes, 1 hour, 24 hours, etc. The time interval can be defined when the trigger is defined and communicated to the user via a notification. Continuing with the previous example, the notification could include “Say ‘pizza’ within the next 24 hours to receive a coupon for your next order.”

[0085] In block 416, the media device can then, in response to detecting the one or more audio segments, facilitate the presentation of an object associated with a video segment. Facilitating the presentation of an object can include, but is not limited to, displaying a representation of the object (e.g., alphanumeric text, images, videos, etc., serial numbers, or product codes associated with the object), sending the object to a device associated with the media device (e.g., a mobile device, tablet, computing device, etc.), sending the object to a user profile associated with a user of the media device (e.g., a user identified using the one or more audio segments, such as an email address, a profile associated with a product or service in the identified video segment, an entity identified in the identified video segment, etc.), or combinations thereof. Returning to the previous example, after detecting that a user says “pizza,” the media device can send a coupon to the user’s mobile device.

[0086] Figure 5 A computing system architecture, including various components that are electrically communicating with each other, is illustrated according to various aspects of this disclosure. According to some implementations, Figure 5The example computing system architecture 500 shown includes a computing device 502 having various components that communicate electrically with each other using a connection 506, such as a bus. The example computing system architecture 500 includes a processing unit 504 that communicates electrically with various system components using the connection 506, and includes a system memory 514. In some embodiments, the system memory 514 includes read-only memory (ROM), random access memory (RAM), and other such memory technologies, including but not limited to those described herein. In some embodiments, the example computing system architecture 500 includes a cache memory 508 that is directly connected to, closely proximate to, or integrated into the processor 504. The system architecture 500 can copy data from memory 514 and / or storage device 510 to the cache 508 for fast access by the processor 504. In this way, the cache 508 can provide performance improvements that reduce or eliminate processor latency in the processor 504 caused by waiting for data. Using modules, methods, and services such as those described herein, the processor 504 can be configured to perform various actions. In some implementations, cache 508 may include various types of caches, including, for example, level 1 (L1) and level 2 (L2) caches. Memory 514 may be referred to herein as system memory or computer system memory. Memory 514 may at different times include elements of the operating system, one or more applications, data associated with the operating system or said one or more applications, or other such data associated with computing device 502.

[0087] Other system memory 514 may also be available. Memory 514 may include various different types of memory with different performance characteristics. Processor 504 may include any general-purpose processor and one or more hardware or software services, such as service 512 stored in storage device 510, which are configured to control processor 504 and dedicated processors in which software instructions are incorporated into the actual processor design. Processor 504 may be a fully self-contained computing system, including multiple cores or processors, connectors (e.g., buses), memory, memory controllers, caches, etc. In some embodiments, such a self-contained computing system with multiple cores is symmetric. In some embodiments, such a self-contained computing system with multiple cores is asymmetric. In some embodiments, processor 504 may be a microprocessor, microcontroller, digital signal processor (“DSP”), or a combination of these and / or other types of processors. In some implementations, processor 504 may include multiple elements such as cores, one or more registers, and one or more processing units such as arithmetic logic unit (ALU), floating-point unit (FPU), graphics processing unit (GPU), physical processing unit (PPU), digital system processing (DSP) unit, or combinations of these and / or other such processing units.

[0088] To enable user interaction with the computing system architecture 500, input device 516 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, a pen, and other such input devices. Output device 518 can also be one or more of a variety of output mechanisms known to those skilled in the art, including but not limited to monitors, speakers, printers, haptic devices, and other such output devices. In some cases, a multimodal system allows the user to provide multiple types of input to communicate with the computing system architecture 500. In some embodiments, input device 516 and / or output device 518 can be coupled to computing device 502 using a remotely connected device, such as a communication interface like the network interface 520 described herein. In such embodiments, the communication interface can control and manage the inputs and outputs received from the attached input device 516 and / or output device 518. It is conceivable that there are no limitations on operation on any particular hardware arrangement, and therefore the basic features described herein can be readily replaced by other hardware, software, or firmware arrangements as they evolve.

[0089] In some embodiments, storage device 510 may be described as non-volatile storage or non-volatile memory. Such non-volatile memory or non-volatile storage may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as magnetic tape cassettes, flash memory cards, solid-state storage devices, digital multifunction disks, magnetic tape cartridges, RAM, ROM, and mixtures thereof.

[0090] As described above, storage device 510 may include hardware and / or software services, such as service 512, that can control or configure processor 504 to perform one or more functions, including but not limited to the methods, processes, functions, systems, and services described herein in various embodiments. In some embodiments, the hardware or software service may be implemented as a module. As shown in the example computing system architecture 500, storage device 510 may be connected to other parts of computing device 502 using system connection 506. In some embodiments, hardware services or hardware modules performing functions, such as service 512, may include software components stored in a non-transitory computer-readable medium that, in conjunction with necessary hardware components such as processor 504, connection 506, cache 508, storage device 510, memory 514, input device 516, output device 518, etc., can perform functions such as those described herein.

[0091] The disclosed systems and services can be used in ways such as Figure 5 The example computing system illustrated executes using one or more components of the example computing system architecture 500. The example computing system may include a processor (e.g., a central processing unit), memory, non-volatile memory, and interface devices. The memory may store data and / or one or more sets of code, software, scripts, etc. The components of the computer system may be coupled together via a bus or through some other known or convenient device.

[0092] In some examples, a processor may be configured to execute some or all of the media device-related methods and systems described herein by means of, for example, executing code using a processor such as processor 504, wherein said code is stored in memory such as memory 514 described herein. One or more of a user device, provider server or system, database system, or other such device, service, or system may include some or all of the components of a computing system, such as one or more components using the example computing system architecture 500 shown herein. Figure 5 The example computing system shown is illustrated. It is conceivable that variations of such systems are considered to be within the scope of this disclosure.

[0093] This disclosure envisions computer systems taking any suitable physical form. By way of example, and not limitation, a computer system may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (such as, for example, a modular computer (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a tablet computer system, a wearable computer system or interface, an interactive self-service terminal, a mainframe, a computer system grid, a mobile phone, a personal digital representative (PDA), a server, or a combination of two or more of these. Where appropriate, a computer system may include one or more computer systems; may be single or distributed; span multiple locations; span multiple machines; and / or reside in a cloud computing system that may include one or more cloud components in one or more networks as described herein by the associated computing resource provider 528. Where appropriate, one or more computer systems may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example, and not limitation, one or more computer systems may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.

[0094] Processor 504 can be a conventional microprocessor, such as an Intel® microprocessor, an AMD® microprocessor, a Motorola® microprocessor, or another such microprocessor. Those skilled in the art will recognize that the terms "machine-readable (storage) medium" or "computer-readable (storage) medium" include any type of device accessible to the processor.

[0095] Memory 514 may be coupled to processor 504 via, for example, a connector such as connector 506 or a bus. As used herein, a connector or bus such as connector 506 is a communication system for transferring data between components within computing device 502, and in some embodiments, may be used for transferring data between computing devices. Connector 506 may be a data bus, memory bus, system bus, or other such data transfer mechanism. Examples of such connectors include, but are not limited to, Industry Standard Architecture (ISA) buses, Extended ISA (EISA) buses, Parallel AT Accessory (PATA) buses (e.g., Integrated Drive Electronics (IDE) or Extended IDE (EIDE) buses), or various types of Parallel Component Interconnect (PCI) buses (e.g., PCI, PCIe, PCI-104, etc.).

[0096] Memory 514 may include RAM, including but not limited to dynamic RAM (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile random access memory (NVRAM), and other types of RAM. DRAM may include error correction codes (EEC). Memory may also include ROM, including but not limited to programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, mask ROM (MROM), and other types of ROM. Memory 514 may also include magnetic or optical data storage media, including read-only (e.g., CD ROM and DVD ROM) or others (e.g., CD or DVD). Memory may be local, remote, or distributed.

[0097] As described above, connector 506 (or bus) may also couple processor 504 to storage device 510, which may include non-volatile memory or storage, drive units, and / or the like. In some embodiments, the non-volatile memory or storage is a floppy disk or hard disk, a magneto-optical disk, an optical disk, a ROM (e.g., CD-ROM, DVD-ROM, EPROM, or EEPROM), a magnetic card or optical card, or another form of data storage. Some of this data may be written to memory during software execution in the computer system via a direct memory access procedure. The non-volatile memory or storage may be local, remote, or distributed. In some embodiments, the non-volatile memory or storage is optional. It is conceivable that a computing system with all applicable data available in memory can be created. A typical computer system will generally include at least one processor, memory, and devices (e.g., buses) that couple the memory to the processor.

[0098] Software and / or data associated with the software may be stored in non-volatile memory and / or drive units. In some implementations (e.g., for large programs), it may not be possible to store the entire program and / or data in memory at any one time. In such implementations, the program and / or data may be moved into and out of memory from, for example, an additional storage device (such as storage device 510). However, it should be understood that, for the purpose of software operation, it may be moved to a computer-readable location suitable for processing, and for illustrative purposes, this location is referred to herein as memory. Even when the software is moved to memory for execution, the processor may utilize hardware registers to store values ​​associated with the software, as well as a local cache, ideally for accelerating execution. As used herein, when a software program is referred to as “implemented in a computer-readable medium,” it is assumed that the software program is stored in any known or convenient location (from non-volatile memory to hardware registers). When at least one value associated with the program is stored in a processor-readable register, the processor is considered “configured to execute the program.”

[0099] Connection 506 can also couple processor 504 to a network interface device such as network interface 520. The interface may include one or more modems or other such network interfaces, including but not limited to those described herein. It should be understood that network interface 520 may be considered part of computing device 502 or may be separate from computing device 502. Network interface 520 may include one or more analog modems, Integrated Services Digital Network (ISDN) modems, cable modems, token ring interfaces, satellite transmission interfaces, or other interfaces for coupling a computer system to other computer systems. In some embodiments, network interface 520 may include one or more input and / or output (I / O) devices. I / O devices may include, by way of example and not limitation, input devices such as input device 516 and / or output devices such as output device 518. For example, network interface 520 may include a keyboard, mouse, printer, scanner, display device, and other such components. Other examples of input and output devices are described herein. In some embodiments, the communication interface device may be implemented as a complete and stand-alone computing device.

[0100] In operation, a computer system can be controlled by operating system software, including a file management system such as a disk operating system. An example of operating system software with an associated file management system is the Windows® operating system family and its associated file management system. Another example of operating system software with its associated file management system is the Linux™ operating system and its associated file management system, including but not limited to various types and implementations of the Linux® operating system and its associated file management system. The file management system can be stored in non-volatile memory and / or drive units and can enable the processor to perform various actions required by the operating system to input and output data and store data in memory, including storing files on non-volatile memory and / or drive units. Other types of operating systems, such as, for example, MacOS®, other types of UNIX® operating systems (e.g., BSD), are also conceivable. ™ and its derivatives, Xenix ™ SunOS ™ Mobile operating systems (e.g., iOS® and its variants, Chrome®, Ubuntu Touch®, watchOS®, Windows 10 Mobile®, Blackberry® OS, etc.) and real-time operating systems (e.g., VxWorks®, QNX®, eCos®, RTLinux®, etc.) may be considered within the scope of this disclosure. It is conceivable that the names of operating systems, mobile operating systems, real-time operating systems, languages, and devices listed herein may be registered trademarks, service marks, or designs of various associated entities.

[0101] In some embodiments, computing device 502 may be connected to one or more additional computing devices, such as computing device 524, via network 522 using a connection such as network interface 520. In such embodiments, computing device 524 may perform one or more services 526 to perform one or more functions under the control of or on behalf of programs and / or services operating on computing device 502. In some embodiments, computing devices such as computing device 524 may include one or more components of the type described in conjunction with computing device 502, including but not limited to processors such as processor 504, connections such as connection 506, caches such as cache 508, storage devices such as storage device 510, memory such as memory 514, input devices such as input device 516, and output devices such as output device 518. In such embodiments, computing device 524 may perform functions such as those described herein in conjunction with computing device 502. In some embodiments, computing device 502 may be connected to multiple computing devices, such as computing device 524, wherein each computing device may also be connected to multiple computing devices, such as computing device 524. Such embodiments may be referred to herein as distributed computing environments.

[0102] Network 522 can be any network, including the Internet, intranet, extranet, cellular network, Wi-Fi network, local area network (LAN), wide area network (WAN), satellite network, Bluetooth® network, virtual private network (VPN), public switched telephone network, infrared (IR) network, Internet of Things (IoT) network, or any other such network or combination of networks. Communication via Network 522 can be wired, wireless, or a combination thereof. Communication via Network 522 can be conducted via a variety of communication protocols, including but not limited to Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), protocols at each layer of the Open Systems Interconnection (OSI) model, File Transfer Protocol (FTP), Universal Plug and Play (UPnP), Network File System (NFS), Server Message Block (SMB), Universal Internet File System (CIFS), and other such communication protocols.

[0103] Communication via network 522, within computing device 502, within computing device 524, or within computing resource provider 528 may include information, which may also be referred to herein as content. Information may include text, graphics, audio, video, haptic feedback, and / or any other information that may be provided to a user of a computing device such as computing device 502. In some embodiments, information may be delivered using transport protocols such as Hypertext Markup Language (HTML), Extensible Markup Language (XML), JavaScript®, Cascading Style Sheets (CSS), JavaScript® Object Notation (JSON), and other such protocols and / or structured languages. Information may first be processed by computing device 502 and presented to a user of computing device 502 in a form perceptible through visual, auditory, olfactory, gustatory, tactile, or other such mechanisms. In some embodiments, communication via network 522 may be received and / or processed by a computing device configured as a server. Such communication may be sent and received using PHP: Hypertext Preprocessor (“PHP”), Python ™ Ruby, Perl® and its variants, Java®, HTML, XML or another such server-side processing language.

[0104] In some implementations, computing devices 502 and / or 524 may connect to computing resource provider 528 via network 522 using network interfaces such as those described herein (e.g., network interface 520). In such implementations, one or more systems (e.g., services 530 and 532) hosted within computing resource provider 528 (also referred to herein as "within the computing resource provider environment") may perform one or more services to perform one or more functions under the control of or on behalf of programs and / or services operating on computing devices 502 and / or 524. Systems such as services 530 and 532 may include one or more computing devices such as those described herein to execute computer code to perform said one or more functions under the control of or on behalf of programs and / or services operating on computing devices 502 and / or 524.

[0105] For example, when the amount of data on computing device 502 exceeds the capacity of storage device 510, computing resource provider 528 can provide services operating on service 530 to store data for computing device 502. In another example, computing resource provider 528 can provide services such as: first, instantiating a virtual machine (VM) on service 532; using the VM to access data stored on service 532; performing one or more operations on the data; and providing the results of the one or more operations to computing device 502. Such operations (e.g., data storage and VM instantiation) may be referred to herein as "in the cloud," "within a cloud computing environment," or "within a managed virtual machine environment," and computing resource provider 528 may also be referred to herein as "the cloud." Examples of such computing resource providers include, but are not limited to, Amazon® Web Services (AWS®), Microsoft Azure®, IBM Cloud®, Google Cloud®, Oracle Cloud®, etc.

[0106] Services provided by computing resource provider 528 include, but are not limited to, data analytics, data storage, archival storage, big data storage, virtual computing (including various scalable VM architectures), blockchain services, containers (e.g., application encapsulation), database services, development environments (including sandbox development environments), e-commerce solutions, gaming services, media and content management services, security services, serverless hosting, and combinations thereof. Various technologies that facilitate such services include, but are not limited to, virtual machines, virtual storage, database services, system schedulers (e.g., hypervisors), resource management systems, and various types of short-term, medium-term, long-term, and archival storage devices.

[0107] It is conceivable that systems such as services 530 and 532 may represent versions of various services (e.g., service 512 or service 526) implemented under the control of computing devices 502 and / or 524. These implementations of various services may involve one or more virtualization technologies, such that, for example, to a user of computing device 502, when service 512 is being executed on, for example, service 530, it may appear as if that service is being executed on computing device 502. Similarly, it is conceivable that various services operating within the environment of computing resource provider 528 may be distributed across various systems within the environment, or partially distributed across computing devices 524 and / or 502.

[0108] The following examples illustrate various aspects of this disclosure. As used below, any reference to a series of examples should be understood as a separate reference to each of these examples (e.g., "Example 1 through Example 4" should be understood as "Example 1, Example 2, Example 4, or Example 4").

[0109] Example 1 is a method comprising: receiving an identification of a video segment from an automatic content identification service, wherein the video segment is being displayed by a display device; sending a notification to the display device based on the identification of the video segment, the notification including information associated with the video segment and a request for audio input; detecting one or more audio segments associated with the notification; and facilitating the presentation of an object associated with the video segment in response to the detection of the one or more audio segments.

[0110] Example 2 is a method of any of Examples 1 and 3 to 7, wherein facilitating the presentation of an object associated with the video clip includes displaying the object by the display device.

[0111] Example 3 is a method according to any one of Examples 1 to 2 and Examples 4 to 7, wherein facilitating the presentation of objects associated with the video segment includes executing an application by the display device, the application being configured to display new video segments associated with the video segment.

[0112] Example 4 is a method of any one of Examples 1 to 3 and Examples 5 to 7, further comprising: sending the one or more audio segments to a natural language processor, the natural language processor being configured to identify an intent corresponding to at least one of the one or more audio segments, wherein facilitating the presentation of an object associated with the video segment is also performed in response to the identification of the intent.

[0113] Example 5 is a method of any one of Examples 1 to 4 and Examples 6 to 7, wherein one or more audio segments are detected within a predetermined time interval, wherein the time interval begins when recognition of the video segment is received.

[0114] Example 6 is a method of any one of Examples 1 to 5 and Examples 7 to 14, wherein the notification is displayed adjacent to the video segment.

[0115] Example 7 is a method of any one of Examples 1 to 6 and Examples 8 to 14, wherein the one or more audio segments are received from a microphone embedded in a control device configured to operate the display device.

[0116] Example 8 is a method of any one of Examples 1 to 7 and Examples 9 to 14, wherein the one or more audio segments are received from a microphone embedded in the display device.

[0117] Example 9 is a method of any one of Examples 1 to 8 and Examples 9 to 14, wherein the object is a coupon associated with a featured product or service in the video clip.

[0118] Example 10 is a method of any of Examples 1 to 9 and Examples 10 to 14, wherein the object includes additional information associated with the video segment.

[0119] Example 11 is a method of any one of Examples 1 to 10 and Examples 12 to 14, the method further comprising: receiving an indication that the video segment is being presented again by the display device; and suppressing the sending of the notification to the display device.

[0120] Example 12 is a method of any one of Examples 1 to 11 and Examples 13 to 14, wherein facilitating the presentation of an object associated with the video segment includes: sending an instruction that, when received by the display device, causes the display device to display the object.

[0121] Example 13 is a method of any of Examples 1 to 12 and Example 14, wherein facilitating the presentation of an object associated with the video clip includes: sending an instruction that, when received by an application on a mobile device, causes the application to generate a push notification associated with the object.

[0122] Example 14 is a method of Examples 1 to 13, wherein facilitating the presentation of an object associated with the video clip includes sending the object via text message or email.

[0123] Example 15 is a system comprising: one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method described in any one of Examples 1-14.

[0124] Example 16 is a non-transitory computer-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of Examples 1 to 14.

[0125] Client devices, user devices, computer resource provider devices, network devices, and other devices can be computing systems including one or more integrated circuits, input devices, output devices, data storage devices, and / or network interfaces, etc. Integrated circuits can include, for example, one or more processors, volatile memory, and / or non-volatile memory, etc., as described herein. Input devices can include, for example, keyboards, mice, keypads, touch interfaces, microphones, cameras, and / or other types of input devices, including but not limited to those described herein. Output devices can include, for example, displays, speakers, haptic feedback systems, printers, and / or other types of output devices, including but not limited to those described herein. Data storage devices, such as hard disk drives or flash memory, enable computing devices to temporarily or permanently store data. Network interfaces, such as wireless or wired interfaces, enable computing devices to communicate with a network. Examples of computing devices (e.g., computing device 902) include, but are not limited to, desktop computers, laptop computers, server computers, handheld computers, tablet computers, smartphones, personal digital representatives, digital home representatives, wearable devices, smart devices, and combinations of these and / or other such computing devices, as well as machines and apparatuses in which computing devices are incorporated and / or virtually implemented.

[0126] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handsets, or integrated circuit devices with multiple uses, including applications in wireless communication handsets and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media such as those described herein. Additionally or alternatively, these techniques can be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures, and can be accessed, read, and / or executed by a computer, such as a propagating signal or wave.

[0127] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor), multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within a dedicated software or hardware module configured to implement a system for suspending database updates.

[0128] As used herein, the term "machine-readable medium" and its equivalents "machine-readable storage medium," "computer-readable medium," and "computer-readable storage medium" refer to media including, but not limited to, portable or non-portable storage devices, optical storage devices, removable or non-removable storage devices, and a variety of other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include transient electronic signals propagated via carrier waves and / or wireless or wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), solid-state drives (SSDs), flash memory, memory, or storage devices.

[0129] Machine-readable media or machine-readable storage media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including memory sharing, messaging, token passing, network transmission, etc. Further examples of machine-readable storage media, machine-readable media, or computer-readable (storage) media include, but are not limited to, recordable media (such as volatile and non-volatile storage devices, floppy disks and other removable disks, hard disk drives, optical disks (e.g., CDs, DVDs, etc.)), and transport media (such as digital and analog communication links).

[0130] It is conceivable that, while the examples in this document may describe or refer to a single machine-readable medium or machine-readable storage medium, the terms "machine-readable medium" and "machine-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) storing the one or more sets of instructions. The terms "machine-readable medium" and "machine-readable storage medium" should also be considered to include any medium capable of storing, encoding, or carrying a set of instructions that are executed by the system and cause the system to perform one or more methods or modules disclosed herein.

[0131] Certain portions of the specific implementations described herein can be presented based on algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate the substance of their work to others skilled in the art. Algorithms herein are, and generally are, considered as a self-consistent sequence of operations leading to a desired result. These operations are those that require physical manipulation of physical quantities. Typically, though not always necessary, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. For general reasons, it has proven convenient to sometimes refer to these signals as bits, values, elements, symbols, characters, items, numbers, etc.

[0132] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. As will be apparent from the following discussion, unless otherwise specifically stated, it should be understood that throughout the description, discussions using terms such as “processing,” “calculating,” “computing,” “determining,” “displaying,” or “generating” refer to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities within the registers and memory of the computer system and convert them into other data similarly represented as physical quantities within the computer system's memory or registers or other such information storage, transmission, or display devices.

[0133] It should also be noted that the various implementations can be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams (e.g., Figure 4(Example process). Although flowcharts, process diagrams, data flow diagrams, structure diagrams, or block diagrams can describe operations as a sequential process, many operations can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process shown in a diagram terminates when its operations are completed, but may have additional steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.

[0134] In some implementations, one or more implementations of algorithms such as those described herein can be implemented using machine learning or artificial intelligence algorithms. Such machine learning or artificial intelligence algorithms can be trained using supervised, unsupervised, reinforcement, or other such training techniques. For example, one of a variety of machine learning algorithms can be used to analyze a dataset to identify correlations between different elements of that dataset without supervision and feedback (e.g., unsupervised training techniques). Machine learning data analysis algorithms can also be trained using samples or real-time data to identify possible correlations. Such algorithms can include k-means clustering, fuzzy c-means (FCM), expectation-maximization (EM), hierarchical clustering, density-based spatial clustering with noise (DBSCAN), etc. Other examples of machine learning or artificial intelligence algorithms include, but are not limited to, genetic algorithms, backpropagation, reinforcement learning, decision trees, linear classification, artificial neural networks, anomaly detection, etc. More generally, machine learning or artificial intelligence methods can include regression analysis, dimensionality reduction, meta-learning, reinforcement learning, deep learning, and other such algorithms and / or methods. It is conceivable that the terms “machine learning” and “artificial intelligence” are often used interchangeably due to the overlap between these fields and the fact that many publicly available techniques and algorithms have similar approaches.

[0135] As an example of supervised training techniques, a dataset can be selected to train a machine learning model, facilitating the identification of correlations among the members of that dataset. The machine learning model can be evaluated to determine whether it is generating accurate correlations among the members of the dataset based on the sample inputs provided to it. Based on this evaluation, the machine learning model can be modified to increase the likelihood that it will identify the desired correlations. The machine learning model can also be dynamically trained by soliciting feedback from system users regarding the effectiveness of the correlations provided by the machine learning or artificial intelligence algorithm (i.e., supervision). This feedback can be used to improve the algorithm used to generate correlations (e.g., the feedback can be used to further train the machine learning or artificial intelligence to provide more accurate correlations).

[0136] The flowcharts, flowcharts, data flow diagrams, structural diagrams, or block diagrams of the various examples discussed herein can also be implemented in hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) used to perform the necessary tasks can be stored in computer-readable or machine-readable storage media (e.g., media for storing program code or code segments) such as those described herein. A processor implemented in an integrated circuit can perform the necessary tasks.

[0137] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the implementations disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above according to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Skilled artisans can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure.

[0138] However, it should be noted that the algorithms and demonstrations presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used in conjunction with the teachings and procedures herein, or it can be demonstrated that it is convenient to construct more specialized devices to perform certain examples. The necessary structures for various such systems will emerge from the descriptions below. Furthermore, these techniques are not described with reference to any particular programming language, and therefore various examples can be implemented using various programming languages.

[0139] In various implementations, the system operates as a standalone device or can be connected (e.g., networked) to other systems. In a networked deployment, the system can operate as a server or client system in a client-server network environment, or as a peer system in a peer-to-peer (or distributed) network environment.

[0140] The system can be a server computer, client computer, personal computer (PC), tablet PC (e.g., iPad®, Microsoft Surface®, Chromebook®, etc.), laptop computer, set-top box (STB), personal digital representative (PDA), mobile device (e.g., cellular phone, iPhone®, Android® device, Blackberry®, etc.), wearable device, embedded computer system, e-book reader, processor, telephone, web application device, network router, switch or bridge, or any system capable of executing a set of instructions (sequentially or otherwise) specifying the actions to be taken by the device. The system can also be a virtual system, such as a virtual version of one of the aforementioned devices, which may be hosted on another computer device such as computer device 902.

[0141] Typically, routines executed to implement the embodiments of this disclosure may be implemented as part of an operating system or a particular application, component, program, object, module, or a sequence of instructions referred to as a "computer program." A computer program typically includes one or more instructions set at different times in various memories and storage devices in a computer, which, when read and executed by one or more processing units or processors in the computer, cause the computer to perform operational elements relating to various aspects of this disclosure.

[0142] Furthermore, although the examples are described in the context of fully operational computers and computer systems, those skilled in the art will understand that various examples can be distributed as program objects in various forms, and this disclosure applies equally to any particular type of machine or computer-readable medium on which the distribution is actually performed.

[0143] In some cases, the operation of a storage device (e.g., a state change from binary one to binary zero or vice versa) may include transformations, such as physical transformations. For certain types of storage devices, such physical transformations may include physical transformations of an item into a different state or thing. For example, but not limited to, for some types of storage devices, state changes may involve the accumulation and storage of charge or the release of stored charge. Similarly, in other storage devices, state changes may include physical changes or transformations of magnetic orientation, or physical changes or transformations of molecular structure, such as from a crystalline state to an amorphous state or vice versa. The foregoing is not intended to be an exhaustive list of all examples in which state changes from binary one to binary zero or vice versa in storage devices may include transformations such as physical transformations. Rather, the foregoing is intended as an illustrative example.

[0144] Storage media can typically be nontransitory or include nontransitory devices. In this context, a nontransitory storage medium can include a tangible device, meaning that the device has a concrete physical form, although the device can change its physical state. Thus, for example, nontransitory means that the device remains tangible despite such changes in state.

[0145] The above description and accompanying drawings are illustrative and should not be construed as limiting or restricting the subject matter to the precise forms disclosed. Those skilled in the art will understand that many modifications and variations are possible in light of the above disclosure, and can be made without departing from the broader scope of the embodiments set forth herein. Numerous specific details have been described to provide a thorough understanding of this disclosure. However, in some cases, well-known or conventional details have not been described to avoid obscuring the description.

[0146] As used herein, when applied to modules of a system, the terms “connection,” “coupling,” or any variation thereof refer to any direct or indirect connection or coupling between two or more elements; the coupling connection between elements can be physical, logical, or any combination thereof. Furthermore, when the words “in this application,” “above,” “below,” and similar terms are used in this application, they should refer to the entire application, not any particular part of it. Where the context permits, singular or plural terms used in the above specific embodiments may also include either the plural or the singular, respectively. When referring to a list of two or more items, the word “or” covers all of the following interpretations: any item in the list, all items in the list, or any combination of items in the list.

[0147] As used herein, the terms “a” and “one” as well as “the” and other such singular references should be interpreted to include both singular and plural, unless otherwise indicated herein or the context clearly contradicts them.

[0148] As used herein, the terms “including,” “having,” “comprising,” and “containing” should be interpreted as open-ended (e.g., “comprising” should be interpreted as “including but not limited to”) unless otherwise indicated or the context clearly contradicts it.

[0149] As used herein, descriptions of ranges of values ​​are intended as a shorthand for each individual value falling within that range, unless otherwise indicated herein or explicitly contradicted by the context. Therefore, each individual value within that range is incorporated into the specification as if it were described separately herein.

[0150] As used herein, the terms “set” (e.g., “set of items”) and “subset” (e.g., “subset of items combined”) should be interpreted as non-empty sets that include one or more members, unless otherwise indicated or the context explicitly contradicts them. Furthermore, unless otherwise indicated or the context explicitly contradicts them, the term “subset” for a given set does not necessarily mean a proper subset of the given set, but rather that the subset and the set may include the same elements (i.e., the set and the subset may be identical).

[0151] As used herein, the use of relevance language such as “at least one of A, B, and C” should be interpreted as referring to one or more of A, B, and C (e.g., any of the following non-empty subsets of the set {A, B, C}: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, or {A, B, C}), unless otherwise indicated or the context explicitly contradicts it. Therefore, relevance language (such as “at least one of A, B, and C”) does not imply a requirement for at least one of A, at least one of B, and at least one of C.

[0152] As used herein, the use of example or exemplary language (e.g., "such as" or "as an example") is intended to illustrate the implementation more clearly and, unless otherwise required, does not impose a limitation on the scope. Such language in the specification should not be construed as indicating that any unclaimed element is necessary for the practice of the implementations described and claimed in this disclosure.

[0153] As used herein, when a component is described as being “configured” to perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0154] Those skilled in the art will understand that the disclosed subject matter may be implemented in other forms and manners not shown below. It should be understood that the use of relational terms (if any), such as first, second, top, and bottom, is only for distinguishing one entity or action from another, and does not necessarily require or imply any such actual relationship or order between these entities or actions.

[0155] Although procedures or blocks are presented in a given order, alternative implementations may execute routines with steps or employ systems with blocks in a different order, and certain procedures or blocks may be deleted, moved, added, subdivided, replaced, combined, and / or modified to provide alternatives or subcombinations. Each of these procedures or blocks can be implemented in a variety of different ways. Furthermore, while procedures or blocks are sometimes shown as executing sequentially, they may alternatively execute in parallel or at different times. Moreover, any specific numbers mentioned herein are merely examples: alternative implementations may take on different values ​​or ranges.

[0156] The teachings provided herein can be applied to other systems, not necessarily those described above. The elements and actions of the various examples described above can be combined to provide additional examples.

[0157] Any patents and applications mentioned above, as well as other references, including any references that may be listed in the accompanying application documents, are incorporated herein by reference. If necessary, aspects of this disclosure may be modified to incorporate the systems, functions, and concepts of the various references described above to provide further examples of this disclosure.

[0158] Based on the specific embodiments described above, these and other modifications can be made to these examples. While the above description illustrates certain examples and depicts the expected best practices, these teachings can be practiced in a variety of ways, no matter how detailed the foregoing text appears. The details of the system can vary significantly in their implementation details while still being covered by the subject matter disclosed herein. As noted above, specific terms used in describing certain features or aspects of this disclosure should not be construed as implying that such terms are redefined herein to be limited to any particular characteristic, feature, or aspect of the disclosure associated with that term. Generally, the terms used in the following claims should not be construed as limiting this disclosure to the specific implementations disclosed in the specification, unless such terms are expressly defined in the foregoing detailed embodiments section. Therefore, the actual scope of this disclosure includes not only the disclosed implementations but also all equivalent ways of practicing or implementing this disclosure under the claims.

[0159] While certain aspects of this disclosure are presented in certain of the following claim forms, the inventors contemplate that all aspects of this disclosure may exist in any number of claim forms. Any claim intended to be addressed under 45 USC §112(f) will begin with the phrase “means for”. Therefore, the applicant reserves the right to add supplementary claims after filing the application to pursue such supplementary claim forms for other aspects of this disclosure.

[0160] The terms used in this specification generally have their common meaning in the art, their meaning in the context of this disclosure, and their meaning in the specific context in which each term is used. Certain terms used to describe this disclosure are discussed above or elsewhere in the specification to provide additional guidance to practitioners regarding the description of this disclosure. For convenience, certain terms may be highlighted, for example, using capital letters, italics, and / or quotation marks. The use of highlighting does not affect the scope and meaning of the terms; in the same context, the scope and meaning of the terms are the same whether or not they are highlighted. It should be understood that the same element may be described in more than one way.

[0161] Therefore, alternative languages ​​and synonyms may be used for any one or more terms discussed herein, and no particular meaning is given to the terms as to whether they are stated or discussed herein. Synonyms for some terms are provided. The recitation of one or more synonyms does not preclude the use of other synonyms. Examples used anywhere in this specification (including examples of any terms discussed herein) are merely illustrative and are not intended to further limit the scope and meaning of this disclosure or any exemplary terms. Likewise, this disclosure is not limited to the various examples given in this specification.

[0162] Without intending to further limit the scope of this disclosure, examples of instruments, apparatus, methods, and related results are given based on the examples provided herein. Note that headings or subheadings may be used in the examples for the convenience of the reader, but this should in no way limit the scope of this disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In case of conflict, this document (including the definitions) shall prevail.

[0163] This description uses examples of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the art of data processing to effectively communicate the substance of their work to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent circuits, microcode, etc. Furthermore, it has sometimes proven convenient to arrange these operations as modules without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.

[0164] Any step, operation, or process described herein may be performed or implemented by one or more hardware or software modules, individually or in combination with other devices. In some examples, the software module is implemented as a computer program object, which includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0165] Examples may also relate to apparatus for performing the operations described herein. Such apparatus may be specifically constructed for the desired purpose, and / or it may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory, tangible, computer-readable storage medium, or in any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing system mentioned in the specification may include a single processor, or may be an architecture employing a multiprocessor design to increase computing power.

[0166] Examples may also involve objects generated by the computational processes described herein. Such objects may include information resulting from the computational processes, wherein such information is stored on a non-transitory, tangible, computer-readable storage medium, and may include any implementation of computer program objects or other combinations of data described herein.

[0167] The language used in this specification has been chosen primarily for readability and pedagogical purposes and may not have been chosen to describe or define the subject matter. Therefore, the scope of this disclosure is not intended to be limited to this specific embodiment, but rather to any claims published in any application based thereon. Thus, the exemplary disclosure is intended to illustrate, and not limit, the scope of the subject matter set forth in the appended claims.

[0168] Specific details have been provided in the foregoing description to offer an understanding of various implementations of the systems and components used in the context-connected system. However, those skilled in the art will understand that the implementations described above can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as block diagrams to avoid obscuring the implementation with unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the implementation.

[0169] The foregoing detailed description of the present technology has been presented for illustrative and descriptive purposes. It is not intended to be exhaustive or to limit the technology to the precise forms disclosed. In view of the foregoing teachings, many modifications and variations are possible. The described embodiments were chosen to best explain the principles of the present technology, its practical application, and to enable others skilled in the art to utilize various embodiments of the present technology and various modifications suitable for the intended particular use. The scope of the present technology is intended to be defined by the claims.

Claims

1. A method, the method comprising: Receive identification of a video segment from an automatic content identification service, wherein the video segment is being displayed by a display device; Based on the identification of the video segment, a notification is sent to the display device, the notification including information associated with the video segment and a request for audio input; Detect one or more audio segments associated with the notification; and In response to the detection of one or more audio segments, the rendering of objects associated with the video segments is facilitated.

2. The method according to claim 1, wherein, Facilitating the presentation of objects associated with the video clip includes displaying the objects by the display device.

3. The method according to claim 1, wherein, Facilitating the presentation of objects associated with the video clip includes an application executed by the display device, the application being configured to display new video clips associated with the video clip.

4. The method according to claim 1, further comprising: The one or more audio segments are sent to a natural language processor, which is configured to identify an intent corresponding to at least one of the one or more audio segments, wherein facilitating the presentation of an object associated with the video segment is also performed in response to the identification of the intent.

5. The method according to claim 1, wherein, The one or more audio segments are detected within a predetermined time interval, wherein the time interval begins upon receiving recognition of the video segment.

6. The method according to claim 1, wherein, The notification is displayed adjacent to the video clip.

7. The method according to claim 1, wherein, The one or more audio segments are received from a microphone embedded in the display device.

8. A system comprising: One or more processors; as well as A non-transitory computer-readable medium storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations including: Receive identification of a video segment from an automatic content identification service, wherein the video segment is being displayed by a display device; Based on the identification of the video segment, a notification is sent to the display device, the notification including information associated with the video segment and a request for audio input; Detect one or more audio segments associated with the notification; and In response to the detection of one or more audio segments, the rendering of objects associated with the video segments is facilitated.

9. The system according to claim 8, wherein, Facilitating the presentation of objects associated with the video clip includes displaying the objects by the display device.

10. The system according to claim 8, wherein, Facilitating the presentation of objects associated with the video clip includes an application executed by the display device, the application being configured to display new video clips associated with the video clip.

11. The system according to claim 8, wherein, The operation also includes: The one or more audio segments are sent to a natural language processor, which is configured to identify an intent corresponding to at least one of the one or more audio segments, wherein facilitating the presentation of an object associated with the video segment is also performed in response to the identification of the intent.

12. The system according to claim 8, wherein, The one or more audio segments are detected within a predetermined time interval, wherein the time interval begins upon receiving recognition of the video segment.

13. The system according to claim 8, wherein, The notification is displayed adjacent to the video clip.

14. The system according to claim 8, wherein, The one or more audio segments are received from a microphone embedded in the display device.

15. A non-transitory computer-readable medium storing instructions, which, when executed by one or more processors, cause the one or more processors to perform operations including: Receive the identification of video segments from the automatic content identification service, whereby... The video clip is being displayed by a display device; Based on the identification of the video segment, a notification is sent to the display device, the notification including information associated with the video segment and a request for audio input; Detect one or more audio segments associated with the notification; and In response to the detection of one or more audio segments, the rendering of objects associated with the video segments is facilitated.

16. The non-transitory computer-readable medium according to claim 15, wherein, Facilitating the presentation of objects associated with the video clip includes displaying the objects by the display device.

17. The non-transitory computer-readable medium according to claim 15, wherein, Facilitating the presentation of objects associated with the video clip includes an application executed by the display device, the application being configured to display new video clips associated with the video clip.

18. The non-transitory computer-readable medium according to claim 15, wherein, The operation also includes: The one or more audio segments are sent to a natural language processor, which is configured to identify an intent corresponding to at least one of the one or more audio segments, wherein facilitating the presentation of an object associated with the video segment is also performed in response to the identification of the intent.

19. The non-transitory computer-readable medium according to claim 15, wherein, The one or more audio segments are detected within a predetermined time interval, wherein the time interval begins upon receiving recognition of the video segment.

20. The non-transitory computer-readable medium according to claim 15, wherein, The one or more audio segments are received from a microphone embedded in the display device.