Keyword-based object insertion into video streams

The system enhances live media streams by inserting objects related to detected keywords, addressing the lack of enhancement time in simultaneous audio-video playback, thereby improving viewer retention and content relevance.

JP2025531099APending Publication Date: 2025-09-19QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514374
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-19
Filing Date
2023-09-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Computing devices often lack sufficient time to enhance live media streams with relevant content due to simultaneous audio and video playback, leading to a degraded viewer experience.

Method used

A system and method for keyword-based object insertion into video streams, utilizing keyword detection and adaptive classification to insert objects associated with detected keywords into the video stream, either from a database or generated on the fly.

Benefits of technology

Enhances viewer retention and adds relevant content to live media streams by inserting objects that maintain viewer interest, such as background images or local restaurant information, improving the overall viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531099000001_ABST
    Figure 2025531099000001_ABST
Patent Text Reader

Abstract

The device includes one or more processors configured to obtain an audio stream and detect one or more keywords within the audio stream. The one or more processors are also configured to adaptively classify one or more objects associated with the one or more keywords. The one or more processors are further configured to insert the one or more objects into a video stream.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS)

[0001] This application claims the benefit of priority from commonly owned U.S. Non-Provisional Patent Application No. 17 / 933,425, filed September 19, 2022, the contents of which are expressly incorporated by reference in their entirety into this specification.

[0002] FIELD OF THE DISCLOSURE

[0002] This disclosure relates generally to inserting one or more objects into a video stream based on one or more keywords.

[0003] 2. Description of Related Art

[0003] Advances in technology have led to smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices that are small, lightweight, and easily carried by users, including wireless telephones such as mobile phones and smartphones, tablet computers, and laptop computers. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications such as web browser applications that can be used to access the Internet. Thus, these devices can contain significant computing power.

[0004]

[0004] Such computing devices often incorporate functionality for receiving audio captured by a microphone and playing the audio through a speaker. The devices often also incorporate functionality for displaying video captured by a camera. In some examples, the devices incorporate functionality for receiving a media stream and playing the audio of the media stream through a speaker simultaneously with displaying the video of the media stream. With live media streams being displayed simultaneously with reception or capture, there is typically insufficient time for a user to edit the video before display. Thus, enhancements that might otherwise be made to improve viewer retention, such as adding related content, are not available when presenting a live media stream, which can result in a degraded viewer experience. Summary of the Invention

[0005] According to one implementation of the present disclosure, a device includes one or more processors configured to acquire an audio stream and detect one or more keywords within the audio stream. The one or more processors are also configured to adaptively classify one or more objects associated with the one or more keywords. The one or more processors are further configured to insert the one or more objects into a video stream.

[0006] According to another implementation of the present disclosure, the method includes acquiring an audio stream at a device. The method also includes detecting, at the device, one or more keywords in the audio stream. The method further includes adaptively classifying, at the device, one or more objects associated with the one or more keywords. The method also includes inserting, at the device, the one or more objects into a video stream.

[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain an audio stream and detect one or more keywords within the audio stream. The instructions, when executed by the one or more processors, also cause the one or more processors to adaptively classify one or more objects associated with the one or more keywords. The instructions, when executed by the one or more processors, further cause the one or more processors to insert the one or more objects into a video stream.

[0008] According to another implementation of the present disclosure, an apparatus includes means for acquiring an audio stream. The apparatus also includes means for detecting one or more keywords in the audio stream. The apparatus further includes means for adaptively classifying one or more objects associated with the one or more keywords. The apparatus also includes means for inserting the one or more objects into a video stream.

[0009]

[0009] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections: Brief Description of the Drawings, Form for Implementing the Invention, and Claims. [Brief explanation of the drawings]

[0010] [Figure 1]

[0010] A block diagram of certain exemplary aspects of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure, and an exemplary example of keyword-based object insertion into a video stream. [Figure 2]

[0011] A diagram of a specific implementation of a method for keyword-based object insertion into a video stream and an illustrative example of keyword-based object insertion into a video stream that may be performed by the device of Figure 1 in accordance with some examples of the present disclosure. [Figure 3]

[0012] A diagram of another specific implementation of a method for keyword-based object insertion into a video stream according to some examples of the present disclosure, and a diagram of an illustrative example of keyword-based object insertion into a video stream that can be performed by the device of Figure 1. [Figure 4]

[0013] 2 is a diagram of an exemplary embodiment of an example keyword detection unit of the system of FIG. 1, in accordance with some examples of the present disclosure. [Figure 5]

[0014] 1 is a diagram of an example aspect of operations associated with keyword detection, according to some examples of the present disclosure. [Figure 6]

[0015] 10 is a diagram of another particular implementation of a method of object generation that may be performed by the device of FIG. 1 and an illustrative example of object generation, according to some examples of the present disclosure. [Figure 7]

[0016] 2 is a diagram of an illustrative embodiment of one or more example components of the object determination unit of the system of FIG. 1, in accordance with some examples of the present disclosure. [Figure 8]

[0017] 1 is a diagram of an example aspect of operations associated with object classification, according to some examples of the present disclosure. [Figure 9A]

[0018] 10 is a diagram of another example aspect of operations associated with the object classification neural network of the system of FIG. 1, in accordance with some examples of the present disclosure. [Figure 9B]

[0019] 2 is a diagram of an example aspect of operations associated with feature extraction performed by the object classification neural network of the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 9C]

[0020] 2A-2C are diagrams of example aspects of operations associated with classification and probability distributions performed by the object classification neural network of the system of FIG. 1, in accordance with some examples of the present disclosure. [Figure 10A]

[0021] 2A-2C are diagrams of specific implementations of a method of determining an insertion position that may be performed by the device of FIG. 1 and examples of determining an insertion position, according to some examples of the present disclosure. [Figure 10B]

[0022] 2 is a diagram of an example aspect of operations performed by a position neural network of the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 11]

[0023] 1 is a block diagram of an example aspect of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 12]

[0024] FIG. 1 is a block diagram of another example aspect of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 13]

[0025] FIG. 1 is a block diagram of another example aspect of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 14]

[0026] FIG. 1 is a block diagram of another example aspect of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 15]

[0027] FIG. 1 is a block diagram of another example aspect of a system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 16]

[0028] 2 is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 17]

[0029] 1 illustrates an example integrated circuit operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 18]

[0030] FIG. 1 is a diagram of a mobile device operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 19]

[0031] FIG. 1 is a diagram of a headset operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 20]

[0032] FIG. 1 is a diagram of a wearable electronic device operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 21]

[0033] FIG. 1 illustrates a diagram of a voice-controlled speaker system operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 22]

[0034] 1 is a diagram of a camera operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 23]

[0035] 1 is a diagram of a headset, such as an extended reality headset, operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 24]

[0036] FIG. 1 is a diagram of an extended reality glasses device operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 25]

[0037] 1 is a diagram of a first example of a vehicle operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 26]

[0038] FIG. 10 is a diagram of a second example vehicle operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure. [Figure 27]

[0039] 2 is a diagram of a particular implementation of a method of keyword-based object insertion into a video stream that may be performed by the device of FIG. 1 in accordance with some examples of the present disclosure. [Figure 28]

[0040] 1 is a block diagram of a particular illustrative example of a device operable to perform keyword-based object insertion into a video stream, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011]

[0041] Computing devices often incorporate functionality for playing back media streams by providing an audio stream to a speaker while simultaneously displaying a video stream. With live media streams being displayed as they are received or captured, there is typically insufficient time for a user to perform enhancements to improve viewer retention, add content related to the video stream, etc., before display.

[0012]

[0042] A system and method for performing keyword-based object insertion into a video stream are disclosed. For example, a video stream updater performs keyword detection in an audio stream to generate keywords and determines whether a database contains any objects associated with the keywords. In response to determining that the database contains objects associated with the keywords, the video stream updater inserts the objects into the video stream. Alternatively, in response to determining that the database does not contain any objects associated with the keywords, the video stream updater applies an object generation neural network to the keywords to generate objects associated with the keywords and inserts the objects into the video stream. Optionally, in some examples, the video stream updater designates the newly generated objects as associated with the keywords and adds the objects to the database. Thus, the video stream updater can enhance the video stream with existing or newly generated objects associated with keywords detected in the audio stream.

[0013]

[0043] This enhancement can improve viewer retention, add relevant content, and so on. For example, maintaining viewer interest during playback of a video stream of a person speaking at a podium can be a challenge. Adding objects to the video stream can make the video stream more interesting to viewers during playback. Illustratively, adding a background image showing the results of tree planting to a live media stream discussing climate change can increase viewer retention of the live media stream. As another example, adding an image of a local restaurant to a video stream about traveling to an area that has the same type of food served at the restaurant can entice viewers to visit the local restaurant or result in an increase in orders placed with the restaurant. In some examples, enhancements can be made to the video stream based on an audio stream acquired separately from the video stream. Illustratively, the video stream can be updated based on user utterances contained in the audio stream received from one or more microphones.

[0014]

[0044] Certain aspects of the present disclosure are described below with reference to the drawings. In this description, common features are indicated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 illustrates a device 130 that includes one or more processors ("processor" 102 in FIG. 1), indicating that in some implementations, the device 130 includes a single processor 102 and in other implementations, the device 130 includes multiple processors 102.

[0015]

[0045] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When features as a group or type are referred to herein (e.g., when no specific one of the features is referenced), the reference number is used without the distinguishing letter. However, when one specific feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, with reference to FIG. 1, multiple objects are shown and associated with reference numbers 122A and 122B. When referring to a specific one of these objects, such as object 122A, the distinguishing letter "A" is used. However, when referring to any one of these objects or these objects as a group, the reference number 122 is used without the distinguishing letter.

[0016]

[0046] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” denotes an example, implementation, and / or aspect and should not be construed as limiting or as indicating a preferred or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, component, operation, etc., do not in themselves indicate a priority or order of the element with respect to other elements, but merely distinguish the element from other elements having the same name (apart from the use of ordinal terms). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.

[0017]

[0047] As used herein, "coupled" may include "communicatively coupled," "electrically coupled," or "physically coupled," as well as (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronic components, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.

[0018]

[0048] In this disclosure, terms such as "determining," "calculating," "estimating," "shifting," "adjusting," and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generating," "calculating," "estimating," "using," "selecting," "accessing," and "determining" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device.

[0019]

[0049] Referring to Figure 1, a particular exemplary embodiment of a system 100 is disclosed. The system 100 is configured to perform keyword-based object insertion into a video stream. Figure 1 also illustrates examples 190 and 192 of keyword-based object insertion into a video stream.

[0020]

[0050] The system 100 includes a device 130 including one or more processors 102 coupled to a memory 132 and a database 150. The one or more processors 102 include a video stream updater 110 configured to perform keyword-based object insertion into a video stream 136, and the memory 132 is configured to store instructions 109 executable by the one or more processors 102 to implement the functionality described with reference to the video stream updater 110.

[0021]

[0051] The video stream updater 110 includes a keyword detection unit 112 coupled to the object insertion unit 116 via an object determination unit 114. Optionally, in some implementations, the video stream updater 110 also includes a position determination unit 170 coupled to the object insertion unit 116.

[0022]

[0052] Device 130 also includes database 150 accessible to one or more processors 102. However, in other aspects, database 150 may be external to device 130, such as stored on a storage device, a network device, cloud-based storage, or a combination thereof. Database 150 is configured to store a set of objects 122, such as object 122A, object 122B, one or more additional objects, or a combination thereof. As used herein, "object" refers to a visual digital element such as one or more of an image, clip art, a photograph, a drawing, a graphics interchange format (GIF) file, a portable network graphics (PNG) file, or a video clip, as illustrative, non-limiting examples. "Objects" are primarily or entirely image-based and therefore distinct from text-based additions such as subtitles.

[0023]

[0053] In some implementations, database 150 is configured to store object keyword data 124 that indicates one or more keywords 120, if any, associated with one or more objects 122. In a particular example, object keyword data 124 indicates that object 122A (e.g., an image of the Statue of Liberty) is associated with one or more keywords 120A (e.g., "New York" and "Statue of Liberty"). In another example, object keyword data 124 indicates that object 122B (e.g., clip art depicting a clock) is associated with one or more keywords 120B (e.g., "clock," "alarm," "time").

[0024]

[0054] The video stream updater 110 is configured to process the audio stream 134 to detect one or more keywords 180 in the audio stream 134 and to insert objects associated with the detected keywords 180 into the video stream 136. In some examples, the media stream (e.g., a live media stream) includes the audio stream 134 and the video stream 136, as further described with reference to FIG. 11 . Optionally, in some examples, at least one of the audio stream 134 or the video stream 136 corresponds to decoded data generated by a decoder by decoding encoded data received from another device, as further described with reference to FIG. 12 . Optionally, in some examples, the video stream updater 110 is configured to receive the audio stream 134 from one or more microphones coupled to the device 130, as further described with reference to FIG. 13 . Optionally, in some examples, the video stream updater 110 is configured to receive the video stream 136 from one or more cameras coupled to the device 130, as further described with reference to FIG. 14 . Optionally, in some embodiments, audio stream 134 is obtained separately from video stream 136. For example, audio stream 134 is received from one or more microphones coupled to device 130, and video stream 136 is received from another device or generated at device 130, as further described with reference to at least Figures 13, 23, and 26.

[0025]

[0055] To illustrate, the keyword detection unit 112 is configured to determine one or more detected keywords 180 in at least a portion of the audio stream 134, as further described with reference to FIG. 5. As used herein, a "keyword" may refer to a single word or a phrase including multiple words. In some implementations, the keyword detection unit 112 is configured to apply a keyword detection neural network 160 to at least a portion of the audio stream 134 to generate the one or more detected keywords 180, as further described with reference to FIG. 4.

[0026]

[0056] The object determination unit 114 is configured to determine (e.g., select or generate) one or more objects 182 associated with the one or more detected keywords 180. The object determination unit 114 is configured to select one or more of the objects 122 stored in the database 150 that are indicated by the object keyword data 124 as being associated with the one or more detected keywords 180 for inclusion in the one or more objects 182. In certain aspects, the selected objects correspond to existing objects and pre-categorized objects associated with the one or more detected keywords 180.

[0027]

[0057] The object determination unit 114 includes an adaptive classifier 144 configured to adaptively classify one or more objects 182 associated with one or more detected keywords 180. Classifying the object 182 includes generating the object 182 (e.g., a newly generated object) based on the one or more detected keywords 180, performing classification of the object 182 to designate the object 182 as associated with the one or more keywords 180 (e.g., the newly classified object), and / or determining whether any of the keywords 120 match any of the keywords 120. In some aspects, the adaptive classifier 144 is configured to refrain from classifying the object 182 in response to determining that existing and pre-classified objects are associated with at least one of the one or more detected keywords 180. Alternatively, the adaptive classifier 144 is configured to classify (e.g., generate, perform classification, or both) an object 182 in response to determining that none of the existing objects are indicated by the object keyword data 124 as being associated with any of the one or more detected keywords 180.

[0028]

[0058] In some aspects, the adaptive classifier 144 includes an object generation neural network 140, an object classification neural network 142, or both. The object generation neural network 140 is configured to generate objects 122 (e.g., newly generated objects) associated with one or more objects 182. For example, the object generation neural network 140 is configured to process one or more detected keywords 180 (e.g., "alarm clock") to generate one or more objects 122 (e.g., clock clip art) associated with the one or more detected keywords 180, as further described with reference to FIGS. 6 and 7. The adaptive classifier 144 is configured to add the one or more objects 122 (e.g., newly generated objects) to the one or more objects 182 associated with the one or more detected keywords 180. In certain aspects, the adaptive classifier 144 is configured to update the object keyword data 124 to indicate that one or more objects 122 (e.g., newly generated objects) are associated with one or more keywords 120 (e.g., one or more detected keywords 180).

[0029]

[0059] The object classification neural network 142 is configured to classify objects 122 (e.g., existing objects) stored in the database 150. For example, the object classification neural network 142 is configured to process an object 122A (e.g., an image of the Statue of Liberty) to generate one or more keywords 120A (e.g., "New York" and "Statue of Liberty") associated with the object 122A, as further described with reference to Figures 9A-9C. As another example, the object classification neural network 142 is configured to process an object 122B (e.g., clip art of a clock) to generate one or more keywords 120B (e.g., "clock," "alarm," and "time"). The adaptive classifier 144 is configured to update the object keyword data 124 to indicate that the object 122A (e.g., an image of the Statue of Liberty) and the object 122B (e.g., clip art of a clock) are associated with one or more keywords 120A (e.g., "New York" and "Statue of Liberty") and one or more keywords 120B (e.g., "clock," "alarm," and "time"), respectively.

[0030]

[0060] After generating (e.g., updating) one or more keywords 120 associated with the set of objects 122, the adaptive classifier 144 is configured to determine whether the set of objects 122 includes at least one object 122 associated with one or more detected keywords 180. In response to determining that at least one of the one or more keywords 120A (e.g., "New York" and "Statue of Liberty") matches at least one of the one or more detected keywords 180 (e.g., "New York City"), the adaptive classifier 144 is configured to add the object 122A (e.g., a newly classified object) to the one or more objects 182 associated with the one or more detected keywords 180.

[0031]

[0061] In some aspects, the adaptive classifier 144 determines that the object 122 is associated with one or more detected keywords 180 in response to determining that the object keyword data 124 indicates that the object 122 is associated with at least one keyword 120 that matches at least one of the one or more detected keywords 180.

[0032]

[0062] In some implementations, adaptive classifier 144 is configured to determine that keyword 120 matches detected keyword 180 in response to determining that keyword 120 is the same as detected keyword 180 or that keyword 120 is a synonym of detected keyword 180. Optionally, in some implementations, adaptive classifier 144 is configured to generate a first vector representing keyword 120 and generate a second vector representing detected keyword 180. In these implementations, adaptive classifier 144 is configured to determine that keyword 120 matches detected keyword 180 in response to determining that the vector distance between the first vector and the second vector is less than a distance threshold.

[0033]

[0063] The adaptive classifier 144 is configured to adaptively classify one or more objects 182 associated with one or more detected keywords 180. For example, in particular implementations, the adaptive classifier 144 is configured to refrain from classifying one or more objects 182 in response to selecting one or more of the objects 122 (e.g., existing objects and pre-classified objects) stored in the database 150 for inclusion in the one or more objects 182. Alternatively, the adaptive classifier 144 is configured to classify one or more objects 182 associated with one or more detected keywords 180 in response to determining that none of the objects 122 (e.g., existing objects and pre-classified objects) are associated with the one or more detected keywords 180.

[0034]

[0064] In some examples, classifying the one or more objects 182 includes using the object generation neural network 140 to generate at least one of the one or more objects 182 (e.g., a newly generated object) associated with at least one of the one or more detected keywords 180. In some examples, classifying the one or more objects 182 includes using the object classification neural network 142 to designate one or more of the objects 122 (e.g., a newly classified object) as associated with the one or more keywords 120, and adding to the one or more objects 182 at least one of the objects 122 having keywords 120 that match the at least one detected keyword 180.

[0035]

[0065] Optionally, in some examples, adaptive classifier 144 uses object generation neural network 140, but not object classification neural network 142, to classify one or more objects 182. Illustratively, in these examples, adaptive classifier 144 includes object generation neural network 140, and object classification neural network 142 may be deactivated or, optionally, omitted from adaptive classifier 144.

[0036]

[0066] Optionally, in some examples, adaptive classifier 144 uses object classification neural network 142, and not object generation neural network 140, to classify one or more objects 182. Illustratively, in these examples, adaptive classifier 144 includes object classification neural network 142, and object generation neural network 140 may be deactivated or, optionally, omitted from adaptive classifier 144.

[0037]

[0067] Optionally, in some examples, adaptive classifier 144 uses object generation neural network 140 and object classification neural network 142 to classify one or more objects 182. Illustratively, in these examples, adaptive classifier 144 includes object generation neural network 140 and object classification neural network 142.

[0038]

[0068] Optionally, in some examples, adaptive classifier 144 uses object generation neural network 140 in response to determining that using object classification neural network 142 did not result in any of objects 122 being classified as associated with one or more detected keywords 180. Illustratively, in these examples, object generation neural network 140 is adaptively used based on the results of using object classification neural network 142.

[0039]

[0069] The adaptive classifier 144 is configured to provide one or more objects 182 associated with the one or more detected keywords 180 to the object insertion unit 116. The one or more objects 182 include one or more existing and pre-classified objects selected by the adaptive classifier 144, one or more objects newly generated by the object generation neural network 140, one or more objects newly classified by the object classification neural network 142, or a combination thereof. Optionally, in some implementations, the adaptive classifier 144 is also configured to provide the one or more objects 182 (or at least type information of the one or more objects 182) to the location determination unit 170.

[0040]

[0070] Optionally, in some implementations, the position determination unit 170 is configured to determine one or more insertion positions 164 and provide the one or more insertion positions 164 to the object insertion unit 116. In some implementations, the position determination unit 170 is configured to determine the one or more insertion positions 164 based at least in part on an object type of the one or more objects 182, as further described with reference to FIGS. 2-3. In some implementations, the position determination unit 170 is configured to apply a position neural network 162 to at least a portion of the video stream 136 to determine the one or more insertion positions 164, as further described with reference to FIG. 10.

[0041]

[0071] In certain aspects, the insertion locations 164 correspond to particular locations within an image frame of the video stream 136 (e.g., background, foreground, top, bottom, particular coordinates, etc.) or particular content within an image frame of the video stream 136 (e.g., a shirt, a photo frame, etc.). For example, during live media processing, one or more insertion locations 164 may indicate a location (e.g., foreground), content (e.g., a shirt), or both (e.g., a shirt in the foreground) within each of one or more particular frames of the video stream 136 that are presented substantially simultaneously as the corresponding detected keyword 180 is played. In some aspects, the one or more particular image frames are time-aligned with one or more audio frames of the audio stream 134 that are processed to determine the one or more detected keywords 180, as further described with reference to FIG. 16 .

[0042]

[0072] In some implementations without a position determination unit 170 for determining the one or more insertion positions 164, the one or more insertion positions 164 correspond to one or more predetermined insertion positions that may be used by the object insertion unit 116. Non-limiting illustrative examples of predetermined insertion positions include background, bottom right, scrolling at the bottom, or a combination thereof. In certain aspects, the one or more predetermined positions are based on default data, configuration settings, user input, or a combination thereof.

[0043]

[0073] The object insertion unit 116 is configured to insert one or more objects 182 at one or more insertion locations 164 in the video stream 136. In some examples, the object insertion unit 116 is configured to perform round-robin insertion of the one or more objects 182 when the one or more objects 182 include multiple objects to be inserted at the same insertion location 164. For example, the object insertion unit 116 performs round-robin insertion of a first subset of the one or more objects 182 (e.g., multiple images) at a first insertion location 164 (e.g., background), performs round-robin insertion of a second subset of the one or more objects 182 (e.g., multiple clip art, GIF files, etc.) at a second insertion location 164 (e.g., shirt), and so on. In another example, the object insertion unit 116 is configured, in response to determining that the one or more objects 182 include multiple objects and that the one or more insertion locations 164 include multiple locations, to insert object 122A of the one or more objects 182 into a first insertion location (e.g., the background) of the one or more insertion locations 164, insert object 122B of the one or more objects 182 into a second insertion location (e.g., bottom right), etc. The object insertion unit 116 is configured to output the video stream 136 (with the inserted one or more objects 182).

[0044]

[0074] In some implementations, device 130 corresponds to or is included in one of various types of devices. In an illustrative example, one or more processors 102 are integrated into a headset device, as further described with reference to FIG. 19. In other examples, one or more processors 102 are integrated into at least one of a mobile phone or tablet computing device such as described with reference to FIG. 18, a wearable electronic device such as described with reference to FIG. 20, a voice-controlled speaker system such as described with reference to FIG. 21, a camera device such as described with reference to FIG. 22, an extended reality (XR) headset such as described with reference to FIG. 23, or an XR glasses device such as described with reference to FIG. 24. In another illustrative example, one or more processors 102 are integrated into a vehicle, as further described with reference to FIGS. 25 and 26.

[0045]

[0075] During operation, video stream updater 110 obtains audio stream 134 and video stream 136. In certain aspects, audio stream 134 is a live stream that video stream updater 110 receives in real time from a microphone, a network device, another device, or a combination thereof. In certain aspects, video stream 136 is a live stream that video stream updater 110 receives in real time from a camera, a network device, another device, or a combination thereof.

[0046]

[0076] Optionally, in some implementations, the media stream (e.g., a live media stream) includes an audio stream 134 and a video stream 136, as further described with reference to FIG. 11. Optionally, in some implementations, at least one of the audio stream 134 or the video stream 136 corresponds to decoded data generated by a decoder by decoding encoded data received from another device, as further described with reference to FIG. 12. Optionally, in some implementations, the video stream updater 110 receives the audio stream 134 from one or more microphones coupled to the device 130, as further described with reference to FIG. 13. Optionally, in some implementations, the video stream updater 110 receives the video stream 136 from one or more cameras coupled to the device 130, as further described with reference to FIG. 14.

[0047]

[0077] The keyword detection unit 112 processes the audio stream 134 to determine one or more detected keywords 180 in the audio stream 134. In some examples, the keyword detection unit 112 processes a predetermined count of audio frames of the audio stream 134, audio frames of the audio stream 134 corresponding to a predetermined play time, or both. In particular aspects, the predetermined count of audio frames, the predetermined play time, or both are based on default data, configuration settings, user input, or a combination thereof.

[0048]

[0078] In some implementations, the keyword detection unit 112 omits (or does not use) the keyword detection unit 112 and instead uses speech recognition techniques to determine one or more words represented in the audio stream 134 and uses semantic analysis techniques to process the one or more words to determine one or more detected keywords 180. Optionally, in some implementations, the keyword detection unit 112 applies a keyword detection neural network 160 to process one or more audio frames of the audio stream 134 to determine (e.g., detect) one or more detected keywords 180 in the audio stream 134, as further described with reference to FIG. 4 . In some aspects, applying the keyword detection neural network 160 includes extracting acoustic features of the one or more audio frames to generate input values ​​and using the keyword detection neural network 160 to process the input values ​​to determine one or more detected keywords 180 that correspond to the acoustic features. The technical effects of applying the keyword detection neural network 160 may include using fewer resources (e.g., time, computing cycles, memory, or a combination thereof) compared to using speech recognition and semantic analysis, and improved accuracy in determining one or more detected keywords 180.

[0049]

[0079] In an example, the adaptive classifier 144 first performs a database search or lookup operation based on a comparison of one or more database keywords 120 and one or more detected keywords 180 to determine whether the set of objects 122 includes any objects associated with the one or more detected keywords 180. In response to determining that the set of objects 122 includes at least one object 122 associated with the one or more detected keywords 180, the adaptive classifier 144 refrains from classifying the one or more objects 182 associated with the one or more detected keywords 180.

[0050]

[0080] In example 190, keyword detection unit 112 determines one or more detected keywords 180 (e.g., "New York City") in audio stream 134 associated with video stream 136A. In response to determining that set of objects 122 includes object 122A (e.g., an image of the Statue of Liberty) associated with one or more keywords 120A (e.g., "New York" and "Statue of Liberty") and determining that at least one of the one or more keywords 120A matches at least one of the one or more detected keywords 180 (e.g., "New York City"), adaptive classifier 144 determines that object 122A is associated with the one or more detected keywords 180. In response to determining that object 122A is associated with the one or more detected keywords 180, adaptive classifier 144 includes object 122A in the one or more objects 182 and refrains from classifying the one or more objects 182 associated with the one or more detected keywords 180.

[0051]

[0081] In example 192, keyword detection unit 112 determines one or more detected keywords 180 (e.g., "alarm clock") in audio stream 134 associated with video stream 136A. Keyword detection unit 112 provides the one or more detected keywords 180 to adaptive classifier 144. In response to determining that set of objects 122 includes object 122B (e.g., clip art of a clock) associated with one or more keywords 120B (e.g., "clock," "alarm," and "time") and determining that at least one of the one or more keywords 120B matches at least one of the one or more detected keywords 180 (e.g., "alarm clock"), adaptive classifier 144 determines that object 122B is associated with the one or more detected keywords 180. In response to determining that object 122B is associated with one or more detected keywords 180, adaptive classifier 144 includes object 122B in one or more objects 182 and refrains from classifying one or more objects 182 associated with one or more detected keywords 180.

[0052]

[0082] In an alternative example where the database search or lookup operation does not detect any objects associated with one or more detected keywords 180 (e.g., "New York City" in example 190 or "alarm clock" in example 192), the adaptive classifier 144 classifies one or more objects 182 associated with the one or more detected keywords 180.

[0053]

[0083] Optionally, in some aspects, classifying the one or more objects 182 includes using an object classification neural network 142 to determine whether any of the set of objects 122 can be classified as associated with one or more detected keywords 180, as further described with reference to FIGS. 9A-9C . For example, using the object classification neural network 142 may include performing feature extraction of the objects 122 of the set of objects 122 to determine input values ​​representative of the objects 122, performing classification based on the input values ​​to determine one or more potential keywords that are likely to be associated with the objects 122, and generating a probability distribution indicating the likelihood that each of the one or more potential keywords is associated with the object 122. The adaptive classifier 144 designates one or more of the potential keywords as one or more keywords 120 associated with the object 122 based on the probability distribution. The adaptive classifier 144 updates the object keyword data 124 to indicate that the object 122 is associated with the one or more keywords 120 generated by the object classification neural network 142.

[0054]

[0084] As an example, the adaptive classifier 144 uses the object classification neural network 142 to process an object 122A (e.g., an image of the Statue of Liberty) and generate one or more keywords 120A (e.g., "New York" and "Statue of Liberty") associated with the object 122A. The adaptive classifier 144 updates the object keyword data 124 to indicate that the object 122A (e.g., an image of the Statue of Liberty) is associated with one or more keywords 120A (e.g., "New York" and "Statue of Liberty"). As another example, the adaptive classifier 144 uses the object classification neural network 142 to process an object 122B (e.g., clip art of a clock) and generate one or more keywords 120B (e.g., "clock," "alarm," and "time") associated with the object 122B. The adaptive classifier 144 updates the object keyword data 124 to indicate that the object 122B (e.g., clip art of the clock) is associated with one or more keywords 120B (e.g., "clock," "alarm," and "time").

[0055]

[0085] After updating the object keyword data 124 (e.g., after applying the object classification neural network 142 to each of the objects 122), the adaptive classifier 144 determines whether any object in the set of objects 122 is associated with one or more detected keywords 180. In response to determining that the object 122 is associated with one or more detected keywords 180, the adaptive classifier 144 adds the object 122 to one or more objects 182. In example 190, in response to determining that the object 122A (e.g., an image of the Statue of Liberty) is associated with one or more detected keywords 180 (e.g., "New York City"), the adaptive classifier 144 adds the object 122A to one or more objects 182. In example 192, in response to determining that the object 122B (e.g., clip art of a clock) is associated with one or more detected keywords 180 (e.g., "alarm clock"), the adaptive classifier 144 adds the object 122B to one or more objects 182. In some implementations, in response to determining that at least one object is included in one or more objects 182, the adaptive classifier 144 refrains from applying the object generation neural network 140 to determine one or more objects 182 associated with the one or more detected keywords 180.

[0056]

[0086] Optionally, in some implementations, classifying the one or more objects 182 includes applying an object generation neural network 140 to the one or more detected keywords 180 to generate the one or more objects 182. In some aspects, the adaptive classifier 144 applies the object generation neural network 140 in response to determining that the one or more objects 182 do not include an object. For example, in implementations that do not include applying an object classification neural network 142, or after applying the object classification neural network 142 but not detecting a matching object for the one or more detected keywords 180, the adaptive classifier 144 applies the object generation neural network 140.

[0057]

[0087] In some aspects, the object determination unit 114 applies the object classification neural network 142 to update the classification of the object 122, regardless of whether any existing objects are already included in the one or more objects 182. For example, in these aspects, the adaptive classifier 144 includes the object generation neural network 140, and the object classification neural network 142 is external to the adaptive classifier 144. Illustratively, in these aspects, classifying the one or more objects 182 includes selectively applying the object generation neural network 140 in response to determining that the object (e.g., an existing object) is not included in the one or more objects 182, while the object classification neural network 142 is applied regardless of whether any existing objects are already included in the one or more objects 182. In these aspects, resources are used to classify objects 122 in the database 150, and resources are selectively used to generate new objects.

[0058]

[0088] In some aspects, object determination unit 114 applies object generation neural network 140 to generate one or more additional objects for addition to one or more objects 182, regardless of whether any existing objects are already included in one or more objects 182. For example, in these aspects, adaptive classifier 144 includes object classification neural network 142, and object generation neural network 140 is external to adaptive classifier 144. Illustratively, in these aspects, classifying one or more objects 182 includes selectively applying object classification neural network 142 in response to determining that one or more objects 182 does not include any objects (e.g., existing objects and pre-classified objects), while object generation neural network 140 is applied regardless of whether any existing objects are already included in one or more objects 182. In these aspects, resources are used to add newly generated objects to database 150, and resources are selectively used to classify objects 122 in database 150 that are likely already classified.

[0059]

[0089] In some implementations, the object generation neural network 140 includes stacked generative adversarial networks (GANs). For example, applying the object generation neural network 140 to the detected keyword 180 includes generating an embedding representing the detected keyword 180, using a stage 1 GAN to generate a lower-resolution object based at least in part on the embedding, and using a stage 2 GAN to refine the lower-resolution object to generate a higher-resolution object, as further described with reference to FIG. 7 . The adaptive classifier 144 adds the newly generated higher-resolution object to the set of objects 122, updates the object keyword data 124 to indicate that the high-resolution object is associated with the detected keyword 180, and adds the newly generated object to one or more objects 182.

[0060]

[0090] In example 190, if none of the objects 122 are associated with one or more detected keywords 180 (e.g., "New York City"), the adaptive classifier 144 applies the object generation neural network 140 to the one or more detected keywords 180 (e.g., "New York City") to generate an object 122A (e.g., an image of the Statue of Liberty). The adaptive classifier 144 adds the object 122A (e.g., an image of the Statue of Liberty) to the set of objects 122 in the database 150, updates the object keyword data 124 to indicate that the object 122A is associated with one or more detected keywords 180 (e.g., "New York City"), and adds the object 122A to one or more objects 182.

[0061]

[0091] In example 192, if none of the objects 122 are associated with one or more detected keywords 180 (e.g., "alarm clock"), the adaptive classifier 144 applies the object generation neural network 140 to the one or more detected keywords 180 (e.g., "alarm clock") to generate object 122B (e.g., clock clip art). The adaptive classifier 144 adds object 122B (e.g., clock clip art) to the set of objects 122 in database 150, updates object keyword data 124 to indicate that object 122B is associated with one or more detected keywords 180 (e.g., "alarm clock"), and adds object 122B to one or more objects 182.

[0062]

[0092] The adaptive classifier 144 provides the one or more objects 182 to the object insertion unit 116 to insert the one or more objects 182 at one or more insertion locations 164 in the video stream 136. In some implementations, the one or more insertion locations 164 are predetermined. For example, the one or more insertion locations 164 are based on default data, configuration settings, user input, or a combination thereof. In some aspects, the predetermined insertion locations 164 may include location-specific locations such as the background, foreground, bottom, corner, center, etc. of the video frame.

[0063]

[0093] Optionally, in some implementations in which the video stream updater 110 includes a position determination unit 170, the adaptive classifier 144 also provides one or more objects 182 (or at least type information of the one or more objects 182) to the position determination unit 170 to dynamically determine one or more insertion positions 164. In some examples, the one or more insertion positions 164 may include location-specific positions such as background, foreground, top, center, bottom, corner, diagonal, or a combination thereof. In some examples, the one or more insertion positions 164 may include content-specific positions such as the front of a shirt, a stadium, a television, a whiteboard, a wall, a picture frame, another element drawn on the video frame, or a combination thereof. Use of the position determination unit 170 enables dynamic selection of elements within the content of the video stream 136 as the one or more insertion positions 164.

[0064]

[0094] In some implementations, the position determination unit 170 performs an image comparison between portions of video frames of the video stream 136 and stored images of potential locations to identify one or more insertion locations 164. Optionally, in some implementations in which the position determination unit 170 includes a position neural network 162, the position determination unit 170 applies the position neural network 162 to the video stream 136 to determine one or more insertion locations 164 within the video stream 136. For example, the position determination unit 170 applies the position neural network 162 to video frames of the video stream 136 to determine one or more insertion locations 164, as further described with reference to FIG. 10 . As compared to performing an image comparison to identify insertion locations, a technical effect of using the position neural network 162 to identify insertion locations may include using fewer resources (e.g., time, computing cycles, memory, or a combination thereof) in determining one or more insertion locations 164, having higher accuracy, or both.

[0065]

[0095] The object insertion unit 116 receives one or more objects 182 from the adaptive classifier 144. In some implementations, the object insertion unit 116 uses one or more predetermined positions as the one or more insertion positions 164. In other implementations, the object insertion unit 116 receives the one or more insertion positions 164 from the position determination unit 170.

[0066]

[0096] The object insertion unit 116 inserts one or more objects 182 at one or more insertion locations 164 in the video stream 136. In example 190, in response to determining that the insertion location 164 (e.g., background) is associated with an object 122A (e.g., an image of the Statue of Liberty) included in the one or more objects 182, the object insertion unit 116 inserts the object 122A as a background in one or more video frames of the video stream 136A to generate the video stream 136B. In example 192, in response to determining that the insertion location 164 (e.g., foreground) is associated with an object 122B (e.g., clock clip art) included in the one or more objects 182, the object insertion unit 116 inserts the object 122B as a foreground object in one or more video frames of the video stream 136A to generate the video stream 136B.

[0067]

[0097] In some implementations, the insertion position 164 corresponds to an element depicted in the video frame (e.g., the front of a shirt). The object insertion unit 116 inserts the object 122 at the insertion position 164 (e.g., the shirt), and the insertion position 164 can change its position in one or more video frames of the video stream 136A to follow the movement of the element. For example, the object insertion unit 116 determines a first position of the element (e.g., the shirt) in a first video frame and inserts the object 122 at the first position in the first video frame. As another example, the object insertion unit 116 determines a second position of the element (e.g., the shirt) in a second video frame and inserts the object 122 at the second position in the second video frame. If the element changes position between the first and second video frames, the first position may differ from the second position.

[0068]

[0098] In particular examples, the one or more objects 182 include a single object 122 and the one or more insertion locations 164 include multiple insertion locations 164. In some implementations, the object insertion unit 116 selects one of the insertion locations 164 for insertion of the object 122, while in other implementations, the object insertion unit 116 inserts copies of the object 122 at two or more of the multiple insertion locations 164 in the video stream 136. In some implementations, the object insertion unit 116 performs round-robin insertion of the object 122 at the multiple insertion locations 164. For example, the object insertion unit 116 inserts the object 122 at a first location of the multiple insertion locations 164 in a first set of video frames of the video stream 136, inserts the object 122 at a second location (rather than the first location) of the one or more insertion locations 164 in a second set of video frames of the video stream 136 that is different from the first set of video frames, and so on.

[0069]

[0099] In particular examples, the one or more objects 182 include multiple objects 122, and the one or more insertion locations 164 include multiple insertion locations 164. In some implementations, the object insertion unit 116 performs round-robin insertion of the multiple objects 122 at the multiple insertion locations 164. For example, the object insertion unit 116 inserts a first object 122 at a first insertion location 164 in a first set of video frames of the video stream 136, inserts a second object 122 (without the first object 122 at the first insertion location 164) in a second set of video frames of the video stream 136 that is different from the first set of video frames, and so on.

[0070]

[0100] In particular examples, the one or more objects 182 include multiple objects 122 and the one or more insertion locations 164 include a single insertion location 164. In some implementations, the object insertion unit 116 performs round-robin insertion of the multiple objects 122 at the single insertion location 164. For example, the object insertion unit 116 inserts a first object 122 at an insertion location 164 in a first set of video frames of the video stream 136, inserts a second object 122 (not the first object 122) at an insertion location 164 in a second set of video frames of the video stream 136 that is different from the first set of video frames, and so on.

[0071]

[0101] The object insertion unit 116 inserts one or more objects 182 into the video stream 136 and then outputs the video stream 136. In some implementations, the object insertion unit 116 provides the video stream 136 to a display device, a network device, a storage device, a cloud-based resource, or a combination thereof.

[0072]

[0102] Thus, system 100 enables the enrichment of video stream 136 with one or more objects 182 associated with one or more detected keywords 180. The enrichment of video stream 136 can improve viewer retention, create advertising opportunities, and the like. For example, adding an object to video stream 136 can make video stream 136 more interesting to viewers. To illustrate, adding object 122A (e.g., an image of the Statue of Liberty) can increase viewer retention for video stream 136 if audio stream 134 includes one or more detected keywords 180 (e.g., "New York City") associated with object 122A. In another example, object 122A can be associated with a related entity associated with one or more detected keywords 180 (e.g., an image of a restaurant in New York, a restaurant serving food associated with New York, another business selling New York-related food or services, a travel website, or a combination thereof).

[0073]

[0103] Although the video stream updater 110 is shown as including a position determination unit 170, in some other implementations, the position determination unit 170 is excluded from the video stream updater 110. For example, in implementations in which the position determination unit 170 is deactivated or omitted from the position determination unit 170, the object insertion unit 116 uses one or more predetermined locations as the one or more insertion locations 164. The use of the position determination unit 170 allows for dynamic determination of the one or more insertion locations 164, including content-specific insertion locations.

[0074]

[0104] Although the adaptive classifier 144 is shown as including the object generation neural network 140 and the object classification neural network 142, in some other implementations, the object generation neural network 140 or the object classification neural network 142 are excluded from the video stream updater 110. For example, adaptively classifying one or more objects 182 may include selectively applying the object generation neural network 140. In some implementations, the object determination unit 114 does not include the object classification neural network 142, and therefore resources are not used to reclassify objects that are likely already classified. In other implementations, the object determination unit 114 includes the object classification neural network 142 external to the adaptive classifier 144 such that objects are classified independently of the adaptive classifier 144. In an example, adaptively classifying one or more objects 182 may include selectively applying the object classification neural network 142. In some implementations, the object determination unit 114 does not include the object generation neural network 140, and therefore resources are not used to generate new objects. In other implementations, the object determination unit 114 includes the object generation neural network 140 external to the adaptive classifier 144, and therefore new objects are generated independently of the adaptive classifier 144.

[0075]

[0105] The use of object generation neural network 140 to generate new objects is provided as an illustrative example. In other examples, another type of object generator that does not include a neural network may be used instead of, or in addition to, object generation neural network 140 to generate new objects. The use of object classification neural network 142 to perform object classification is provided as an illustrative example. In other examples, another type of object classifier that does not include a neural network may be used instead of, or in addition to, object classification neural network 142 to perform object classification.

[0076]

[0106] Although the keyword detection unit 112 is shown as including the keyword detection neural network 160, in some other implementations, the keyword detection unit 112 may process the audio stream 134 to determine one or more detected keywords 180 independent of any neural network. For example, the keyword detection unit 112 may use speech analysis and semantic analysis to determine one or more detected keywords 180. Using a keyword detection neural network 160 (e.g., as compared to speech recognition and semantic analysis) may include using fewer resources (e.g., time, computing cycles, memory, or a combination thereof) in determining the one or more detected keywords 180, having higher accuracy, or both.

[0077]

[0107] Although the location determination unit 170 is shown as including the location neural network 162, in some other implementations, the location determination unit 170 can determine one or more insertion locations 164 independently of any neural network. For example, the location determination unit 170 can use image comparison to determine one or more insertion locations 164. Using a location neural network 162 (e.g., as compared to image comparison) can include using fewer resources (e.g., time, computing cycles, memory, or a combination thereof), having greater accuracy, or both, in determining the one or more insertion locations 164.

[0078]

[0108] 2, a particular implementation of a method 200 for keyword-based object insertion into a video stream and an example of keyword-based object insertion into a video stream 250 are shown. In particular aspects, one or more operations of the method 200 are performed by one or more of the keyword detection unit 112, the adaptive classifier 144, the position determination unit 170, the object insertion unit 116, the video stream updater 110, the one or more processors 102, the device 130, the system 100, or combinations thereof of FIG.

[0079]

[0109] The method 200 includes obtaining at least a portion of an audio stream, at 202. For example, the keyword detection unit 112 of FIG. 1 obtains one or more audio frames of the audio stream 134, as described with reference to FIG.

[0080]

[0110] The method 200 also includes detecting keywords, at 204. For example, the keyword detection unit 112 of Figure 1 processes one or more audio frames of the audio stream 134 to determine one or more detected keywords 180, as described with reference to Figure 1. In example 250, the keyword detection unit 112 processes the audio stream 134 to determine one or more detected keywords 180 (e.g., "New York City").

[0081]

[0111] The method 200 further includes determining at 206 whether any background objects correspond to keywords. In an example, the set of objects 122 of FIG. 1 correspond to background objects. Illustratively, each of the set of objects 122 may be inserted into the background of a video frame. The adaptive classifier 144 determines whether any object in the set of objects 122 corresponds to (e.g., is associated with) one or more detected keywords 180, as described with reference to FIG. 1.

[0082]

[0112] Method 200 also includes, in response to determining 206 that the background object corresponds to a keyword, inserting 208 the background object. For example, in response to determining that object 122A corresponds to one or more detected keywords 180, adaptive classifier 144 adds object 122A to one or more objects 182 associated with one or more detected keywords 180. In response to determining that object 122A is included in one or more objects 182 corresponding to one or more detected keywords 180, object insertion unit 116 inserts object 122A into video stream 136. In example 250, object insertion unit 116 inserts object 122A (e.g., an image of the Statue of Liberty) into video stream 136A to generate video stream 136B.

[0083]

[0113] In another method, in response to determining at 206 that there are no background objects corresponding to the keywords, method 200 includes retaining the original background at 210. For example, in response to adaptive classifier 144 determining that set of objects 122 does not include any background objects associated with one or more detected keywords 180, video stream updater 110 bypasses object insertion unit 116 and outputs one or more video frames of video stream 136 unchanged (e.g., without inserting any background objects into one or more video frames of video stream 136A).

[0084]

[0114] Thus, the method 200 enables enhancing the video stream 136 with background objects associated with one or more detected keywords 180. If a background object is not associated with one or more detected keywords 180, the background of the video stream 136 remains unchanged.

[0085]

[0115] 3, a particular implementation of a method 300 for keyword-based object insertion into a video stream and an example diagram 350 of keyword-based object insertion into a video stream are shown. In particular aspects, one or more operations of the method 300 are performed by one or more of the keyword detection unit 112, the adaptive classifier 144, the position determination unit 170, the object insertion unit 116, the video stream updater 110, the one or more processors 102, the device 130, the system 100, or combinations thereof of FIG.

[0086]

[0116] The method 300 includes obtaining at least a portion of an audio stream, at 302. For example, the keyword detection unit 112 of FIG. 1 obtains one or more audio frames of the audio stream 134, as described with reference to FIG.

[0087]

[0117] The method 300 also includes detecting keywords using a keyword detection neural network, at 304. For example, the keyword detection unit 112 of FIG. 1 uses the keyword detection neural network 160 to process one or more audio frames of the audio stream 134 to determine one or more detected keywords 180, as described with reference to FIG.

[0088]

[0118] The method 300 further includes determining whether the keywords map to any objects in the database, at 306. For example, the adaptive classifier 144 of FIG. 1 determines whether any objects in the set of objects 122 stored in the database 150 correspond to (e.g., are associated with) one or more detected keywords 180, as described with reference to FIG.

[0089]

[0119] In response to determining 306 that the keywords map to objects in the database, method 300 includes selecting 308 an object. For example, in response to determining that one or more detected keywords 180 (e.g., "New York City") are associated with object 122A (e.g., an image of the Statue of Liberty), adaptive classifier 144 of FIG. 1 selects object 122A to be added to one or more objects 182 associated with the one or more detected keywords 180, as described with reference to FIG. 1. As another example, in response to determining that one or more detected keywords 180 (e.g., "New York City") are associated with object 122B (e.g., clip art of an apple with the letters "NY"), adaptive classifier 144 of FIG. 1 selects object 122B to be added to one or more objects 182 associated with the one or more detected keywords 180, as described with reference to FIG. 1.

[0090]

[0120] In another method, in response to determining at 306 that the keyword does not map to any object in the database, method 300 includes using an object generation neural network to generate an object at 310. For example, in response to determining that none of the set of objects 122 are associated with one or more detected keywords 180, adaptive classifier 144 of FIG. 1 uses object generation neural network 140 as described with reference to FIG. 1 to generate object 122A (e.g., an image of the Statue of Liberty), object 122B (e.g., clip art of an apple with the letters "NY"), one or more additional objects, or a combination thereof. After generating the object at 310, method 300 includes adding the generated object to the database at 312 and selecting the object at 308. For example, the adaptive classifier 144 of FIG. 1 adds the object 122A, the object 122B, or both, to the database 150, as described with reference to FIG. 1, and selects the object 122A, the object 122B, or both, to add to one or more objects 182 associated with one or more detected keywords 180.

[0091]

[0121] Method 300 also includes determining, at 314, whether the object is a background type. For example, position determination unit 170 of FIG. 1 may determine whether object 122 included in one or more objects 182 is a background type. Based on the determination of whether object 122 is a background type, position determination unit 170 designates insertion position 164 for object 122, as described with reference to FIG. 1. In a particular example, position determination unit 170 of FIG. 1 designates first insertion position 164 (e.g., background) for object 122A in response to determining that object 122A of one or more objects 182 is a background type. As another example, position determination unit 170 designates second insertion position 164 (e.g., foreground) for object 122B in response to determining that object 122B of one or more objects 182 is not a background type. In some implementations, in response to determining that a position (e.g., background) of a video frame of the video stream 136 includes at least one object associated with one or more detected keywords 180, the position determination unit 170 selects another position (e.g., foreground) of the video frame as the insertion position 164.

[0092]

[0122] In a particular implementation, a first subset of the set of objects 122 may be stored in a background database and a second subset of the set of objects 122 may be stored in a foreground database, both of which may be included in database 150. In this implementation, location determination unit 170 determines that object 122A is of a background type in response to determining that object 122A is included in the background database. In an example, location determination unit 170 determines that object 122B is of a foreground type and not of a background type in response to determining that object 122B is included in the foreground database.

[0093]

[0123] In some implementations, the first subset and the second subset do not overlap. For example, an object 122 is included in either the background database or the foreground database, but not both. However, in other implementations, the first subset at least partially overlaps with the second subset. For example, a copy of the object 122 may be included in each of the background database and the foreground database.

[0094]

[0124] In particular implementations, the object type of object 122 is based on the file type of object 122 (e.g., an image file, a GIF file, a PNG file, etc.). For example, in response to determining that object 122A is an image file, position determination unit 170 determines that object 122A is of a background type. In another example, in response to determining that object 122B is not an image file (e.g., object 122B is a GIF file or a PNG file), position determination unit 170 determines that object 122B is of a foreground type and not of a background type.

[0095]

[0125] In particular implementations, the metadata of object 122 indicates whether object 122 is a background type or a foreground type. For example, in response to determining that the metadata of object 122A indicates that object 122A is a background type, position determination unit 170 determines that object 122A is a background type. As another example, in response to determining that the metadata of object 122B indicates that object 122B is a foreground type, position determination unit 170 determines that object 122B is a foreground type and not a background type.

[0096]

[0126] In response to determining at 314 that the object is of a background type, method 300 includes inserting the object into the background at 316. For example, in response to determining that first insertion position 164 (e.g., background) is designated for object 122A of one or more objects 182, object insertion unit 116 of FIG. 1 inserts object 122A at the first insertion position (e.g., background) in one or more video frames of video stream 136, as described with reference to FIG. 1.

[0097]

[0127] In another method, in response to determining at 314 that the object is not of a background type, method 300 includes inserting the object into the foreground at 318. For example, in response to determining that second insertion position 164 (e.g., foreground) has been designated for object 122B of one or more objects 182, object insertion unit 116 of FIG. 1 inserts object 122B at the second insertion position (e.g., foreground) in one or more video frames of video stream 136, as described with reference to FIG. 1.

[0098]

[0128] Thus, the method 300 allows for generating a new object 122 associated with one or more detected keywords 180 if none of the existing objects 122 are associated with the one or more detected keywords 180. The object 122 may be added to the background or foreground of the video stream 136 based on the object type of the object 122. The object type of the object 122 may be based on the file type, storage location, metadata, or a combination thereof of the object 122.

[0099]

[0129] In diagram 350, keyword detection unit 112 processes audio stream 134 using keyword detection neural network 160 to determine one or more detected keywords 180 (e.g., "New York City"). In a particular aspect, adaptive classifier 144 determines that object 122A (e.g., an image of the Statue of Liberty) is associated with one or more detected keywords 180 (e.g., "New York City") and adds object 122A to one or more objects 182. In response to determining that object 122A is of a background type, position determination unit 170 designates object 122A as being associated with first insertion position 164 (e.g., background). In response to determining that object 122A is associated with first insertion position 164 (e.g., background), object insertion unit 116 inserts object 122A into one or more video frames of video stream 136A to generate video stream 136B.

[0100]

[0130] According to an alternative aspect, adaptive classifier 144 may instead determine that object 122B (e.g., clip art of an apple with the letters “NY”) is associated with one or more detected keywords 180 (e.g., “New York City”) and adds object 122B to one or more objects 182. In response to determining that object 122B is not of the background type, position determination unit 170 designates object 122B as being associated with second insertion position 164 (e.g., foreground). In response to determining that object 122B is associated with second insertion position 164 (e.g., foreground), object insertion unit 116 inserts object 122B into one or more video frames of video stream 136A to generate video stream 136C.

[0101]

[0131] 4, a diagram 400 of an example implementation of the keyword detection unit 112 is shown. The keyword detection neural network 160 includes a speech recognition neural network 460 coupled to a keyword selector 464 via a potential keyword detector 462.

[0102]

[0132] The speech recognition neural network 460 is configured to process at least a portion of the audio stream 134 to generate one or more words 461 detected in the portion of the audio stream 134. In particular aspects, the speech recognition neural network 460 includes a recurrent neural network (RNN). In other aspects, the speech recognition neural network 460 may include another type of neural network.

[0103]

[0133] In an exemplary implementation, the speech recognition neural network 460 includes an encoder 402, an RNN transducer (RNN-T) 404, and a decoder 406. In a particular aspect, the encoder 402 is trained as a connectionist temporal classification (CTC) network. During training, the encoder 402 is configured to process one or more acoustic features 412 to predict phonemes 414, graphemes 416, and wordpieces 418 from a long short-term memory (LSTM) layer 420, an LSTM layer 422, and an LSTM layer 426, respectively. The encoder 402 includes a temporal convolutional layer 424 that reduces the encoder temporal sequence length (e.g., by a factor of three). The decoder 406 is trained to predict one or more wordpieces 458 by processing an input embedding 454 of one or more input wordpieces 452 using an LSTM layer 456. According to some aspects, the decoder 406 is trained to reduce the cross-entropy loss.

[0104]

[0134] The RNN-T 404 is configured to process one or more acoustic features 432 of at least a portion of the audio stream 134 using an LSTM layer 434, an LSTM layer 436, and an LSTM layer 440 to provide a first input (e.g., a first word piece) to a feedforward 448 (e.g., a feedforward layer). The RNN-T 404 also includes a temporal convolutional layer 438. The RNN-T 404 is configured to process an input embedding 444 of one or more input word pieces 442 using an LSTM layer 446 to provide a second input (e.g., a second word piece) to the feedforward 448. In certain aspects, the one or more acoustic features 432 correspond to real-time test data, and the one or more input word pieces 442 correspond to existing training data from which the speech recognition neural network 460 is trained. The feedforward 448 is configured to process the first input and the second input to generate word pieces 450. The speech recognition neural network 460 is configured to output one or more words 461 corresponding to one or more wordpieces 450 .

[0105]

[0135] The RNN-T 404 (e.g., the weights of the RNN-T 404) are initialized based on the encoder 402 (e.g., the trained encoder 402) and the decoder 406 (e.g., the trained decoder 406). In an example (shown by the dashed arrow in FIG. 4), the weights of the LSTM layer 434 are initialized based on the weights of the LSTM layer 420, the weights of the LSTM layer 436 are initialized based on the weights of the LSTM layer 422, the weights of the LSTM layer 440 are initialized based on the weights of the LSTM layer 426, the weights of the temporal convolutional layer 438 are initialized based on the weights of the temporal convolutional layer 424, the weights of the LSTM layer 446 are initialized based on the weights of the LSTM layer 456, the weights for generating the input embedding 444 are initialized based on the weights for generating the input embedding 454, or a combination thereof.

[0106]

[0136] LSTM layer 420 including five LSTM layers, LSTM layer 422 including five LSTM layers, LSTM layer 426 including two LSTM layers, and LSTM layer 456 including two LSTM layers are provided as illustrative examples. In other examples, LSTM layer 420, LSTM layer 422, LSTM layer 426, and LSTM layer 456 can include any number of LSTM layers. In certain aspects, LSTM layer 434, LSTM layer 436, LSTM layer 440, and LSTM layer 446 include the same number of LSTM layers as LSTM layer 420, LSTM layer 422, LSTM layer 426, and LSTM layer 456, respectively.

[0107]

[0137] Potential keyword detector 462 is configured to process one or more words 461 to determine one or more potential keywords 463, as will be further described with reference to Figure 5. Keyword selector 464 is configured to select one or more detected keywords 180 from the one or more potential keywords 463, as will be further described with reference to Figure 5.

[0108]

[0138] 5, a diagram 500 of an exemplary aspect of operations associated with keyword detection is shown. In particular aspects, keyword detection is performed by the keyword detection neural network 160, keyword detection unit 112, video stream updater 110, one or more processors 102, device 130, system 100, speech recognition neural network 460, potential keyword detector 462, keyword selector 464 of FIG. 4, or a combination thereof.

[0109]

[0139] The keyword detection neural network 160 obtains at least a portion of the audio stream 134 representing an utterance. The keyword detection neural network 160 uses a speech recognition neural network 460 on the portion of the audio stream 134 to detect one or more words 461 of the utterance (e.g., "wishes for you on your birthday, may you get whatever you ask for, may all your wishes come true on your birthday and always, happy birthday"), as described with reference to FIG.

[0110]

[0140] Potential keyword detector 462 performs a semantic analysis on one or more words 461 to identify one or more potential keywords 463 (e.g., "wish," "request," "birthday"). For example, potential keyword detector 462 ignores conjunctions, articles, prepositions, etc. in one or more words 461. One or more potential keywords 463 are underlined in one or more words 461 in diagram 500. In some implementations, one or more potential keywords 463 may include one or more words (e.g., "wish," "request," "birthday"), one or more phrases (e.g., "New York City," "alarm clock"), or a combination thereof.

[0111]

[0141] The keyword selector 464 selects at least one of the one or more potential keywords 463 (e.g., "wish," "request," "birthday") as one or more detected keywords 180 (e.g., "birthday"). In some implementations, the keyword selector 464 performs a semantic analysis on the one or more words 461 to determine which of the one or more potential keywords 463 correspond to a topic of the one or more words 461 and selects at least one of the one or more potential keywords 463 that correspond to the topic as one or more detected keywords 180. In a particular example, the keyword selector 464 selects the potential keyword 463 (e.g., "birthday") as one or more detected keywords 180 based at least in part on determining that the potential keyword 463 (e.g., "birthday") appears more frequently (e.g., three times) in the one or more words 461 compared to others of the one or more potential keywords 463. The keyword selector 464 selects at least one (e.g., "birthday") of one or more potential keywords 463 (e.g., "wish," "request," "birthday") corresponding to the topic of the one or more words 461 as one or more detected keywords 180.

[0112]

[0142] In particular aspects, object 122A (e.g., genie clip art) is associated with one or more keywords 120A (e.g., "wishes" and "genie"), and object 122B (e.g., an image with balloons and a birthday banner) is associated with one or more keywords 120B (e.g., "balloons," "birthday," "birthday banner"). In particular aspects, in response to determining that one or more keywords 120B (e.g., "balloons," "birthday," "birthday banner") match one or more detected keywords 180 (e.g., "birthday"), adaptive classifier 144 selects object 122B to be included in one or more objects 182 associated with one or more detected keywords 180, as described with reference to FIG. 1 .

[0113]

[0143] 6, shown is a method 600 of object generation, examples 650, 652, and 654. In certain aspects, one or more operations of method 600 are performed by object generation neural network 140, adaptive classifier 144, video stream updater 110, one or more processors 102, device 130, system 100, or combinations thereof of FIG.

[0114]

[0144] The method 600 includes, at 602, preprocessing. For example, the object generation neural network 140 of FIG. 1 preprocesses at least a portion of the audio stream 134. Illustratively, the preprocessing may include reducing noise in at least a portion of the audio stream 134 to increase the signal-to-noise ratio.

[0115]

[0145] The method 600 also includes feature extraction, at 604. For example, the object generation neural network 140 of FIG. 1 extracts features 605 (e.g., acoustic features) from the pre-processed portion of the audio stream 134.

[0116]

[0146] The method 600 further includes, at 606, performing semantic analysis using the language model. For example, the object generation neural network 140 of FIG. 1 may obtain one or more words 461 and one or more detected keywords 180 corresponding to the preprocessed portion of the audio stream 134. Illustratively, the object generation neural network 140 obtains the one or more words 461 based on the operation of the keyword detection unit 112. For example, the keyword detection unit 112 of FIG. 1 may perform preprocessing (e.g., noise removal, one or more additional enhancements, or a combination thereof) of at least a portion of the audio stream 134 to generate a preprocessed portion of the audio stream 134. The speech recognition neural network 460 of FIG. 4 may perform speech recognition on the preprocessed portion to generate one or more words 461 and provide the one or more words 461 to the potential keyword detector 462 of FIG. 4 and also to the object generation neural network 140.

[0117]

[0147] The object-generating neural network 140 may perform semantic analysis on features 605, one or more words 461 (e.g., "a flower with long pink petals and raised orange stamens"), one or more detected keywords 180 (e.g., "flower"), or a combination thereof to generate one or more descriptors 607 (e.g., "long pink petals; raised orange stamens"). In particular aspects, the object-generating neural network 140 performs the semantic analysis using a language model. In some examples, the object-generating neural network 140 performs semantic analysis on one or more detected keywords 180 (e.g., "New York") to determine one or more associated words (e.g., "Statue of Liberty," "harbor," etc.).

[0118]

[0148] The method 600 also includes generating 608 objects using an object generation network. For example, the adaptive classifier 144 of FIG. 1 processes one or more detected keywords 180 (e.g., "flower"), one or more descriptors 607 (e.g., "long pink petals" and "protuberant orange stamens"), associated words, or combinations thereof, using the object generation neural network 140 to generate one or more objects 182, as further described with reference to FIG. 7. Thus, the adaptive classifier 144 enables multiple words corresponding to the one or more detected keywords 180 to be used as inputs to the object generation neural network 140 (e.g., GAN) to generate objects 182 (e.g., images) associated with the multiple words. In some aspects, the object generation neural network 140 generates the object 182 (e.g., an image) in real time as the audio stream 134 of the live media stream is processed, such that the object 182 may be inserted into the video stream 136 substantially simultaneously (e.g., with imperceptible or nearly imperceptible delay) as one or more detected keywords 180 are determined. Optionally, in particular implementations, the object generation neural network 140 selects an existing object (e.g., an image of a flower) that matches one or more detected keywords 180 (e.g., “flower”) and modifies the existing object to generate the object 182. For example, the object generation neural network 140 modifies the existing object based on one or more detected keywords 180 (e.g., “flower”), one or more descriptors 607 (e.g., “long pink petals” and “raised orange stamens”), related words, or a combination thereof, to generate the object 182.

[0119]

[0149] In example 650, the adaptive classifier 144 uses the object generation neural network 140 to process one or more words 461 (e.g., "a flower with long pink petals and raised orange stamens") to generate an object 122 (e.g., a generated image of a flower with various pink petals, orange stamens, or a combination thereof). In example 652, the adaptive classifier 144 uses the object generation neural network 140 to process one or more words 461 ("blue bird") to generate an object 122 (e.g., a generated photorealistic image of a bird). In example 654, the adaptive classifier 144 uses the object generation neural network 140 to process one or more words 461 ("blue bird") to generate an object 122 (e.g., a generated clip art of a bird).

[0120]

[0150] 7, a diagram 700 of one example of one or more components of the object determination unit 114 is shown, including the object generation neural network 140. In certain aspects, the object determination unit 114 may include one or more additional components not shown for ease of illustration.

[0121]

[0151] In particular implementations, the object-generating neural network 140 includes a stacked GAN. Illustratively, the object-generating neural network 140 includes a stage 1 GAN coupled to a stage 2 GAN. The stage 1 GAN includes a refinement augmentor 704 coupled to a stage 1 discriminator 708 via a stage 1 generator 706. The stage 2 GAN includes a refinement augmentor 710 coupled to a stage 2 discriminator 714 via a stage 2 generator 712. The stage 1 GAN generates a low-resolution object based on the embeddings (702). The stage 2 GAN generates a high-resolution object (e.g., a photorealistic image) based on the embeddings 702 and the low-resolution object from the stage 1 GAN.

[0122]

[0152] The object generation neural network 140 generates an embedding of a text description 701 (e.g., "The bird is gray with a white chest and a very short beak") that represents at least a portion of the audio stream 134.

[0123]

number

[0124] 702. In some aspects, the text description 701 corresponds to one or more words 461 of FIG. 4, one or more detected keywords 180 of FIG. 1, one or more descriptors 607 of FIG. 6, related words, or a combination thereof. In particular implementations, some details of the text description 701 that are ignored by the Stage 1 GAN when generating the low-resolution object are considered by the Stage 2 GAN when generating the high-resolution object.

[0125]

[0153] The object generation neural network 140 provides embeddings 702 to each of a training augmentor 704, a stage 1 discriminator 708, a training augmentor 710, and a stage 2 discriminator 714. The training augmentor 704 uses a fully connected layer to generate the embeddings

[0126]

number

[0127] Processing 702 and Gaussian distribution

[0128]

number

[0129] 703 and variance (σ) 705 of

[0130]

number

[0131] is embedded

[0132]

number

[0133] 702 corresponds to a diagonal covariance matrix that is a function of σ0. The variance (σ0) 705 is

[0134]

number

[0135] The adjustment augmenter 704 uses a Gaussian distribution

[0136]

number

[0137] Gaussian-tuned variables for embeddings 702 sampled from

[0138]

number

[0139] 709 to capture the meaning of the embedding 702 with variations. For example, the adjustment variables

[0140]

number

[0141] 709 is based on the following formula:

[0142]

number

[0143]

[0154] where:

[0144]

number

[0145] is the adjustment variable

[0146]

number

[0147] 709 (e.g., adjustment vector), μ corresponds to mean (μ) 703, σ corresponds to variance (σ) 705,

[0148]

number

[0149] corresponds to element-wise multiplication,

[0150]

number

[0151] corresponds to a Gaussian distribution N(0,1). The adjustment augmenter 704 adjusts the adjustment variable

[0152]

number

[0153] 709 (eg, adjustment vectors) to the stage 1 generator 706 .

[0154]

[0155] The stage 1 generator 706 generates low-resolution objects 717 conditioned on the text description 701. For example, the stage 1 generator 706 may generate low-resolution objects 717 conditioned on the text description 701.

[0155]

number

[0156] 709 and a random variable (z) to generate a low-resolution object 717. In an example, the low-resolution object 717 (e.g., an image, clip art, GIF file, etc.) represents a primitive shape and a basic color. In a particular aspect, the random variable (z) corresponds to random noise (e.g., a -dimensional noise vector). In a particular example, the stage 1 generator 706 generates a low-resolution object 717 conditional on the tuning variable

[0157]

number

[0158] 709 and a random variable (z), and the combination is processed by a series of upsampling blocks 715 to produce a lower resolution object 717.

[0159] The stage 1 discriminator 708 uses the embeddings to generate the text tensor.

[0160]

number

[0161] The Stage 1 Discriminator 708 processes the low-resolution object 717 using a downsampling block 719 to generate an object filter map. The object filter map is combined with the text tensor to generate an object text tensor that is fed into a convolutional layer. A fully connected layer 721 with one node is used to generate a decision score.

[0162] In some aspects, the stage 2 generator 712 is designed as an encoder-decoder with a residual block 729. Similar to the adjustment augmenter 704, the adjustment augmenter 710 performs an embedding

[0163]

number

[0164] Process 702 and adjust variables

[0165]

number

[0166] 723, which is spatially replicated in stage 2 generator 712 to form a text tensor. The low-resolution object 717 is processed by a series of downsampling blocks (e.g., encoders) to generate an object filter map. The object filter map is combined with the text tensor to generate an object text tensor, which is processed by residual block 729. In particular aspects, residual block 729 is designed to learn a multi-model representation across the features of the low-resolution object 717 and the features of the text description 701. A series of upsampling blocks 731 (e.g., decoders) are used to generate a higher-resolution object 733. In particular examples, the higher-resolution object 733 corresponds to a photorealistic image.

[0167] The stage 2 discriminator 714 uses the embeddings to generate the text tensor.

[0168]

number

[0169] 702. The stage 2 discriminator 714 processes the higher resolution objects 733 using a downsampling block 735 to generate an object filter map. In certain embodiments, the counts in the downsampling block 735 are greater than the counts in the downsampling block 719 due to the larger size of the higher resolution objects 733 compared to the lower resolution objects 717. The object filter map is combined with the text tensor to generate an object text tensor that is fed to a convolutional layer. A fully connected layer 737 with one node is used to generate a decision score.

[0170] During the training phase, the stage 1 generator 706 and the stage 1 discriminator 708 may be trained together. During training, the stage 1 discriminator 708 is trained (e.g., modified based on feedback) to improve its ability to distinguish between images generated by the stage 1 generator 706 and real images with similar resolution, while the stage 1 generator 706 is trained to improve its ability to generate images that the stage 1 discriminator 708 classifies as real images. Similarly, the stage 2 generator 712 and the stage 2 discriminator 714 may be trained together. During training, the stage 2 discriminator 714 is trained (e.g., modified based on feedback) to improve its ability to distinguish between images generated by the stage 2 generator 712 and real images with similar resolution, while the stage 2 generator 712 is trained to improve its ability to generate images that the stage 2 discriminator 714 classifies as real images. In some implementations, after the training phase is complete, the stage 1 generator 706 and the stage 2 generator 712 can be used in the object generation neural network 140, while the stage 1 discriminator 708 and the stage 2 discriminator 714 can be omitted (or deactivated).

[0171] In certain aspects, the low-resolution object 717 corresponds to an image having basic colors and primitive shapes, and the high-resolution object 733 corresponds to a photorealistic image. In certain aspects, the low-resolution object 717 corresponds to a basic line drawing (e.g., no gradations of shading, single colors, or both), and the high-resolution object 733 corresponds to a detailed drawing (e.g., with gradations of shading, multiple colors, or both).

[0172] In certain aspects, the object determination unit 114 adds the higher resolution object 733 to the database 150 as the object 122A and updates the object keyword data 124 to indicate that the object 122A is associated with one or more keywords 120A (e.g., the text description 701). In certain aspects, the object determination unit 114 adds the low resolution object 717 to the database 150 as the object 122B and updates the object keyword data 124 to indicate that the object 122B is associated with one or more keywords 120B (e.g., the text description 701). In certain aspects, the object determination unit 114 adds the low resolution object 717, the high resolution object 733, or both to the one or more objects 182.

[0173] 8, there is shown a method of object classification 800. In certain aspects, one or more operations of the method 800 are performed by the object classification neural network 142, the object determination unit 114, the adaptive classifier 144, the video stream updater 110, one or more processors 102, the device 130, the system 100, or a combination thereof of FIG.

[0174]

[0163] Method 800 includes retrieving 802 the next object from the database. For example, adaptive classifier 144 of Figure 1 may select an initial object (e.g., object 122A) from database 150 during an initial iteration of a processing loop through all of the objects 122 in database 150, as described further below.

[0175]

[0164] The method 800 also includes determining whether the object is associated with any keywords, at 804. For example, the adaptive classifier 144 of Figure 1 determines whether the object keyword data 124 indicates any keywords 120 associated with the object 122A.

[0176]

[0165] In response to determining 804 that the object is associated with at least one keyword, method 800 includes determining 806 whether there are more objects in the database. For example, adaptive classifier 144 of FIG. 1 determines whether there are any additional objects 122 in database 150 in response to determining that object keyword data 124 indicates that object 122A is associated with one or more keywords 120A. Illustratively, adaptive classifier 144 analyzes objects 122 in sequence based on object identifiers to determine whether there are additional objects in database 150 that correspond to the next identifier following the identifier of object 122A. If there are no more unprocessed objects in the database, method 800 ends at 808. Alternatively, method 800 includes selecting 802 a next object from the database for the next iteration of the processing loop.

[0177]

[0166] In response to determining 804 that the object is not associated with any keywords, the method 800 includes applying 810 an object classification neural network to the object. For example, the adaptive classifier 144 of Figure 1 applies the object classification neural network 142 to the object 122A to generate one or more potential keywords in response to determining that the object keyword data 124 indicates that the object 122A is not associated with any keywords 120, as further described with reference to Figures 9A-9C.

[0178] Method 800 also includes, at 812, associating the object with the generated potential keyword having the highest probability score. For example, each of the potential keywords generated by object classification neural network 142 for the object may be associated with a score indicating the probability that the potential keyword matches the object. Adaptive classifier 144 may designate the keyword with the highest score among the potential keywords as keyword 120A and update object keyword data 124 to indicate that object 122A is associated with keyword 120A, as further described with reference to FIG.

[0179]

[0168] Referring to Figure 9A, a diagram 900 of an exemplary aspect of the operations associated with the object classification neural network 142 of Figure 1 is shown. The object classification neural network 142 is configured to perform feature extraction 902 on the object 122A to generate features 926, as will be further described with reference to Figure 9B.

[0180] 9C, the object classification neural network 142 is configured to perform classification 904 of the features 926 to generate a classification layer output 932. The object classification neural network 142 is configured to process the classification layer output 932 to determine a probability distribution 906 associated with the one or more potential keywords, and to select at least one of the one or more potential keywords as one or more keywords 120A based on the probability distribution 906.

[0181] 9B, a diagram of an exemplary embodiment of feature extraction 902 is shown. In a particular implementation, object classification neural network 142 includes a convolutional neural network (CNN) including multiple convolution stages 922 configured to generate an output feature map 924. The convolution stages 922 include a first set of convolution, ReLU, and pooling layers in a first stage 922A, a second set of convolution, ReLU, and pooling layers in a second stage 922B, and a third set of convolution, ReLU, and pooling layers in a third stage 922C. The output feature map 924 output from the third stage 922C is converted into a vector (e.g., a flattening layer) corresponding to features 926. Although three convolution stages 922 are shown, in other implementations, any other number of convolution stages 922 may be used for feature extraction.

[0182] 9C, a diagram of an exemplary aspect of determining the classification 904 and probability distribution 906 is shown. In a particular aspect, the object classification neural network 142 includes a fully connected layer 928, such as layer 928A, layer 928B, layer 928C, one or more additional layers, or a combination thereof. The object classification neural network 142 performs the classification 904 by processing the features 926 using the fully connected layer 928 to generate a classification layer output 932. For example, the output of the last layer 928D corresponds to the classification layer output 932.

[0183]

[0172] Object classification neural network 142 applies a softmax activation function 930 to classification layer output 932 to generate a probability distribution 906. For example, probability distribution 906 indicates the probability that one or more potential keywords 934 are associated with object 122A. Illustratively, probability distribution 906 indicates a first probability (e.g., 0.5), a second probability (e.g., 0.7), and a third probability (e.g., 0.1) of a first potential keyword 934 (e.g., "bird"), a second potential keyword 934 (e.g., "blue bird"), and a third potential keyword 934 (e.g., "white bird") associated with object 122A (e.g., an image of a blue bird), respectively.

[0184] Object classification neural network 142 selects at least one of one or more potential keywords 934 for inclusion in one or more keywords 120A associated with object 122A (e.g., an image of a blue bird) based on probability distribution 906. In the illustrated example, object classification neural network 142 selects second potential keyword 934 (e.g., “blue bird”) in response to determining that second potential keyword 934 (e.g., “blue bird”) is associated with the highest probability (e.g., 0.7) within probability distribution 906. In another implementation, object classification neural network 142 selects at least one of potential keywords 934 based on the selected one or more potential keywords having at least a threshold probability (e.g., 0.5), as indicated by probability distribution 906. For example, in response to determining that each of the first potential keyword 934 (e.g., “bird”) and the second potential keyword 934 (e.g., “blue bird”) is associated with a first probability (e.g., 0.5) and a second probability (e.g., 0.7), respectively, that are equal to or greater than a threshold probability (e.g., 0.5), the object classification neural network 142 selects the first potential keyword 934 (e.g., “bird”) and the second potential keyword 934 (e.g., “blue bird”) for inclusion in the one or more keywords 120A.

[0185] 10A, shown is a method 1000 and example 1050 of insertion location determination. In certain aspects, one or more operations of the method 1000 are performed by the location neural network 162, the location determination unit 170, the object insertion unit 116, the video stream updater 110, one or more processors 102, the device 130, the system 100, or a combination thereof of FIG.

[0186] Method 1000 includes applying a location neural network to a video frame at 1002. In example 1050, location determination unit 170 applies location neural network 162 to video frame 1036 of video stream 136 to generate features 1046, as further described with reference to FIG.

[0187] The method 1000 also includes performing segmentation, at 1022. For example, the position determination unit 170 performs segmentation based on the features 1046 to generate one or more segmentation masks 1048. In some aspects, performing segmentation includes applying a neural network to the features 1046 according to various techniques to generate the segment masks. Each segmentation mask 1048 corresponds to the outline of a segment of the video frame 1036 that corresponds to a region of interest, such as a person, a shirt, pants, a hat, a picture frame, a television, a sports field, one or more other types of regions of interest, or a combination thereof.

[0188] Method 1000 further includes applying masking, at 1024. For example, position determination unit 170 applies one or more segmentation masks 1048 to video frame 1036 to generate one or more segments 1050. Illustratively, position determination unit 170 applies a first segmentation mask 1048 to video frame 1036 to generate a first segment corresponding to a shirt, applies a second segmentation mask 1048 to video frame 1036 to generate a second segment corresponding to pants, and so on.

[0189] Method 1000 also includes applying detection at 1026. For example, position determination unit 170 performs detection to determine whether any of the one or more segments 1050 match a location criterion. Illustratively, the location criterion may indicate a valid insertion location in video stream 136, such as a person, a shirt, a stadium, or the like. In some examples, the location criterion is based on default data, configuration settings, user input, or a combination thereof. Position determination unit 170 generates detection data 1052 indicating whether any of the one or more segments 1050 match the location criterion. In certain aspects, position determination unit 170 generates detection data 1052 indicating the at least one segment in response to determining that at least one segment of one or more segments 1050 matches the location criterion.

[0190] Optionally, in some implementations, method 1000 includes applying detection for each of the one or more objects 182 based on an object type of the one or more objects 182. For example, one or more objects 182 include object 122A of a particular object type. In some implementations, the location criteria indicate valid locations associated with the object type. For example, the location criteria indicate a first valid location (e.g., shirt, hat, etc.) associated with a first object type (e.g., GIF, clip art, etc.), a second valid location (e.g., wall, stadium, etc.) associated with a second object type (e.g., image), etc. In response to determining that object 122A is the first object type, location determining unit 170 generates detection data 1052 indicating at least one of the one or more segments 1050 that matches the first valid location. Alternatively, in response to determining that object 122A is of the second object type, position determination unit 170 generates detection data 1052 indicating at least one of one or more segments 1050 that matches the second valid position.

[0191] In some implementations, the position criteria indicate that if the one or more objects 182 include an object 122 associated with a keyword 120 and another object associated with the keyword 120 is included in the background of the video frame, the object 122 is included in the foreground of the video frame. For example, in response to determining that the one or more objects 182 include an object 122A associated with one or more keywords 120A, that the video frame 1036 includes an object 122B associated with one or more keywords 120B at a first position (e.g., background), and that at least one of the one or more keywords 120A matches at least one of the one or more keywords 120B, the position determination unit 170 generates detection data 1052 that indicates at least one of the one or more segments 1050 that matches a second position (e.g., foreground) of the video frame 1036.

[0192]

[0181] The method 1000 further includes determining whether a location is identified, at 1008. For example, the location determination unit 170 determines whether the detection data 1052 indicates that any of the one or more segments 1050 meets the location criteria.

[0193] In response to determining at 1008 that a location has been identified, method 1000 includes designating an insertion location at 1010. In example 1050, location determination unit 170 designates segment 1050 (e.g., a shirt) as an insertion location 164 in response to determining that detection data 1052 indicates that segment 1050 satisfies the location criteria. In particular examples, detection data 1052 indicates that multiple segments 1050 satisfy the location criteria. In some aspects, location determination unit 170 selects one of the multiple segments 1050 for designation as the insertion location 164. In other examples, location determination unit 170 selects two or more (e.g., all) of the multiple segments 1050 for addition to the one or more insertion locations 164.

[0194]

[0183] In response to determining at 1008 that a position is not identified, the method 1000 includes skipping the insertion at 1012. For example, in response to determining that the detection data 1052 indicates that none of the segments 1050 match the position criteria, the position determination unit 170 generates a "no position" output indicating that no insertion position is selected. In this example, in response to receiving the no position output, the object insertion unit 116 outputs the video frame 1036 without inserting any object into the video frame 1036.

[0195] 10B, a diagram 1070 of an exemplary aspect of operations performed by the position neural network 162 of the position determination unit 170 is shown. In a particular aspect, the position neural network 162 includes a residual neural network (resnet), such as resnet 152. For example, the position neural network 162 includes multiple convolutional layers (e.g., CONV1, CONV2, etc.) and a pooling layer (“pool”) that are used to process the video frames 1036 and generate features 1046.

[0196] 11, a diagram of a system 1100 including a particular implementation of device 130 is shown. System 1100 is operable to perform keyword-based object insertion into a video stream. In certain aspects, system 100 of FIG. 1 includes one or more components of system 1100. Some components of device 130 of FIG. 1 are not shown in device 130 of FIG. 11 for ease of illustration. In some aspects, device 130 of FIG. 1 may include one or more of the components of device 130 shown in FIG. 11, one or more additional components, one or more fewer components, one or more different components, or a combination thereof.

[0197] System 1100 includes device 130 coupled to device 1130 and one or more display devices 1114. In certain aspects, device 1130 includes a computing device, a server, a network device, a storage device, a cloud storage device, a video camera, a communication device, a broadcast device, or a combination thereof. In certain aspects, one or more display devices 1114 include a touchscreen, a monitor, a television, a communication device, a playback device, a display screen, a vehicle, an XR device, or a combination thereof. In certain aspects, an XR device may include an augmented reality device, a mixed reality device, or a virtual reality device. One or more display devices 1114 are described as being external to device 130 as an illustrative example. In other examples, one or more display devices 1114 may be integrated into device 130.

[0198]

[0187] Device 130 includes a demultiplexer (demux) 1172 coupled to the video stream updater 110. Device 130 is configured to receive the media stream 1164 from device 1130. In an example, device 130 receives the media stream 1164 from device 1130 over a network. The network may include a wired network, a wireless network, or both.

[0199]

[0188] The demux 1172 demultiplexes the media stream 1164 to produce the audio stream 134 and the video stream 136. The demux 1172 provides the audio stream 134 to the keyword detection unit 112 and provides the video stream 136 to the position determination unit 170, the object insertion unit 116, or both. The video stream updater 110 updates the video stream 136 by inserting one or more objects 182 into one or more portions of the video stream 136, as described with reference to Figure 1.

[0200]

[0189] In certain aspects, media stream 1164 corresponds to a live media stream. Video stream updater 110 updates video stream 136 of the live media stream and provides video stream 136 (e.g., an updated version of video stream 136) to one or more display devices 1114, one or more storage devices, or a combination thereof.

[0201] In some examples, video stream updater 110 selectively updates a first portion of video stream 136, as described with reference to FIG. 1. Video stream updater 110 provides the first portion (e.g., after selective updating) to one or more display devices 1114, one or more storage devices, or a combination thereof. Optionally, in some aspects, device 130 outputs the updated portion of video stream 136 to one or more display devices 1114 while receiving subsequent portions of video stream 136 included in media stream 1164 from device 1130. Optionally, in some aspects, video stream updater 110 provides audio stream 134 to one or more speakers simultaneously with providing video stream 136 to one or more display devices 1114.

[0202] 12, there is shown a diagram of a system 1200 operable to perform keyword-based object insertion into a video stream. The system 1200 includes a device 130 coupled to a device 1206 and one or more display devices 1114.

[0203]

[0192] In certain aspects, device 1206 includes a computing device, a server, a network device, a storage device, a cloud storage device, a video camera, a communication device, a broadcast device, or a combination thereof. Device 130 includes a decoder 1270 coupled to the video stream updater 110 and configured to receive the encoded data 1262 from device 1206. In an example, device 130 receives the encoded data 1262 from device 1206 over a network. The network may include a wired network, a wireless network, or both.

[0204]

[0193] The decoder 1270 decodes the encoded data 1262 to generate decoded data 1272. In certain aspects, the decoded data 1272 includes the audio stream 134 and the video stream 136. In certain aspects, the decoded data 1272 includes one of the audio stream 134 or the video stream 136. In this aspect, the video stream updater 110 obtains the decoded data 1272 (e.g., one of the audio stream 134 or the video stream 136) from the decoder 1270 and obtains the other of the audio stream 134 or the video stream 136 separately from the decoded data 1272, such as from another component or device. The video stream updater 110 selectively updates the video stream 136, as described with reference to FIG. 1, and provides the video stream 136 (e.g., after the selective updating) to one or more display devices 1114, one or more storage devices, or a combination thereof.

[0205] 13, there is shown a diagram of a system 1300 operable to perform keyword-based object insertion into a video stream. The system 1300 includes a device 130 coupled to one or more microphones 1302 and one or more display devices 1114.

[0206]

[0195] The one or more microphones 1302 are shown as being external to the device 130 as an illustrative example. In other examples, the one or more microphones 1302 may be integrated into the device 130. The video stream updater 110 receives the audio stream 134 from the one or more microphones 1302 and obtains the video stream 136 separately from the audio stream 134. In certain aspects, the audio stream 134 includes a user's speech. The video stream updater 110 selectively updates the video stream 136 as described with reference to FIG. 1 and provides the video stream 136 to one or more display devices 1114. In certain aspects, the video stream updater 110 provides the video stream 136 to a display screen of one or more authorized devices (e.g., one or more display devices 1114). For example, device 130 captures the speech of a performer while the performer is backstage at a concert and transmits enhanced video content (e.g., video stream 136) to the devices of premium ticket holders.

[0207] 14, there is shown a diagram of a system 1400 operable to perform keyword-based object insertion into a video stream. The system 1400 includes a device 130 coupled to one or more cameras 1402 and one or more display devices 1114.

[0208]

[0197] The one or more cameras 1402 are shown as being external to the device 130 as an illustrative example. In other examples, the one or more cameras 1402 may be integrated into the device 130. The video stream updater 110 receives the video streams 136 from the one or more cameras 1402 and obtains the audio streams 134 separately from the video streams 136. The video stream updater 110 selectively updates the video streams 136 and provides the video streams 136 to one or more display devices 1114, as described with reference to FIG. 1 .

[0209] 15 is a block diagram of an exemplary aspect of a system 1500 operable to perform keyword-based object insertion into a video stream, according to some examples of the present disclosure, where one or more processors 102 include an always-on power domain 1503 and a second power domain 1505, such as an on-demand power domain. In some implementations, a first stage 1540 and a buffer 1560 of the multi-stage system 1520 are configured to operate in an always-on mode, and a second stage 1550 of the multi-stage system 1520 is configured to operate in an on-demand mode.

[0210] The always-on power domain 1503 includes a buffer 1560 and a first stage 1540. Optionally, in some implementations, the first stage 1540 includes a position determination unit 170. The buffer 1560 is configured to store at least a portion of the audio stream 134 and at least a portion of the video stream 136 to be accessible for processing by the components of the multi-stage system 1520. For example, the buffer 1560 stores one or more portions of the audio stream 134 to be accessible for processing by the components of the second stage 1550, and stores one or more portions of the video stream 136 to be accessible for processing by the components of the first stage 1540, the second stage 1550, or both.

[0211] The second power domain 1505 includes a second stage 1550 of the multi-stage system 1520, and also includes an activation circuit 1530. Optionally, in some implementations, the second stage 1550 includes the keyword detection unit 112, the object determination unit 114, the object insertion unit 116, or a combination thereof.

[0212] The first stage 1540 of the multi-stage system 1520 is configured to generate at least one of a wake-up signal 1522 or an interrupt 1524 to initiate one or more operations in the second stage 1550. In an example, the wake-up signal 1522 is configured to transition the second power domain 1505 from a low power mode 1532 to an active mode 1534 to activate one or more components of the second stage 1550.

[0213] For example, the activation circuit 1530 may include or be coupled to a power management circuit, a clock circuit, a headswitch or footswitch circuit, a buffer control circuit, or any combination thereof. The activation circuit 1530 may be configured to initiate power-up of the second stage 1550, such as by selectively applying or increasing the voltage of a power supply for the second stage 1550, the second power domain 1505, or both. As another example, the activation circuit 1530 may be configured to selectively gate or ungate a clock signal to the second stage 1550, such as to prevent or enable circuit operation without removing power.

[0214] In some implementations, the first stage 1540 includes the position determination unit 170, and the second stage 1550 includes the keyword detection unit 112, the object determination unit 114, the object insertion unit 116, or a combination thereof. In these implementations, the first stage 1540 is configured to generate at least one of a wake-up signal 1522 or an interrupt 1524 to initiate operation of the keyword detection unit 112 of the second stage 1550 in response to the position determination unit 170 detecting the at least one insertion position 164.

[0215] In some implementations, the first stage 1540 includes the keyword detection unit 112, and the second stage 1550 includes the location determination unit 170, the object determination unit 114, the object insertion unit 116, or a combination thereof. In these implementations, the first stage 1540 is configured to generate at least one of a wake-up signal 1522 or an interrupt 1524 to initiate operation of the location determination unit 170, the object determination unit 114, or both of the second stage 1550 in response to the keyword detection unit 112 determining one or more detected keywords 180.

[0216] Output 1552 generated by second stage 1550 of multi-stage system 1520 is provided to application 1554. Application 1554 may be configured to output video stream 136 to one or more display devices, audio stream 134 to one or more speakers, or both. Illustratively, application 1554 may correspond to, as illustrative, non-limiting examples, a voice interface application, an integrated assistant application, a vehicle navigation and entertainment application, a gaming application, a social networking application, or a home automation system.

[0217]

[0206] By selectively activating the second stage 1550 based on the results of processing data in the first stage 1540 of the multi-stage system 1520, the overall power consumption associated with keyword-based object insertion into the video stream can be reduced.

[0218] 16, a diagram 1600 of an exemplary aspect of operation of components of the system 100 of FIG. 1 is shown, in accordance with some examples of the present disclosure. The keyword detection unit 112 is configured to receive a sequence 1610 of audio data samples, such as a sequence of consecutively captured frames of the audio stream 134, shown as a first frame (A1) 1612, a second frame (A2) 1614, and one or more additional frames including an Nth frame (AN) 1616 (where N is an integer greater than 2). The keyword detection unit 112 is configured to output a sequence 1620 of sets of detected keywords 180, including one or more additional sets including a first set (K1) 1622, a second set (K2) 1624, and an Nth set (KN) 1626.

[0219] The object determination unit 114 is configured to receive a sequence 1620 of sets of detected keywords 180. The object determination unit 114 is configured to output a sequence 1630 of sets of one or more objects 182, including a first set (O1) 1632, a second set (O2) 1634, and one or more additional sets including an N set (ON) 1636.

[0220] The position determination unit 170 is configured to receive a sequence 1640 of video data samples, such as a sequence of consecutively captured frames of the video stream 136, shown as a first frame (V1) 1642, a second frame (V2) 1644, and one or more additional frames including an Nth frame (VN) 1646. The position determination unit 170 is configured to output a sequence 1650 of sets of one or more insertion locations 164, including a first set (L1) 1652, a second set (L2) 1654, and one or more additional sets including an Nth set (LN) 1656.

[0221]

[0210] The object insertion unit 116 is configured to receive the sequence 1630, the sequence 1640, and the sequence 1650. The object insertion unit 116 is configured to output a sequence 1660 of video data samples, such as frames of the video stream 136, for example a first frame (V1) 1642, a second frame (V2) 1644, and one or more additional frames including an Nth frame (VN) 1646.

[0222] During operation, the keyword detection unit 112 processes the first frame 1612 to generate a first set 1622 of detected keywords 180. In some examples, the keyword detection unit 112, in response to determining that no keywords were detected in the first frame 1612, generates the first set 1622 indicating that no keywords were detected (e.g., an empty set). The position determination unit 170 processes the first frame 1642 to generate a first set 1652 of insertion positions 164. In some examples, the position determination unit 170, in response to determining that no insertion positions are detected in the first frame 1642, generates the first set 1652 indicating the detected insertion positions (e.g., an empty set).

[0223] Optionally, in some aspects, the first frame 1612 is time-aligned with the first frame 1642. For example, a particular time (e.g., capture time, playback time, receipt time, creation time, etc.) indicated by a first timestamp associated with the first frame 1612 is within a threshold duration of the corresponding time of the first frame 1642.

[0224] The object determination unit 114 processes the first set 1622 of the detected keywords 180 to generate a first set 1632 of one or more objects 182. In some examples, the object determination unit 114 generates the first set 1632 (e.g., an empty set) indicating that there are no objects associated with the first set 1622 of the detected keywords 180 in response to determining that the first set 1622 (e.g., an empty set) indicates that there are no detected keywords, that there are no objects associated with the first set 1622 (e.g., no existing objects and no created objects), or both.

[0225] The object insertion unit 116 processes the first frame 1642 of the video stream 136, the first set 1652 of insertion locations 164, and the first set 1632 of one or more objects 182 to selectively update the first frame 1642. The sequence 1660 includes the selectively updated version of the first frame 1642. By way of example, the object insertion unit 116 adds the first frame 1642 to the sequence 1660 (without inserting an object) in response to determining that the first set 1652 (e.g., an empty set) does not indicate a detected insertion location, that the first set 1632 (e.g., an empty set) does not indicate an object (e.g., no existing object and no created object), or both. Alternatively, if the first set 1632 includes one or more objects and the first set 1652 indicates one or more insertion positions 164, the object insertion unit 116 updates the first frame 1642 by inserting one or more objects from the first set 1652 into the one or more insertion positions 164 indicated by the first set 1632, and adds the updated version of the first frame 1642 to the sequence 1660.

[0226] Optionally, in some examples, the object insertion unit 116 updates one or more additional frames of the sequence 1640 in response to updating the first frame 1642. For example, the first set 1632 of objects 182 may be inserted into multiple frames of the sequence 1640 such that the objects persist across two or more video frames during playout. Optionally, in some aspects, the object insertion unit 116 instructs the keyword detection unit 112 to skip processing of one or more frames of the sequence 1610 in response to updating the first frame 1642. For example, one or more detected keywords 180 may remain the same for at least a threshold count of frames of the sequence 1610 such that updates to frames of the sequence 1660 correspond to the same keywords 180 for at least a threshold count of frames.

[0227] In an example, the insertion location 164 indicates a particular location within the first frame 1642, and generating an updated version of the first frame 1642 includes inserting at least one object of the first set 1632 at the particular location within the first frame 1642. In another example, the insertion location 164 indicates particular content (e.g., a shirt) represented in the first frame 1642. In this example, generating an updated version of the first frame 1642 includes performing image recognition to detect the location of the content (e.g., the shirt) within the first frame 1642, and inserting at least one object of the first set 1632 at the detected location within the first frame 1642. In some examples, the insertion location 164 indicates one or more particular image frames (e.g., a threshold count of image frames). Illustratively, in response to updating the first frame 1642, the object insertion unit 116 selects up to a threshold count of image frames subsequent to the first frame 1642 in the sequence 1640 as one or more additional frames for insertion. Updating the one or more additional frames includes performing image recognition to detect a location of content (e.g., a shirt) in each of the one or more additional frames. In response to determining that content is detected in the additional frame, the object insertion unit 116 inserts at least one object in the additional frame at the detected location of the content. Alternatively, in response to determining that content is not detected in the additional frame, the object insertion unit 116 skips the insertion in that additional frame and processes the next additional frame for insertion. Illustratively, the inserted object changes position as the content (e.g., a shirt) changes position in the additional frames, and the object is not inserted into any of the additional frames in which content is not detected.

[0228]

[0217] Such processing continues, including the keyword detection unit 112 processing the Nth frame 1616 of the audio stream 134 to generate an Nth set 1626 of detected keywords 180; the object determination unit 114 processing the Nth set 1626 of the detected keywords 180 to generate an Nth set 1636 of objects 182; the position determination unit 170 processing the Nth frame 1646 of the video stream 136 to generate an Nth set 1656 of insertion positions 164; and the object insertion unit 116 selectively updating the Nth frame 1646 of the video stream 136 based on the Nth set 1636 of objects 182 and the Nth set 1656 of insertion positions 164 to generate the Nth frame 1646 of the sequence 1660.

[0229] 17 shows an implementation 1700 of device 130 as an integrated circuit 1702 that includes one or more processors 102. The one or more processors 102 include a video stream updater 110. The integrated circuit 1702 also includes an audio input 1704, such as one or more bus interfaces, to allow an audio stream 134 to be received for processing. The integrated circuit 1702 includes a video input 1706, such as one or more bus interfaces, to allow a video stream 136 to be received for processing. The integrated circuit 1702 includes a video output 1708, such as a bus interface, to allow transmission of an output signal such as video stream 136 (e.g., following insertion of one or more objects 182 of FIG. 1). The integrated circuit 1702 enables implementation of keyword-based object insertion into a video stream as a component within a system such as a mobile phone or tablet as shown in FIG. 18, a headset as shown in FIG. 19, a wearable electronic device as shown in FIG. 20, a voice-controlled speaker system as shown in FIG. 21, a camera as shown in FIG. 22, an XR headset as shown in FIG. 23, XR glasses as shown in FIG. 24, or a vehicle as shown in FIG. 25 or 26.

[0230] 18 shows, as an illustrative, non-limiting example, an implementation 1800 in which the device 130 includes a mobile device 1802, such as a phone or tablet. The mobile device 1802 includes one or more microphones 1302, one or more cameras 1402, and a display screen 1804. Components of one or more processors 102, including the video stream updater 110, are integrated into the mobile device 1802 and are shown using dashed lines to indicate internal components that are generally not visible to a user of the mobile device 1802. In a particular example, the video stream updater 110 operates to detect user voice activity in the audio stream 134, which is then processed to perform one or more actions on the mobile device 1802, such as inserting one or more objects 182 of FIG. 1 into the video stream 136 and launching a graphical user interface or otherwise displaying the video stream 136 (e.g., with the inserted object(s) 182) on the display screen 1804 (e.g., via an integrated “smart assistant” application).

[0231] 19 shows an implementation 1900 in which device 130 includes a headset device 1902. Headset device 1902 includes one or more microphones 1302, one or more cameras 1402, or a combination thereof. One or more components of processor 102, including video stream updater 110, are integrated into headset device 1902. In a particular example, video stream updater 110 operates to detect user voice activity in audio stream 134, which is then processed at headset device 1902 to perform one or more operations, such as inserting one or more objects 182 of FIG. 1 into video stream 136 and transmitting video data corresponding to video stream 136 (e.g., with inserted objects 182) to a second device (not shown), such as one or more display devices 1114 of FIG. 11 for display.

[0232] 20 shows an implementation 2000 in which the device 130 includes a wearable electronic device 2002, shown as a “smart watch.” The video stream updater 110, one or more microphones 1302, one or more cameras 1402, or a combination thereof, are integrated into the wearable electronic device 2002. In a particular example, the video stream updater 110 operates to detect user voice activity in the audio stream 134, which is then processed to perform one or more actions on the wearable electronic device 2002, such as inserting one or more objects 182 of FIG. 1 into the video stream 136 and launching a graphical user interface or otherwise displaying the video stream 136 (e.g., with the inserted objects 182) on a display screen 2004 of the wearable electronic device 2002. To illustrate, the wearable electronic device 2002 may include a display screen configured to display notifications based on user utterances detected by the wearable electronic device 2002. In a particular example, the wearable electronic device 2002 includes a haptic device that provides a haptic notification (e.g., vibration) in response to detecting the user's voice activity. For example, the haptic notification may cause the user to look at the wearable electronic device 2002 to see a displayed notification (e.g., an object 182 inserted into the video stream 136) that corresponds to a detected keyword 180 spoken by the user. Thus, the wearable electronic device 2002 can alert a user who is hearing impaired or wearing a headset that the user's voice activity has been detected.

[0233] 21 is an implementation 2100 in which device 130 includes a wireless speaker and voice-activated device 2102. The wireless speaker and voice-activated device 2102 can have wireless network connectivity and is configured to perform Assistant operations. One or more processors 102, including a video stream updater 110, one or more microphones 1302, one or more cameras 1402, or a combination thereof, are included in the wireless speaker and voice-activated device 2102. The wireless speaker and voice-activated device 2102 also includes a speaker 2104. In operation, in response to receiving a verbal command (e.g., one or more detected keywords 180) identified in audio stream 134 via operation of video stream updater 110, wireless speaker and voice-activated device 2102 can perform an assistant action, such as inserting one or more objects 182 into video stream 136 and presenting video stream 136 (e.g., with inserted objects 182) to another device, such as one or more display devices 1114 of Figure 11. For example, wireless speaker and voice-activated device 2102 performs an assistant action, such as displaying an image associated with a restaurant, in response to receiving a key phrase (e.g., "hello, assistant") followed by one or more detected keywords 180 (e.g., "I'm hungry").

[0234] 22 illustrates an implementation 2200 in which device 130 includes a portable electronic device corresponding to a camera device 2202. Video stream updater 110, one or more microphones 1302, or a combination thereof are included in camera device 2202. In a particular aspect, one or more cameras 1402 of FIG. 14 include camera device 2202. In operation, in response to receiving a verbal command (e.g., one or more detected keywords 180) identified in audio stream 134 via operation of video stream updater 110, camera device 2202 can perform an action responsive to the verbal user command, such as inserting one or more objects 182 into video stream 136 captured by camera device 2202 and displaying video stream 136 (e.g., with the inserted objects 182) on one or more display devices 1114 of FIG. 11. In some aspects, the one or more display devices 1114 may include a display screen of the camera device 2202, another device, or both.

[0235] FIG. 23 illustrates an implementation 2300 in which the device 130 includes a portable electronic device corresponding to an XR headset 2302. The XR headset 2302 may include a virtual reality, mixed reality, or augmented reality headset. The video stream updater 110, one or more microphones 1302, one or more cameras 1402, or a combination thereof, are integrated into the XR headset 2302. In particular aspects, user voice activity detection may be performed on an audio stream 134 received from the one or more microphones 1302 of the XR headset 2302. The visual interface device is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the XR headset 2302 is worn. In particular examples, the video stream updater 110 inserts one or more objects 182 into the video stream 136, and the visual interface device is configured to display the video stream 136 (e.g., with the inserted objects 182). In certain aspects, the video stream updater 110 provides the video stream 136 (e.g., along with inserted objects 182) to a shared environment that is displayed by the XR headset 2302, one or more additional XR devices, or a combination thereof.

[0236] FIG. 24 shows an implementation 2400 in which the device 130 includes a portable electronic device corresponding to XR glasses 2402. The XR glasses 2402 can include virtual reality, augmented reality, or mixed reality glasses. The XR glasses 2402 include a holographic projection unit 2404 configured to project visual data onto a surface of lenses 2406 or reflect visual data from the surface of lenses 2406 onto the wearer's retinas. The video stream updater 110, one or more microphones 1302, one or more cameras 1402, or a combination thereof, are integrated into the XR glasses 2402. The video stream updater 110 can function to insert one or more objects 182 into the video stream 136 based on one or more detected keywords 180 detected in the audio stream 134 received from the one or more microphones 1302. In a particular example, the holographic projection unit 2404 is configured to display the video stream 136 (e.g., with the inserted objects 182). In certain aspects, the video stream updater 110 provides the video stream 136 (e.g., along with inserted objects 182) to a shared environment that is displayed by the holographic projection unit 2404, one or more additional XR devices, or a combination thereof.

[0237] In particular examples, the holographic projection unit 2404 is configured to display one or more of the inserted objects 182 indicative of the detected audio events. For example, the one or more objects 182 may be superimposed on the user's field of view at specific positions that coincide with the positions of sources of sounds associated with the detected audio events in the audio stream 134. To illustrate, sounds may be perceived by the user as emanating from the direction of the one or more objects 182. In example implementations, the holographic projection unit 2404 is configured to display one or more objects 182 associated with the detected audio events (e.g., one or more detected keywords 180).

[0238] FIG. 25 illustrates an implementation 2500 in which the device 130 corresponds to or is integrated into a vehicle 2502, which may be depicted as a manned or unmanned aerial device (e.g., a package delivery drone). The video stream updater 110, one or more microphones 1302, one or more cameras 1402, or a combination thereof, are integrated into the vehicle 2502. User voice activity detection may be performed based on an audio stream 134 received from the one or more microphones 1302 of the vehicle 2502, such as for delivery instructions from an authorized user of the vehicle 2502. In certain aspects, the video stream updater 110 updates the video stream 136 with one or more objects 182 (e.g., assembly instructions) based on one or more detected keywords 180 detected in the audio stream 134 and provides the video stream 136 (e.g., with the inserted objects 182) to one or more display devices 1114 of FIG. 11 . The one or more display devices 1114 may include a display screen of the vehicle 2502, a user device, or both.

[0239] 26 shows another implementation 2600 in which the device 130 corresponds to or is integrated into a vehicle 2602, shown as a car. The vehicle 2602 includes one or more processors 102 that include a video stream updater 110. The vehicle 2602 also includes one or more microphones 1302, one or more cameras 1402, or a combination thereof.

[0240] In some examples, the one or more microphones 1302 are positioned to capture the vocalizations of the operator of the vehicle 2602. User voice activity detection may be performed based on the audio stream 134 received from the one or more microphones 1302 of the vehicle 2602. In some implementations, user voice activity detection may be performed based on the audio stream 134 received from an internal microphone (e.g., one or more microphones 1302), such as for voice commands from authorized occupants. For example, user voice activity detection may be used to detect voice commands from the operator of the vehicle 2602 (e.g., from a parent requesting the location of a sushi restaurant) and ignore the voice of another occupant (e.g., a child requesting the location of an ice cream shop).

[0241] In particular implementations, the video stream updater 110, in response to determining one or more detected keywords 180 in the audio stream 134, inserts one or more objects 182 into the video stream 136 and presents the video stream 136 (e.g., with the inserted objects 182) to the display 2620. In particular aspects, the audio stream 134 includes speech of an occupant of the vehicle 2602 (e.g., "Sushi is my favorite"). The video stream updater 110 determines the one or more detected keywords 180 (e.g., "sushi") based on the audio stream 134 and determines a first position of the vehicle 2602 at a first time based on global positioning system (GPS) data.

[0242] 1, the video stream updater 110 determines one or more objects 182 corresponding to the one or more detected keywords 180. Optionally, in some aspects, the video stream updater 110 adaptively classifies the one or more objects 182 associated with the one or more detected keywords 180 and the first location using the adaptive classifier 144. For example, in response to determining that the set of objects 122 includes an object 122A (e.g., a sushi restaurant image) associated with one or more keywords 120A (e.g., “sushi,” “restaurant”) that match the one or more detected keywords 180 (e.g., “sushi”) and that is associated with a particular location within a threshold distance of the first location, the video stream updater 110 adds the object 122A to the one or more objects 182 (e.g., without classifying the one or more objects 182).

[0243] In certain aspects, in response to determining that the set of objects 122 does not include any objects associated with the one or more detected keywords 180 and a location within a threshold distance of the first location, the video stream updater 110 uses the adaptive classifier 144 to classify the one or more objects 182. In certain aspects, classifying the one or more objects 182 includes using the object generation neural network 140 to determine the one or more objects 182 associated with the one or more detected keywords 180 and the first location. For example, the video stream updater 110 retrieves, from a navigation database, addresses of restaurants within a threshold range of the first location, applies the object generation neural network 140 to the addresses and the one or more detected keywords 180 (e.g., “sushi”) to generate object 122A (e.g., clip art showing a sushi roll and the address), and adds object 122A to the one or more objects 182.

[0244] In certain aspects, classifying the one or more objects 182 includes using an object classification neural network 142 to determine one or more objects 182 associated with the one or more detected keywords 180 and the first location. For example, the video stream updater 110 uses the object classification neural network 142 to process an object 122A (e.g., an image showing a sushi roll and an address) and determine that the object 122A is associated with a keyword 120A (e.g., "sushi") and the address. In response to determining that the keyword 120A (e.g., "sushi") matches the one or more detected keywords 180 and that the address is within a threshold distance of the first location, the video stream updater 110 adds the object 122A to the one or more objects 182.

[0245] The video stream updater 110 inserts one or more objects 182 into the video stream 136 and provides the video stream 136 (e.g., with the inserted objects 182) to the display 2620. For example, the inserted objects 182 are overlaid on navigation information shown on the display 2620. In particular aspects, the video stream updater 110 determines a second position of the vehicle 2602 based on the GPS data at a second time. In particular implementations, the video stream updater 110 dynamically updates the video stream 136 based on changes in the position of the vehicle 2602. The video stream updater 110 uses the adaptive classifier 144 to classify one or more second objects associated with the one or more detected keywords 180 and the second position and inserts the one or more second objects into the video stream 136.

[0246]

[0235] In certain aspects, the fleet of vehicles includes vehicle 2602 and one or more additional vehicles, and the video stream updater 110 provides the video stream 136 (e.g., along with inserted objects 182) to the display devices of one or more vehicles in the fleet.

[0247] 27, shown is a particular implementation of a method 2700 for keyword-based object insertion into a video stream. In a particular aspect, one or more operations of the method 2700 are performed by at least one of the keyword detection unit 112, the object determination unit 114, the adaptive classifier 144, the object insertion unit 116, the video stream updater 110, one or more processors 102, the device 130, the system 100, or a combination thereof of FIG.

[0248]

[0237] The method 2700 includes obtaining an audio stream, at 2702. For example, the keyword detection unit 112 of Figure 1 obtains the audio stream 134, as described with reference to Figure 1 .

[0249]

[0238] The method 2700 also includes detecting one or more keywords in the audio stream, at 2704. For example, the keyword detection unit 112 of Figure 1 detects one or more detected keywords 180 in the audio stream 134, as described with reference to Figure 1.

[0250] Method 2700 further includes, at 2706, adaptively classifying one or more objects associated with the one or more keywords. For example, adaptive classifier 144 of FIG. 1 may classify (e.g., for identification, generation, or both via neural network-based classification) one or more objects 182 associated with the one or more detected keywords 180 in response to determining that none of the set of objects 122 stored in database 150 are associated with the one or more detected keywords 180. Alternatively, adaptive classifier 144 of FIG. 1 may designate at least one of the set of objects 122 as one or more objects 182 associated with the one or more detected keywords 180 (e.g., without classifying the one or more objects 182) in response to determining that at least one of the set of objects 122 is associated with at least one of the one or more detected keywords 180.

[0251] Optionally, in some implementations, adaptively classifying at 2706 includes using an object generation neural network to generate one or more objects based on the one or more keywords at 2708. For example, the adaptive classifier 144 of FIG. 1 uses the object generation neural network 140 to generate at least one of the one or more objects 182 based on the one or more detected keywords 180, as described with reference to FIG.

[0252] Optionally, in some implementations, adaptively classifying at 2706 includes determining, at 2710, using an object classification neural network to determine that one or more objects are associated with the one or more detected keywords 180. For example, the adaptive classifier 144 of FIG. 1 uses the object classification neural network 142, as described with reference to FIG. 1, to determine that at least one of the objects 122 is associated with the one or more detected keywords 180 and adds the at least one of the objects 122 to the one or more objects 182.

[0253]

[0242] The method 2700 includes inserting one or more objects into the video stream, at 2712. For example, the object insertion unit 116 of Figure 1 inserts one or more objects 182 into the video stream 136, as described with reference to Figure 1.

[0254]

[0243] Thus, method 2700 enables enhancing video stream 136 with one or more objects 182 associated with one or more detected keywords 180. Enrichment to video stream 136 can improve viewer retention, create advertising opportunities, and the like. For example, adding an object to video stream 136 can make video stream 136 more interesting to a viewer. By way of illustration, adding object 122A (e.g., an image of the Statue of Liberty) can increase viewer retention of video stream 136 if audio stream 134 includes one or more detected keywords 180 (e.g., "New York City") associated with object 122A. In another example, object 122A can correspond to a visual element representing a related entity associated with one or more detected keywords 180 (e.g., an image associated with a restaurant in New York, a restaurant serving food associated with New York, another business selling New York-related goods or services, a travel website, or a combination thereof).

[0255] 27 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a neural processing unit (NPU), a controller, another hardware device, a firmware device, or any combination thereof. By way of example, the method 2700 of FIG. 27 may be performed by a processor executing instructions such as those described with reference to FIG. 28.

[0256] 28, a block diagram of a particular example implementation of a device is shown, generally designated 2800. In various implementations, device 2800 may have more or fewer components than those shown in FIG. 28. In an example implementation, device 2800 may correspond to device 130. In an example implementation, device 2800 may perform one or more of the operations described with reference to FIGS. 1-27.

[0257] In particular implementations, device 2800 includes a processor 2806 (e.g., a CPU). Device 2800 may include one or more additional processors 2810 (e.g., one or more DSPs). In particular aspects, one or more processors 102 of FIG. 1 correspond to processor 2806, processor 2810, or a combination thereof. Processor 2810 may include a speech and music coder-decoder (CODEC) 2808, including a voice coder (“vocoder”) encoder 2836, a vocoder decoder 2838, a video stream updater 110, or a combination thereof.

[0258] The device 2800 may include a memory 2886 and a codec 2834. The memory 2886 may include instructions 109 executable by one or more additional processors 2810 (or processor 2806) to implement the functionality described with reference to the video stream updater 110. The device 2800 may include a modem 2870 coupled to an antenna 2852 via a transceiver 2850.

[0259] In certain aspects, the modem 2870 is configured to receive data from and transmit data to one or more devices. For example, the modem 2870 is configured to receive the media stream 1164 of FIG. 11 from the device 1130 and provide the media stream 1164 to the demux 1172. In a particular example, the modem 2870 is configured to receive the video stream 136 from the video stream updater 110 and provide the video stream 136 to one or more display devices 1114 of FIG. 11. In another example, the modem 2870 is configured to receive the encoded data 1262 of FIG. 12 from the device 1206 and provide the encoded data 1262 to the decoder 1270. In some implementations, the modem 2870 is configured to receive the audio stream 134 from one or more microphones 1302 of FIG. 13, receive the video stream 136 from one or more cameras 1402, or a combination thereof.

[0260] The device 2800 may include a display 2828 coupled to the display controller 2826. In certain aspects, the one or more display devices 1114 of FIG. 1 include the display 2828. One or more speakers 2892, one or more microphones 1302, or a combination thereof may be coupled to a codec 2834. The codec 2834 may include a digital-to-analog converter (DAC) 2802, an analog-to-digital converter (ADC) 2804, or both. In certain implementations, the codec 2834 can receive analog signals from the one or more microphones 1302, convert the analog signals to digital signals using the analog-to-digital converter 2804, and provide the digital signals to the speech and music codec 2808 (e.g., as audio stream 134). The speech and music codec 2808 may process the digital signal, which may be further processed by the video stream updater 110. In particular implementations, the speech and music codec 2808 may provide the digital signal to a codec 2834. The codec 2834 may convert the digital signal to an analog signal using a digital-to-analog converter 2802 and provide the analog signal to one or more speakers 2892.

[0261] In certain implementations, the device 2800 may be included in a system-in-package or system-on-chip device 2822. In certain implementations, the memory 2886, the processor 2806, the processor 2810, the display controller 2826, the codec 2834, and the modem 2870 are included in the system-in-package or system-on-chip device 2822. In certain implementations, the input device 2830, the one or more cameras 1402, and the power supply 2844 are coupled to the system-in-package or system-on-chip device 2822. Moreover, in certain implementations, as shown in FIG. 28 , the display 2828, the input device 2830, the one or more cameras 1402, the one or more speakers 2892, the one or more microphones 1302, the antenna 2852, and the power supply 2844 are external to the system-in-package or system-on-chip device 2822. In particular implementations, each of the display 2828, input device 2830, one or more cameras 1402, one or more speakers 2892, one or more microphones 1302, antenna 2852, and power source 2844 may be coupled to a component of the system-in-package or system-on-chip device 2822, such as an interface or controller.

[0262]

[0251] Device 2800 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a playback device, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of-things (IoT) device, an extended reality (XR) device, a base station, a mobile device, or any combination thereof.

[0263] In connection with the described implementation, the apparatus includes means for acquiring an audio stream. For example, the means for acquiring may be the keyword detection unit 112, the video stream updater 110, one or more processors 102, the device 130, the system 100, the speech recognition neural network 460 of FIG. 4, the demux 1172 of FIG. 11, the decoder 1270 of FIG. 12, the buffer 1560 of FIG. 15, the first stage 1540, the always-on power domain 1503, the second stage 1550, the second power domain 1505, the integrated circuit 1702, the audio input 1704 of FIG. 17, the mobile device 1802 of FIG. 18, the headset device 1902 of FIG. 19, the wireless LAN 1902 of FIG. 20, or the like. 21, the camera device 2202 of FIG. 22, the XR headset 2302 of FIG. 23, the XR glasses 2402 of FIG. 24, the vehicle 2502 of FIG. 25, the vehicle 2602 of FIG. 26, the codec 2834, the ADC 2804, the speech and music codec 2808, the vocoder decoder 2838, the processor 2810, the processor 2806, the device 2800 of FIG. 28, one or more other circuits or components configured to obtain an audio stream, or any combination thereof.

[0264] The apparatus also includes means for detecting one or more keywords in the audio stream. For example, the means for detecting may be the keyword detection unit 112 of FIG. 1, the video stream updater 110, the one or more processors 102, the device 130, the system 100, the speech recognition neural network 460 of FIG. 4, the potential keyword detector 462, the keyword selector 464, the first stage 1540 of FIG. 15, the always-on power domain 1503, the second stage 1550, the second power domain 1505, the integrated circuit 1702 of FIG. 17, the mobile device 1802 of FIG. 18, the headset device 19 of FIG. 19, or the like. 20, the wearable electronic device 2002 of FIG. 21, the camera device 2202 of FIG. 22, the XR headset 2302 of FIG. 23, the XR glasses 2402 of FIG. 24, the vehicle 2502 of FIG. 25, the vehicle 2602 of FIG. 26, the processor 2810, the processor 2806, the device 2800 of FIG. 28, one or more other circuits or components configured to detect one or more keywords, or any combination thereof.

[0265] The apparatus further includes means for adaptively classifying one or more objects associated with the one or more keywords. For example, the means for adaptively classifying may include the object determination unit 114, the adaptive classifier 144, the object generation neural network 140, the object classification neural network 142, the video stream updater 110, the one or more processors 102, the device 130, the system 100 of FIG. 1 , the second stage 1550, the second power domain 1505 of FIG. 15, the integrated circuit 1702 of FIG. 17, the mobile device 1802 of FIG. 18, the headset device 19 of FIG. 19 ... 902, the wearable electronic device 2002 of FIG. 20, the voice-activated device 2102 of FIG. 21, the camera device 2202 of FIG. 22, the XR headset 2302 of FIG. 23, the XR glasses 2402 of FIG. 24, the vehicle 2502 of FIG. 25, the vehicle 2602 of FIG. 26, the processor 2810, the processor 2806, the device 2800 of FIG. 28, one or more other circuits or components configured to adaptively classify, or any combination thereof.

[0266]

[0255] The apparatus also includes means for inserting one or more objects into the video stream. 24, the vehicle 2502 of FIG. 25, the vehicle 2602 of FIG. 26, the processor 2810, the processor 2806, the device 2800 of FIG. 28, one or more other circuits or components configured to selectively insert one or more objects into the video stream, or any combination thereof.

[0267] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 2886) includes instructions (e.g., instructions 109) that, when executed by one or more processors (e.g., one or more processors 2810 or processor 2806), cause the one or more processors to obtain an audio stream (e.g., audio stream 134) and detect one or more keywords (e.g., one or more detected keywords 180) within the audio stream. The instructions, when executed by the one or more processors, also cause the one or more processors to adaptively classify one or more objects (e.g., one or more objects 182) associated with the one or more keywords. The instructions, when executed by the one or more processors, further cause the one or more processors to insert the one or more objects into a video stream (e.g., video stream 136).

[0268]

[0257] Certain aspects of the present disclosure are described below in a set of interrelated examples.

[0269]

[0258] According to Example 1, the device includes one or more processors configured to acquire an audio stream, detect one or more keywords in the audio stream, adaptively classify one or more objects associated with the one or more keywords, and insert the one or more objects into a video stream.

[0270]

[0259] Example 2 includes the device of Example 1, wherein the one or more processors are configured to classify one or more objects associated with the one or more keywords based on determining that none of the set of objects is indicated as being associated with the one or more keywords.

[0271]

[0260] Example 3 includes the device of example 1 or example 2, wherein classifying the one or more objects includes using an object generation neural network to generate the one or more objects based on the one or more keywords.

[0272] Example 4 includes the device of example 3, wherein the object generation neural network includes stacked generative adversarial networks (GANs).

[0273]

[0262] Example 5 includes the device of any of Examples 1 to 4, wherein classifying the one or more objects includes using an object classification neural network to determine that the one or more objects are associated with one or more keywords.

[0274] Example 6 includes the device of example 5, wherein the object classification neural network includes a convolutional neural network (CNN).

[0275]

[0264] Example 7 includes the device of any of Examples 1 to 6, wherein the one or more processors are configured to apply a keyword detection neural network to the audio stream to detect one or more keywords.

[0276] Example 8 includes the device of example 7, in which the keyword detection neural network includes a recurrent neural network (RNN).

[0277]

[0266] Example 9 includes the device of any of Examples 1 to 8, wherein the one or more processors are configured to apply a position neural network to the video stream to determine one or more insertion positions within one or more video frames of the video stream, and insert one or more objects at the one or more insertion positions within the one or more video frames.

[0278] Example 10 includes the device of example 9, wherein the position neural network includes a residual neural network (resnet).

[0279]

[0268] Example 11 includes the device of any of Examples 1 to 10, wherein the one or more processors are configured to insert a particular object into the foreground or background of the video stream based at least on a file type of the particular object of the one or more objects.

[0280]

[0269] Example 12 includes the device of any of Examples 1 to 11, wherein the one or more processors are configured to insert one or more objects into a foreground of the video stream in response to determining that the background of the video stream includes at least one object associated with one or more keywords.

[0281]

[0270] Example 13 includes the device of any of examples 1 to 12, wherein the one or more processors are configured to perform round-robin insertion of one or more objects in the video stream.

[0282]

[0271] Example 14 includes any of the devices of Examples 1 to 13, wherein the one or more processors are integrated into at least one of a mobile device, a vehicle, an augmented reality device, a communication device, a playback device, a television, or a computer.

[0283]

[0272] Example 15 includes the device of any of Examples 1 to 14, in which the audio stream and the video stream are included in the live media stream received at the one or more processors.

[0284] Example 16 includes the device of example 15, wherein the one or more processors are configured to receive a live media stream from the network device.

[0285]

[0274] Example 17 includes the device of example 16, further including a modem, wherein the one or more processors are configured to receive the live media stream via the modem.

[0286]

[0275] Example 18 includes the device of any of Examples 1 to 17, further including one or more microphones, and the one or more processors are configured to receive an audio stream from the one or more microphones.

[0287]

[0276] Example 19 includes the device of any of Examples 1 to 18, further including a display device, wherein the one or more processors are configured to provide the video stream to the display device.

[0288]

[0277] Example 20 includes the device of any of Examples 1 to 19, further including one or more speakers, and wherein the one or more processors are configured to output an audio stream via the one or more speakers.

[0289]

[0278] Example 21 includes the device of any of Examples 1 to 20, wherein the one or more processors are integrated into a vehicle, the audio stream includes speech from an occupant of the vehicle, and the one or more processors are configured to provide the video stream to a display device of the vehicle.

[0290]

[0279] Example 22 includes the device of Example 21, wherein the one or more processors are configured to determine a first location of the vehicle at a first time and adaptively classify one or more objects associated with the one or more keywords and the first location.

[0291]

[0280] Example 23 includes the device of Example 22, wherein the one or more processors are configured to determine a second position of the vehicle at a second time, adaptively classify one or more second objects associated with the one or more keywords and the second position, and insert the one or more second objects into the video stream.

[0292]

[0281] Example 24 includes the device of any of Examples 21 to 23, wherein the one or more processors are configured to transmit the video stream to a display device of one or more second vehicles.

[0293]

[0282] Example 25 includes the device of any of Examples 1 to 24, wherein the one or more processors are integrated into an extended reality (XR) device, the audio stream includes speech of a user of the XR device, and the one or more processors are configured to provide a video stream to a shared environment displayed by at least the XR device.

[0294]

[0283] Example 26 includes the device of any of Examples 1 to 25, wherein the audio stream includes a user's speech, and the one or more processors are configured to transmit the video stream to a display of one or more authorized devices.

[0295]

[0284] According to Example 27, a method includes obtaining an audio stream at a device; detecting one or more keywords in the audio stream at the device; selectively applying a neural network at the device to determine one or more objects associated with the one or more keywords; and inserting the one or more objects into a video stream at the device.

[0296]

[0285] Example 28 includes the method of Example 27, and further includes classifying one or more objects associated with the one or more keywords based on determining that none of the set of objects includes any object indicated as being associated with the one or more keywords.

[0297]

[0286] Example 29 includes the method of example 27 or example 28, wherein classifying the one or more objects includes using an object generation neural network to generate the one or more objects based on the one or more keywords.

[0298] Example 30 includes the method of example 29, in which the object generation neural network includes stacked generative adversarial networks (GANs).

[0299]

[0288] Example 31 includes any of the methods of Examples 27 to 30, wherein classifying the one or more objects includes using an object classification neural network to determine that the one or more objects are associated with the one or more keywords.

[0300] Example 32 includes the method of example 31, in which the object classification neural network includes a convolutional neural network (CNN).

[0301]

[0290] Example 33 includes the method of any of Examples 27 to 32, further including applying a keyword detection neural network to the audio stream to detect one or more keywords.

[0302] Example 34 includes the method of example 33, in which the keyword detection neural network includes a recurrent neural network (RNN).

[0303]

[0292] Example 35 includes any of the methods of Examples 27 to 34, and further includes applying a position neural network to the video stream to determine one or more insertion positions within one or more video frames of the video stream, and inserting one or more objects at the one or more insertion positions within the one or more video frames.

[0304] Example 36 includes the method of example 35, in which the position neural network includes a residual neural network (resnet).

[0305]

[0294] Example 37 includes any of the methods of Examples 27 to 36, and further includes inserting a particular object into the foreground or background of the video stream based at least on a file type of the particular object among the one or more objects.

[0306]

[0295] Example 38 includes any of the methods of Examples 27 to 37, and further includes inserting one or more objects into the foreground of the video stream in response to determining that the background of the video stream includes at least one object associated with one or more keywords.

[0307]

[0296] Example 39 includes the method of any of Examples 27 to 38, and further includes performing round-robin insertion of the one or more objects in the video stream.

[0308]

[0297] Example 40 includes any of the methods of Examples 27 to 39, wherein the device is integrated into at least one of a mobile device, a vehicle, an augmented reality device, a communication device, a playback device, a television, or a computer.

[0309]

[0298] Example 41 includes the method of any of Examples 27 to 40, in which the audio stream and the video stream are included in the live media stream received at the device.

[0310]

[0299] Example 42 includes the method of example 41, further including receiving a live media stream from the network device.

[0311]

[0300] Example 43 includes the method of example 42, further including receiving the live media stream via a modem.

[0312]

[0301] Example 44 includes the method of any of Examples 27 to 43, and further includes receiving an audio stream from one or more microphones.

[0313]

[0302] Example 45 includes the method of any of Examples 27 to 44, further including providing the video stream to a display device.

[0314]

[0303] Example 46 includes the method of any of Examples 27 to 45, further including providing the audio stream to one or more speakers.

[0315]

[0304] Example 47 includes the method of any of Examples 27 to 46, and further includes providing the video stream to a display device of the vehicle, and the audio stream includes speech of an occupant of the vehicle.

[0316]

[0305] Example 48 includes the method of Example 47, further including determining a first location of the vehicle at a first time and adaptively classifying one or more objects associated with the one or more keywords and the first location.

[0317]

[0306] Example 49 includes the method of example 48, and further includes determining a second position of the vehicle at a second time, adaptively classifying one or more second objects associated with the one or more keywords and the second position, and inserting the one or more second objects into the video stream.

[0318]

[0307] Example 50 includes the method of any of Examples 47 to 49, further including transmitting the video stream to a display device of one or more second vehicles.

[0319]

[0308] Example 51 includes any of the methods of Examples 27 to 50, and further includes providing a video stream to a shared environment displayed by at least an extended reality (XR) device, and the audio stream includes speech of a user of the XR device.

[0320]

[0309] Example 52 includes any of the methods of Examples 27 to 51, and further includes transmitting the video stream to a display of one or more authorized devices, and the audio stream includes the user's speech.

[0321]

[0310] According to Example 53, a device includes a memory configured to store instructions and a processor configured to execute the instructions to perform any of the methods of Examples 27 to 52.

[0322]

[0311] According to Example 54, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods of Examples 27 to 52.

[0323]

[0312] According to Example 55, an apparatus includes means for performing the method of any one of Examples 27 to 52.

[0324]

[0313] According to Example 56, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to obtain an audio stream, detect one or more keywords in the audio stream, adaptively classify one or more objects associated with the one or more keywords, and insert the one or more objects into a video stream.

[0325]

[0314] According to Example 57, the apparatus comprises means for acquiring an audio stream, means for detecting one or more keywords in the audio stream, means for adaptively classifying one or more objects associated with the one or more keywords, and means for inserting the one or more objects into a video stream.

[0326] Those skilled in the art will further appreciate that the various exemplary logical blocks, configurations, modules, circuits, and algorithm steps described with respect to the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been described above generally with respect to their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0327] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC), which may reside in a computing device or user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.

[0328]

[0317] The foregoing description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed embodiments. Various modifications of these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest possible scope consistent with the principles and novel features defined by the following claims.

Claims

1. A device, one or more processors, wherein the one or more processors: Get the audio stream, Detecting one or more keywords within the audio stream; adaptively classifying one or more objects associated with the one or more keywords; A device configured to insert the one or more objects into a video stream.

2. 10. The device of claim 1, wherein the one or more processors are configured to classify the one or more objects associated with the one or more keywords based on determining that none of a set of objects is indicated as being associated with the one or more keywords.

3. 10. The device of claim 1, wherein classifying the one or more objects comprises using an object generation neural network to generate the one or more objects based on the one or more keywords.

4. The device of claim 3 , wherein the object-generating neural network comprises stacked generative adversarial networks (GANs).

5. 10. The device of claim 1, wherein classifying the one or more objects comprises using an object classification neural network to determine that the one or more objects are associated with the one or more keywords.

6. The device of claim 5 , wherein the object classification neural network comprises a convolutional neural network (CNN).

7. 10. The device of claim 1, wherein the one or more processors are configured to apply a keyword detection neural network to the audio stream to detect the one or more keywords.

8. The device of claim 7 , wherein the keyword detection neural network comprises a recurrent neural network (RNN).

9. the one or more processors: applying a position neural network to the video stream to determine one or more insertion locations within one or more video frames of the video stream; The device of claim 1 , configured to insert the one or more objects at the one or more insertion locations in the one or more video frames.

10. The device of claim 9 , wherein the position neural network comprises a residual neural network (resnet).

11. 2. The device of claim 1, wherein the one or more processors are configured to insert a particular object of the one or more objects into the foreground or background of the video stream based at least on a file type of the particular object.

12. 2. The device of claim 1, wherein the one or more processors are configured to, in response to determining that a background of the video stream includes at least one object associated with the one or more keywords, insert the one or more objects into a foreground of the video stream.

13. The device of claim 1 , wherein the one or more processors are configured to perform round-robin insertion of the one or more objects into the video stream.

14. The device of claim 1 , wherein the one or more processors are integrated into at least one of a mobile device, a vehicle, an augmented reality device, a communication device, a playback device, a television, or a computer.

15. The device of claim 1 , wherein the audio stream and the video stream are included in a live media stream received at the one or more processors.

16. The device of claim 15 , wherein the one or more processors are configured to receive the live media stream from a network device.

17. 17. The device of claim 16, further comprising a modem, wherein the one or more processors are configured to receive the live media stream via the modem.

18. The device of claim 1 , further comprising one or more microphones, the one or more processors configured to receive the audio streams from the one or more microphones.

19. The device of claim 1 , further comprising a display device, the one or more processors configured to provide the video stream to the display device.

20. 10. The device of claim 1, further comprising one or more speakers, the one or more processors configured to output the audio stream through the one or more speakers.

21. 10. The device of claim 1, wherein the one or more processors are integrated into a vehicle, the audio stream includes speech from an occupant of the vehicle, and the one or more processors are configured to provide the video stream to a display device of the vehicle.

22. the one or more processors: determining a first position of the vehicle at a first time; The device of claim 21 , configured to adaptively categorize the one or more objects associated with the one or more keywords and the first location.

23. the one or more processors: determining a second position of the vehicle at a second time; adaptively classifying one or more second objects associated with the one or more keywords and the second location; The device of claim 22 , configured to insert the one or more second objects into the video stream.

24. 22. The device of claim 21, wherein the one or more processors are configured to transmit the video streams to one or more second vehicle display devices.

25. 10. The device of claim 1, wherein the one or more processors are integrated into an extended reality (XR) device, the audio stream includes speech of a user of the XR device, and the one or more processors are configured to provide the video stream to a shared environment displayed by at least the XR device.

26. 10. The device of claim 1, wherein the audio stream includes a user's speech, and the one or more processors are configured to transmit the video stream to a display of one or more authorized devices.

27. 1. A method comprising: acquiring an audio stream at a device; detecting, at the device, one or more keywords in the audio stream; adaptively classifying, at the device, one or more objects associated with the one or more keywords; and inserting, at the device, the one or more objects into a video stream.

28. 28. The method of claim 27, wherein classifying the one or more objects comprises using an object generation neural network to generate the one or more objects based on the one or more keywords.

29. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: Get the audio stream, detecting one or more keywords within the audio stream; adaptively classifying one or more objects associated with the one or more keywords; A non-transitory computer-readable medium for causing the one or more objects to be inserted into a video stream.

30. 1. An apparatus comprising: means for obtaining an audio stream; means for detecting one or more keywords within said audio stream; means for adaptively classifying one or more objects associated with said one or more keywords; means for inserting the one or more objects into a video stream.