Methods, systems, and circuitry for feature and parameter extraction for audio and video processing

By generating features and parameters in the media processing pipeline that are not protected by SVP, the challenge of real-time processing of large-scale media data by machine learning processors is addressed, achieving improvements in both security and performance.

CN116055782BActive Publication Date: 2026-06-30AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
Filing Date
2022-09-13
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

In digital video systems, machine learning processors face challenges in real-time processing of large-scale media data, and existing technologies struggle to simultaneously meet both high security requirements and the performance demands of machine learning algorithms.

Method used

Feature extraction is performed in the media processing pipeline to generate features and parameters that are not protected by SVP, which can be used by the machine learning processor, avoiding direct access to raw media data.

Benefits of technology

It significantly reduces the amount of data that machine learning processors need to process, improves system security and performance, and meets the security requirements of content providers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116055782B_ABST
    Figure CN116055782B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and apparatus for improving the security of media streams. The system receives a decoded media stream from a media decoding pipeline that receives and decodes an encoded media stream. The system identifies a set of features to be generated based on the decoded media stream and generates the set of features using the decoded media stream. Such features may include luminance histograms, pixel intensity data, motion vectors, edge detection, audio gain, pitch information, and other types of features. The system provides the set of features to a processor executing a machine learning model, wherein the apparatus prevents the processor executing the machine learning model from accessing the decoded media stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application generally relate to the field of media processing, and more specifically, to methods, systems and apparatus for improving the security of media streams. Background Technology

[0002] In digital video systems, encoded bitstreams of media data are decoded by a portion of the media processing system (e.g., a set-top box or TV processor) and subsequently output at the appropriate audio or video device. Performing real-time or near real-time processing of information from these streams is challenging because it requires handling extremely large volumes of data. Similarly, providers of such content often require adherence to stringent processing guidelines to prevent unauthorized access to or distribution of the content stream. These challenges make implementing machine learning architectures in media processing pipelines, such as set-top boxes or smart TVs, impractical using conventional methods. Summary of the Invention

[0003] The systems and methods disclosed herein address these and other problems by providing techniques for extracting features and parameters from decoded bitstreams in a media stream decoding pipeline. By generating these features within the media processing pipeline, the total amount of information provided to configurable machine learning architectures is significantly reduced, which substantially improves the performance of these machine learning systems. Similarly, by providing features to configurable machine learning, while preventing configurable machine learning from accessing the raw decoded media stream, the requirements imposed by content providers are satisfied, yet configurable processing within the media processing pipeline is still permitted.

[0004] At least one aspect of this disclosure relates to a method for improving the security of a media stream. The method may be performed, for example, by one or more processors of a device. The method may include receiving a decoded media stream from a media decoding pipeline that receives and decodes an encoded media stream. The method may include identifying a set of features to be generated based on the decoded media stream. The method may include generating the set of features using the decoded media stream. The method may include providing the set of features to a processor executing a machine learning model. The device prevents the processor executing the machine learning model from accessing the decoded media stream.

[0005] In some embodiments, the decoded media stream may contain video data. In some embodiments, the set of features may include a luminance histogram of the video data. In some embodiments, identifying the set of features may include determining the type of media in the decoded media stream. In some embodiments, identifying the set of features may include identifying the set of features based on the type of media in the decoded media stream. In some embodiments, identifying the set of features may include extracting raw video data from the decoded media stream. In some embodiments, identifying the set of features may include identifying the set of features based on attributes of the raw video data.

[0006] In some embodiments, generating the set of features may include preprocessing the decoded media stream. In some embodiments, generating the set of features may include applying edge detection or spectral analysis techniques to the decoded media stream. In some embodiments, generating the set of features may include determining one or more of a minimum pixel intensity value, a maximum pixel intensity value, an average pixel intensity value, or a median pixel intensity value based on one or more color channels of the decoded media stream. In some embodiments, generating the set of features may include detecting motion vectors in the decoded media stream. In some embodiments, providing the set of features to the machine learning model may include storing the set of features in a memory area accessible by the processor executing the machine learning model. In some embodiments, the method may include generating an output media stream comprising the output of the machine learning processor and graphics generated from the decoded media stream. In some embodiments, the method may include providing the output media stream for rendering at a display.

[0007] At least one aspect of this disclosure relates to a system for improving the security of a media stream. The system may include one or more processors. The system may receive a decoded media stream from a media decoding pipeline that receives and decodes an encoded media stream. The system may identify a set of features to be generated based on the decoded media stream. The system may generate the set of features using the decoded media stream. The system may provide the set of features to a processor executing a machine learning model. The apparatus prevents the processor executing the machine learning model from accessing the decoded media stream.

[0008] In some implementations, the decoded media stream may contain video data. In some implementations, the set of features includes a luminance histogram of the video data. In some implementations, to identify the set of features, the system may determine the type of media in the decoded media stream. In some implementations, the system may identify the set of features based on the type of media in the decoded media stream. In some implementations, to identify the set of features, the system may extract raw video data from the decoded media stream. In some implementations, the system may identify the set of features based on attributes of the raw video data.

[0009] In some implementations, the system may preprocess the decoded media stream to generate the set of features. In some implementations, the system may apply edge detection or spectral analysis techniques to the decoded media stream to generate the set of features. In some implementations, the system may determine one or more of a minimum pixel intensity value, a maximum pixel intensity value, an average pixel intensity value, or a median pixel intensity value based on one or more color channels of the decoded media stream to generate the set of features. In some implementations, the system may detect motion vectors in the decoded media stream to generate the set of features. In some implementations, the system may store the set of features in a memory area accessible by the processor executing the machine learning model to provide the set of features to the machine learning model.

[0010] At least one aspect of this disclosure relates to circuitry for improving the security of a media stream. The circuitry may include a bitstream decoder that receives an encoded bitstream and generates a decoded media stream. The circuitry may include a processing component that receives the decoded media stream from the bitstream decoder. The circuitry may identify a set of features to be generated based on the decoded media stream. The circuitry may generate the set of features using the decoded media stream. The circuitry may provide the set of features to a second processor executing a machine learning model. Access to the decoded media stream is prevented by the second processor executing the machine learning model.

[0011] In some embodiments, the circuit may include a second processor that executes the machine learning model. In some embodiments, the circuit is configured such that the second processor cannot access the decoded media stream.

[0012] These and other aspects and embodiments are discussed in detail below. The foregoing information and the following detailed description contain illustrative examples of the aspects and embodiments, and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and embodiments. Drawings provide illustration and further understanding of the aspects and embodiments, and are incorporated in and constitute a part of this specification. Aspects can be combined, and it will be readily understood that features described in the context of one aspect of the invention can be combined with other aspects. Aspects can be implemented in any convenient form. For example, they can be transmitted on a suitable carrier medium (computer-readable medium) by means of a suitable computer program, which can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). Aspects can also be implemented using suitable equipment, which can take the form of a programmable computer running a computer program arranged to implement the aspects. As used in the specification and claims, the singular forms 'a / an' and 'the' include a plural referent unless the context clearly indicates otherwise. Attached Figure Description

[0013] The present disclosure is described by way of example with reference to the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as background art, the figures represent aspects of the present disclosure. For clarity, not every component may be labeled in every figure. In the figures:

[0014] Figure 1 A block diagram illustrating an example system for improving the security of media streaming, based on one or more implementation schemes;

[0015] Figure 2 An example data flow diagram illustrating a technique for generating features from a media stream according to one or more implementation schemes;

[0016] Figure 3 Example flowchart illustrating a method for improving the security of media streams according to one or more implementation schemes;

[0017] Figure 4A and 4B The illustration depicts a block diagram illustrating an implementation scheme of a computing device available for the methods and systems described herein. Detailed Implementation

[0018] The following is a detailed description of various concepts and implementation schemes related to techniques, approaches, methods, devices, and systems for improving the security of media streaming. The various concepts introduced above and discussed in more detail below can be implemented in any of many ways, as the described concepts are not limited to any particular implementation. Examples of specific implementation schemes and applications are provided primarily for illustrative purposes.

[0019] Machine learning processors, specifically those running on media content such as video and audio, are useful for extracting insightful features from media streams. For example, machine learning techniques can be performed on media content to detect the presence of objects, identify themes or features, detect edges, or perform other types of insightful processing. There is increasing focus on executing such machine learning models on end-user devices (e.g., set-top boxes) used to display media content. This arrangement will be useful for: new patterns, such as scene and object detection, tracking and classification, identifying user preferences and generating predictions and recommendations, and audio or visual enhancement, etc.

[0020] One approach to media processing on a set-top box or other device receiving an encoded media stream is to send the audio and / or video data directly to a machine learning processor and execute the desired machine learning algorithm on the raw audio and / or video data on that processor. However, this approach has several problems. First, especially when considering high-resolution video data, the data volume is large and the cost of redirecting from the decoder to a separate processor (such as a software processor or a separate chip) is high. For example, high dynamic range video with 4K (3840x2160) resolution at sixty (60) frames per second results in a bit rate of approximately 14.9 Gbps. This amount of data is challenging for a machine learning processor to directly access and process in real time.

[0021] Additionally, content providers protect audio / video data transmitted to set-top boxes or similar streaming devices. Generally, set-top boxes are configured to securely protect this data and prevent access by other external processing components. For example, SVP (Secure Video Processing) requires strict prohibition of access to video / audio bitstreams or decoded raw pixels by other non-secure processing components, as doing so poses a significant security risk to the video and / or audio data. These security issues arise at the interface used to transfer data by copying it to processor-accessible memory and within the machine learning processor itself (e.g., programmed to only redirect / copy the signal). Therefore, any data that could be used to reconstruct the source media stream will be blocked according to these requirements.

[0022] While some methods can perform downsampling or scaling techniques to reduce the size of the bitstream, these solutions are insufficient for the machine learning processors and security requirements of media processing systems. For example, such methods might use a large source buffer (e.g., 4K HDR video at 60 frames per second) and convert HDR (10-bit) to SDR (8-bit) while scaling the video resolution down to a predetermined size (e.g., 512x512) and producing a standard RGB pixel format typically used by existing machine learning models. However, even with such techniques, the resulting media stream is impractical for most machine learning models in end-user devices (e.g., 512x512 pixels * 3 components / pixel * 8 bits / component * 60 frames / second = 377 Mbps bitrate). Furthermore, security vulnerabilities remain because direct pixel data is still exposed to external components. Even a down-resolution version of the media stream may violate security protocols implemented by such devices. Data must be further protected or at least obfuscated to prevent direct access in order to comply with requirements; however, doing so would also render conventional machine learning processes infeasible.

[0023] The systems and methods disclosed herein address these and other problems by implementing additional feature extraction processing within the media decoding pipeline. The resulting features are in a format unprotected by SVP (e.g., unusable for reconstructing the media stream), thus allowing them to be used by insecure operations. To this end, feature extraction techniques can compute or otherwise generate new types of data from the raw media stream, allowing other processors (e.g., machine learning processors) to run additional processing algorithms with minimal processing power. This improves both the security of the media stream (e.g., by complying with SVP requirements) and the overall performance of the system. By extracting only features useful to machine learning algorithms from the raw media data, the amount of information processed by the machine learning algorithms is significantly reduced without compromising overall algorithmic effectiveness. The performance of such algorithms is improved not only due to the reduced data size but also because the data is derived from the algorithm and therefore requires less complex preprocessing steps.

[0024] The systems and methods described herein implement hardware components (e.g., hardware blocks in a programmable processor, such as field-programmable gate arrays or other programmable logic) after the media decoder in a media processing pipeline. These hardware components directly receive the decoded media content and generate features and parameters compatible with machine learning processors. Some examples of such features may include luminance histograms, edge detection techniques, spectral analysis, motion vectors, or pixel parameters generated from the raw video content. Similarly, for audio content, bandgap, histogram, peak, and other audio analyses (e.g., pitch, tone, or harmonic analysis) may be performed. These features cannot be used to “reconstruct” the media stream and therefore do not require SVP protection. Likewise, the generated features are significantly smaller than the original pixel data and can be used to target and simplify specific machine learning algorithms or other useful processing algorithms. For example, the input data to a machine learning model (e.g., a neural network) is often specified and predetermined, and therefore the features can be generated or customized to the desired machine learning model or framework. These and other improvements are described in more detail herein.

[0025] Now for reference Figure 1 This diagram illustrates an example system 100 for improving the security of media streams according to one or more embodiments. System 100 may include at least one feature extraction system 105, at least one stream decoder 110, at least one machine learning processor, at least one display 145, and in some embodiments at least one output processor 150. Feature extraction system 105 may include at least one stream receiver 115, at least one feature recognizer 120, at least one feature generator 125, one or more features 130 (e.g., stored in a memory area at feature extraction system 105), and at least one feature provider 135.

[0026] Each of the components of system 100 (e.g., feature extraction system 105, stream decoder 110, machine learning processor 140, etc.) can use a computing system (e.g., regarding...) Figure 4A and 4BThe computing system 400 described herein may be implemented as a hardware component or a combination of software and hardware components. In some embodiments, the stream decoder 110 and the feature extraction system 105 may be hardware components defined in a configurable logic processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In some embodiments, the machine learning processor 140 may also be part of the same configurable logic processor. In some embodiments, the stream decoder 110, the feature extraction system 105, and the machine learning processor 140 may be separate hardware or software components. Each of the components of the feature extraction system 105 (e.g., stream receiver 115, feature recognizer 120, feature generator 125, one or more features 130, and feature provider 135) may be implemented in software, hardware, or any combination of software or hardware, and may perform the functionality described herein.

[0027] The feature extraction system 105 may include at least one processor and memory, such as processing circuitry. The memory may store processor-executable instructions that, when executed by the processor, cause the processor to perform one or more of the operations described herein. The processor may include a microprocessor, an ASIC, an FPGA, programmable logic circuitry, or a combination thereof. The memory may include (but is not limited to) electronic, optical, magnetic, or any other storage or transmission means capable of providing program instructions to the processor. The memory may further include floppy disks, CD-ROMs, DVDs, magnetic disks, memory chips, ASICs, FPGAs, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM (EEPROM), erasable programmable ROM (EPROM), flash memory, optical media, or any other suitable memory from which the processor can read instructions. The instructions may include code from any suitable computer programming language. The feature extraction system 105 may include one or more computing devices or components. The feature extraction system 105 may include any or all of these components and perform a combination thereof. Figure 4A and 4B The computer system 400 described may contain any or all of its functions. The feature extraction system 105 may communicate with other components of the system 100 using one or more communication interfaces (not shown). The communication interfaces may be or include any type of wired or wireless communication system for data communication, including suitable computer or logic interfaces or buses.

[0028] The memory of the feature extraction system 105 may be any suitable computer memory that stores or maintains any information in connection with the description of this technology. The memory may store one or more data structures that may contain, index, or otherwise store each of the values, complex numbers, sets, variables, vectors, or thresholds described herein. The memory may be accessed using one or more memory addresses, index values, or identifiers of any item, structure, or area maintained in the memory. The memory may be accessed and modified by components of the feature extraction system 105. In some embodiments, the memory may be internal to the feature extraction system 105. In some embodiments, the memory may reside external to the feature extraction system 105 and be accessible via a suitable communication interface (e.g., a communication bus). In some embodiments, the memory may be distributed across many different computer systems or storage elements and be accessible via a network or a suitable computer bus interface. The feature extraction system 105 may store the results of any or all calculations, determinations, selections, identifications, generation, constructions, or computations performed on one or more data structures indexed or identified with appropriate values ​​in one or more areas of the memory of the feature extraction system 105.

[0029] The memory may store feature 130, which is generated by components of the feature extraction system 105. The memory may be provided such that portions of the memory are accessible only by certain components of the system 100 via physical connection or by software configuration. For example, the memory area storing feature 130 may be accessed and modified (e.g., read and write access) by components of the feature extraction system 105. However, the machine learning processor 140 may only access the area of ​​the memory containing feature 130 and is prevented from accessing (e.g., by permission or due to lack of hardware connection) areas of the feature extraction system 105's memory containing any other data. This prevents the machine learning processor 140 from accessing any information that may be subject to SVP requirements, while still allowing the machine learning processor 140 to access features and parameters generated from the decoded media stream, for example, in real-time or near real-time. In some embodiments, feature 130 is stored in the memory area of ​​the machine learning processor 140. For example, the feature extraction system 105 (and its components) may make write accesses to one or more memory components of the machine learning processor 140, while the machine learning processor 140 may only be able to access its own memory. Using the techniques described herein, feature extraction system 105 can generate features from a decoded media stream and write features 130 into the memory of machine learning processor 140.

[0030] Feature 130 may include any type of feature that can be derived from the media stream (e.g., audio and / or video data) but not used to fully reconstruct the media stream. Some non-limiting examples of feature 130 may include a luminance histogram, an edge detection map (e.g., a matrix or data structure indicating the location of edges in video data), spectral analysis information (e.g., a data structure indicating spectral bands of one or more pixels in video data), motion vectors (e.g., the location, direction, and magnitude of movement detected in video data), and various parameters such as minimum pixel value (e.g., minimum intensity across one or more color channels of a pixel in video data), maximum pixel value (e.g., minimum intensity across one or more color channels of a pixel in video data), average pixel value (e.g., average intensity across one or more color channels of a pixel), and audio-related information in the media stream (e.g., gain information, frequency bands, histograms of various parameters of the audio stream, peak information, pitch information, cadence information, harmonic analysis information, etc.). Feature 130 may be formatted by components of feature extraction system 105 into one or more data structures to conform to one or more inputs of a model or algorithm executed at machine learning processor 140. In some implementations, configuration settings specifying the type and format of features to be generated from the decoded media stream may be stored in or provided to feature extraction system 105, as described herein.

[0031] In some embodiments, the feature extraction system 105 may generate one or more features 130 using the techniques described herein during an offline process (e.g., after the decoded media stream for display has been provided). In such embodiments, the feature extraction system 105 may store one or more portions of the decoded media stream provided by the stream decoder 110 for later execution of the feature extraction techniques described herein. In some embodiments, one or more of the features 130 may be extracted from the media stream by an external computing system (not shown) using a static or offline process similar to the feature generation described herein. The features 130 may then be provided to the feature extraction system 105 for use in the operations described herein. For example, in one such embodiment, feature extraction may be performed before the encoded media stream is provided to the device for display. In some such embodiments, one or more data structures including features 130 may be provided to the device for analysis during the decoding and display of the content. In other such embodiments, the analysis may also be performed before the encoded media stream is provided to the device for display, and the results of the analysis (e.g., object recognition, facial recognition, etc.) may be provided to the device or another device for display along with the decoding and display of the content.

[0032] Machine learning processor 140 may include at least one processor and memory (e.g., processing circuitry), and may be defined as part of a configurable logical structure defining other components of system 100. The memory may store processor-executable instructions that, when executed by the processor, cause the processor to perform one or more of the operations described herein. The processor may include a microprocessor, ASIC, FPGA, programmable logic circuitry, or a combination thereof. The memory may include (but is not limited to) electronic, optical, magnetic, or any other storage or transmission means capable of providing program instructions to the processor. The memory may further include floppy disks, CD-ROMs, DVDs, magnetic disks, memory chips, ASICs, FPGAs, ROMs, RAMs, EEPROMs, EPROMs, flash memory, optical media, or any other suitable memory from which the processor can read instructions. Instructions may include code from any suitable computer programming language. Machine learning processor 140 may include one or more computing devices or components. Machine learning processor 140 may include any or all of these components and perform a combination of... Figure 4A and 4B The computer system 400 described may contain any or all of its functions. The machine learning processor 140 may communicate with other components of the system 100 using one or more communication interfaces (not shown). The communication interfaces may be or include any type of wired or wireless communication system for data communication, including suitable computer or logic interfaces or buses.

[0033] Machine learning processor 140 may be a hardware processor, a software component, or a processor defined in configurable logic, which implements one or more processing algorithms on feature 130. Machine learning processor 140 may be capable of executing any type of algorithm compatible with feature 130, and may include artificial intelligence models such as linear regression models, logistic regression models, decision tree models, support vector machine (SVM) models, Naive Bayes models, k-nearest neighbor models, k-means models, random forest models, dimensionality reduction models, fully connected neural network models, recurrent neural network (RNN) models (e.g., Long Short-Term Memory (LSTM), independent RNNs, recurrent RNNs, etc.), convolutional neural network (CNN) models, clustering algorithms, or unsupervised learning techniques, etc. Machine learning processor 140 may execute other types of algorithms on feature 130. Machine learning processor 140 may be configured via programmable computer-readable instructions, which may be provided to machine learning processor 140 via one or more suitable communication interfaces. Machine learning processor 140 may access portions of the memory of feature extraction system 105, for example, areas of memory storing any extracted or generated features 130.

[0034] In some implementations, the machine learning processor 140 may include its own memory storing features 130, which are written to the machine learning processor 140's memory by the feature extraction system 105 via a suitable communication interface. Access to the decoded media stream provided by the stream decoder 110 is prevented, for example, by lack of permission to access the memory area storing the decoded media stream or by lack of hardware connection to the memory component storing the decoded media stream. In effect, this prevents the machine learning processor 140 from accessing any information subject to the SVP protocol, thereby improving system security. However, the machine learning processor 140 can still access the memory area storing features 130, which are not subject to SVP requirements, because features 130 cannot be used to fully reconstruct the decoded media stream.

[0035] In some implementations, machine learning processor 140 may provide output that influences the display of the decoded media stream provided for display. Machine learning processor 140 may provide such output to output processor 150, which may be a software or hardware module as part of a media processing pipeline and is therefore “trusted” (e.g., compliant with the SVP protocol) to access the decoded media stream. Output processor 150 may process individual pixels or groups of pixels in the decoded media stream before providing the final output to display 145. Some examples of output processing performed by output processor 150 include (but are not limited to) filtering techniques (e.g., smoothing, sharpening, etc.), scaling (e.g., from lower resolution to higher resolution), or color modification operations that modify the color of one or more pixels in the decoded output stream. Additionally, one or more graphical elements, such as bounding boxes, markers for features to be identified in the media stream, text-based conversions of audio in the decoded media stream, or other graphical content, may be generated by output processor 150 and displayed along with the decoded media stream. These operations may be performed based on the output of machine learning processor 140.

[0036] Output processor 150 can generate a final media stream comprising modified pixels of the decoded media stream and any resulting graphics at defined locations and sizes, and provide the final stream to display 145 for rendering. In some embodiments, feature extraction system 105 can bypass output processor 150 and provide the decoded media stream to display 145 for rendering without performing additional processing. Output processor 150 can also provide the output stream to other devices (not shown) for output. For example, output processor can transmit the decoded and modified media stream to a mobile device or another computing device via one or more suitable computer communication interfaces.

[0037] Stream decoder 110 may receive encoded bitstreams, for example, from a suitable computer communication interface. For instance, the encoded bitstream may be encoded by a subscription content provider (e.g., a cable television provider), and stream decoder 110 may decode the encoded bitstream to produce a decoded media stream (e.g., raw video or audio content). Stream decoder 110 may include one or more frame buffers. Stream decoder 110 may decode compressed or otherwise encoded bitstreams received from a content provider system or network (not shown). Decoded media content may take the form of one or more decoded video frames. Frame buffers or other memory components of stream decoder 110 or feature extraction system 105 may store decoded frames decoded from the encoded bitstream and any audio channels or audio content. Stream decoder 110 may include any number of decoder components or frame buffer components and may decode multiple encoded bitstreams simultaneously or concurrently to produce multiple decoded media streams (each of which may be processed by feature extraction system 105). Stream decoder 110 may include one or more hardware video decoders that perform decoding operations and subsequently provide the resulting decoded frames (and any decoded audio data) to feature extraction system 105. Stream decoder 110 may include additional buffers (e.g., input buffers for storing the compressed or encoded bitstream before it is decoded by stream decoder 110), network interface, controller, memory, input and output devices, conditional access components, and other components for audio, video, or data processing.

[0038] The stream decoder 110 can obtain one or more encoded bitstreams from an encoded content source. The source can be a content provider, service provider, front-end, camera, storage device, computing system, or any other system capable of providing encoded bitstreams of video, audio, or data content. For example, the encoded bitstream may contain video content (e.g., programs, channels, etc.). The video content may be encoded to have one or more predetermined resolutions (e.g., 2160p vs. 1080p), frame rates (e.g., 60 frames per second vs. 30 frames per second), bit precision (e.g., 10 bits vs. 8 bits), or other video characteristics. The encoded bitstream can be received via a cable television network, satellite, the Internet, cellular network, or other networks. For example, a cable television network may operate via coaxial cable, Bayonet Neill–Concelman (BNC) cable, Ethernet cable (e.g., Category 5, 5e, 6, and 7), fiber optic cable, or other high-data-transmission technologies. The cable television network may connect to a local or remote dish satellite.

[0039] Stream decoder 110 can generate a decoded content stream from one or more received encoded bitstreams. Stream decoder 110 then provides the decoded content to feature extraction system 105. In some embodiments, stream decoder 110 can provide the decoded content stream to display 145, which can be any type of display device. In some embodiments, stream decoder 110 can generate additional video streams with different characteristics from the received encoded bitstreams. For example, stream decoder 110 may include a codec that generates decoded content streams at various resolutions. Such resolutions may include, for example, 4K Ultra High Definition (UHD) resolution (e.g., 3,840 x 2,160 pixels or 2160p). Stream decoder 110 receives an encoded bitstream of 4K / 2160p content as input, decodes it, and generates a low-resolution (e.g., 1080p, 720p, etc.) content stream. Similar operations can be performed to reduce the bit rate of audio data. In another example, stream decoder 110 may produce a content stream with other characteristics (e.g., frame rate, bit precision, bit rate, chroma subsampling). Stream decoder 110 may provide one or more of the produced (e.g., decoded) content streams to feature extraction system 105. Stream decoder 110 may be, may include, or may be a portion of a media decoding pipeline.

[0040] Display 145 can be any type of display device capable of displaying the decoded content stream generated by stream decoder 110. Some non-limiting examples of types of display 145 include televisions, computer monitors, mobile computer displays, projectors, flat panel displays, or handheld user device (e.g., smartphone) displays, etc. In some embodiments, the decoded content stream received by feature extraction system 105 may be forwarded by feature extraction system 105 or otherwise provided to the display. In some embodiments, display 145 may receive and present the decoded content from stream decoder 110. Display 145 may include speakers or other types of audio output devices to output audio data from the decoded media stream, and may also include a video display to present video data from the decoded media stream.

[0041] Referring now to the operation of feature extraction system 105, stream receiver 115 may receive a decoded media stream from media that has received and decoded an encoded media stream (e.g., an encoded bitstream). The decoded media stream may contain video data, audio data, metadata, or any combination thereof. Stream receiver 115 may receive decoded content from stream decoder 110 via one or more suitable communication interfaces. Stream receiver 115 may include one or more buffers or memory areas that store the decoded content for processing by other components of feature extraction system 105. Buffers that store time-stamped frames and audio data received from stream decoder 110 may be accessed by other components of feature extraction system 105 but not by machine learning processor 140 to meet SVP requirements. In some embodiments, stream receiver 115 may provide the decoded media content to display 145. Stream receiver 115 may store the received decoded content stream associated with various parameters and metadata corresponding to the decoded content stream (e.g., timestamp information, duration, resolution, bit rate, decoding rate, etc.). This information may be used by other components of the feature extraction system 105 to perform one or more of the operations described herein. This information may also be used by the display 145 to correctly reproduce or display the decoded media content.

[0042] Once the decoded media content has been received by the stream receiver 115, the feature recognizer 120 can identify a set of features 130 to be generated based on the decoded media stream. For example, the feature recognizer 120 can identify the features 130 to be generated based on the type of content in the decoded media stream. If the decoded media stream contains video content, then the feature recognizer 120 can identify features 130 corresponding to the video data to be generated. Similarly, if the decoded media stream contains audio content, then the feature recognizer 120 can identify features 130 corresponding to the audio data to be generated. The features 130 to be generated can also be identified based on configuration settings (e.g., settings provided in a configuration file or via an external computing system or configured by the user). Configuration settings can specify one or more features 130 to be generated for various types of media content. In some embodiments, the feature recognizer 120 can identify specific media content (e.g., a specific program, a specific channel, a specific song, a specific record, etc.) in the decoded media content and use the identifier of the media content to identify the features 130 to be generated (e.g., from the configuration). For example, some configurations can perform different operations on different media programs with different attributes (e.g., generating motion vectors for sports events, while performing color and audio analysis for dramatic television programs). The attributes of media content can include various parameters of the decoded media stream, including the name of the media stream, the type of the media stream (e.g., video, audio, etc.), the theme of the media stream, and any other attributes of the media stream.

[0043] The stored or accessed configuration settings can specify various types of features 130 to be extracted for a specific media content. In some embodiments, the feature recognizer 120 can query or transmit requests to the machine learning processor 140, which may contain configuration settings for the features 130 to be recognized, including one or more formatting requirements of the media content after its generation. The machine learning processor 140 may contain a program or another type of user-configurable settings specifying one or more features 130 to be generated. The feature recognizer 120 can query the machine learning processor 140 to obtain these features 130 and store identifiers of the features 130 to be generated in the memory of the feature extraction system 105 for use by other components when performing this technology.

[0044] Feature recognizer 120 can identify different features 130 to be generated from video and audio content. For example, for video data, feature recognizer 120 can identify the type of the features 130 to be generated as including brightness histograms, edge detection maps, object detection information (e.g., the presence of one or more objects in a video frame), object classification information (e.g., the classification of one or more objects in a video frame), spectral analysis information, motion vectors, and one or more of various parameters (e.g., minimum pixel value, maximum pixel value, or average pixel value). Parameters may also include parameters or properties of the video content, including resolution, bit rate, and other information. Features 130 generated for audio content can be identified as including, for example, gain information, band histograms, other histograms of various parameters of the audio content, peak information (e.g., minimum, maximum, range values, etc.), pitch information, tone information, harmonic analysis information, the pitch of a song or the beats per minute in a song, and other audio features. The identified feature 130 generated using the decoded media content can be provided to the feature generator 125, which can then generate the identified feature 130 to be generated using the decoded media stream.

[0045] Once the features 130 to be generated have been identified, the feature generator 125 can use the decoded media stream to generate a set of features 130 identified by the feature recognizer 120. To this end, the feature generator 125 can perform one or more processing algorithms on the decoded media stream (e.g., on one or more buffers storing the decoded media content). For example, to generate one or more edge detection maps, the feature generator 125 can apply edge detection techniques to one or more frames of the decoded media content. Some examples of edge detection techniques include (but are not limited to) Canny edge detection, Deriche edge detection, differential edge detection, Sobel edge detection, Prewitt edge detection, and Roberts cross-edge detection, etc. In some implementations, the feature generator 125 can preprocess (e.g., normalization, color adjustment, noise reduction, etc.) the decoded media content before performing processing algorithms to generate features 130. Other algorithms may also be executed to generate additional information. For example, an object detection model (such as a neural network) can be performed on frames of video content to detect the presence of one or more objects. In some implementations, several objects (which can be detected as having a specific class or type) can be detected in one or more frames of video information using such an object detection model.

[0046] Object detection models can use the pixels of a frame as input to perform, for example, in a feedforward neural network or similar artificial intelligence model. Similarly, object classification can be performed on a frame (or a portion of a frame) to classify detected objects. Object classification models can include, for example, CNNs, fully connected neural networks, or other similar artificial intelligence models that can be used to generate object classifications. Feature generator 125 can generate a spectral analysis of the frame regarding video information, for example, by applying one or more spectral analysis algorithms to the pixels of the frame. Performing spectral analysis on the color channel values ​​of pixels in a frame of video data can include performing Fourier transforms or fast Fourier transforms on one or more color channels of the frame. The resulting spectral image can indicate the spectral information of colors in the image, but can be intentionally modified or incomplete to prevent it from being used to reconstruct the frame itself.

[0047] Additionally, feature generator 125 can perform various algorithms on the frame sequence in the video data, for example, to generate features 130 corresponding to the media content at different times. For instance, feature generator 125 can generate features 130 containing motion vectors of various pixels or objects detected in the media content by performing one or more motion estimation algorithms on the pixels of the frame sequence of the decoded media content. Some examples of motion estimation algorithms include block matching algorithms, phase correlation and frequency domain analysis algorithms, pixel recursion algorithms, and optical flow algorithms. In some implementations, feature generator 125 can generate motion vectors by performing corner detection on each frame in a sequence and then using a statistical function (e.g., Random Sample Consensus (RANSAC)) to compare the positions of the corners over time (e.g., between frames) to estimate the motion vectors.

[0048] The features 130 generated for the audio content can be identified as including, for example, gain information, a band histogram, other histograms of various parameters of the audio content, peak information (e.g., minimum, maximum, range values, etc.), pitch information, tone information, harmonic analysis information, the pitch of a song or the beats per minute in a song, and other audio features. The feature generator 125 can also generate features containing various color parameters of frames in the video stream. For example, the feature generator 125 can determine the minimum or maximum pixel intensity value of a frame by comparing the color channel intensity values ​​of each pixel in the frame to each other. Similarly, the feature generator 125 can calculate an average pixel intensity value by summing the color intensity values ​​of each pixel and dividing the sum by the total number of pixels. The feature generator 125 can determine a median pixel intensity value by ranking the pixels by intensity value and then determining the median of the pixels based on the intensity. These operations can be performed for a single color channel or for a combination of color channels. For example, the feature generator 125 can generate a parameter for the maximum pixel intensity value of the "red" channel or a combination of the "red" and "blue" channels. These parameters can be stored by feature generator 125 as part of feature 130.

[0049] Feature generator 125 can also generate features of audio content in a content stream (e.g., an audio channel synchronized to video content or one or more audio channels for stereo or surround sound). Some examples of features 130 that can be generated from audio data may include, for example, gain information, pitch information, beat information (e.g., using one or more beat detection algorithms), a frequency band histogram (e.g., calculated by performing a Fourier transform or FFT algorithm on the audio signal), volume information, peak information (e.g., minimum audio signal value, maximum audio signal value, peak-to-peak difference, etc.), pitch information (e.g., by performing one or more pitch detection algorithms), the pitch of an audio sample (e.g., using a pitch-finding algorithm), or the beats per minute of an audio sample (e.g., using a beat detection algorithm). Feature generator 125 can generate features conforming to a format corresponding to machine learning processor 140. As described above, feature recognizer 120 can query machine learning processor 140 or another configuration setting to identify the format from which features 130 should be generated. Feature generator 125 can receive this formatting information, specifying the structure and content of the data structure that makes up feature 130, and generate feature 130 with an appropriate structure using the techniques described herein. In this way, feature generator 125 can generate feature 130 that is compatible with various customizable machine learning models or processes. Feature generator 125 can store the generated features with an appropriate data structure in the memory of feature extraction system 105.

[0050] Feature provider 135 may provide a set of features 130 generated by feature generator 125 to machine learning processor 140. For example, feature provider 135 may pass features to machine learning processor 140 as one or more features are generated, or in batches once a predetermined number of features 130 are generated. Feature provider 135 may pass features 130 to machine learning processor via a suitable communication interface. In some embodiments, feature provider 135 may pass features to machine learning processor 140 by storing a set of features 130 in a memory area accessible by machine learning processor 140. For example, machine learning processor 140 may include one or more memory components. To provide features 130 to machine learning processor, feature provider 135 may write features 130 into the memory of machine learning processor 140, which machine learning processor 140 may then use the features when executing a machine learning model.

[0051] The machine learning model executed by the machine learning processor 140 can be any type of machine learning model, such as a face detection or object detection model. Other processes can also be executed at the machine learning processor 140 using features 130, such as content fingerprinting or content recommendation. For example, the machine learning processor 140 can execute one or more algorithms or machine learning models that can output content recommendations based on features 130 of content seen by a user over time. The machine learning processor 140 can communicate with one or more other external computing devices (also not shown) via a network (not illustrated). The machine learning processor 140 can transmit the output of the machine learning model or other algorithms to the external computing devices via the network.

[0052] Now, a brief reference. Figure 2 This section describes an example data flow diagram 200 illustrating techniques for generating features from a media stream according to one or more embodiments. As shown in stage 205, encoded content is provided, and said encoded content cannot be decrypted by most computing devices and circuits in any other way, except by a specific stream decoder (e.g., stream decoder 110) configured to decode the encoded content. Next, in stage 210, the encoded content is decoded into decoded content. The decoded content may be raw video or audio information and may contain metadata indicating the nature or attributes of the video or audio information. The decoded content is provided for display at display device 145. In stage 215, the decoded content is processed using the techniques described herein to create feature 130, which may be provided to machine learning processor 140. Finally, in stage 220, machine learning processor 140 executes one or more machine learning algorithms to generate insights related to the media content. As shown, machine learning processor 140 does not access the decoded content and prevents it from accessing this content to comply with SVP requirements.

[0053] Now for reference Figure 3 A flowchart depicts a method 300 for improving the security of media streams. Method 300 may be derived from a feature extraction system 105. Figure 4A and 4B The described computer system 400 or any other computing device described herein executes or otherwise implements the feature extraction system. In a brief overview, the feature extraction system may receive a decoded media stream, identify a set of n features to be generated, generate the k-th feature, determine whether a counter k is equal to the number of features n, increment the counter k, and provide the feature to a machine learning processor.

[0054] In step 302, the feature extraction system may receive a decoded media stream from a media decoding pipeline that receives and decodes the encoded media stream. The decoded media stream may contain video data, audio data, metadata, or any combination thereof. The feature extraction system may receive the decoded content from a stream decoder (e.g., stream decoder 110) via one or more suitable communication interfaces. The feature extraction system may include one or more buffers or memory areas that store the decoded content for processing by other components of the feature extraction system. Buffers that store timestamped frames and audio data received from the stream decoder may be accessed by other components of the feature extraction system but not by the machine learning processor 140 to comply with SVP requirements. In some embodiments, the feature extraction system may provide the decoded media content to a display (e.g., display 145). The feature extraction system may store the received decoded content stream associated with various parameters and metadata (e.g., timestamp information, duration, resolution, bit rate, decoding rate, etc.) corresponding to the decoded content stream. This information may be used by other components of the feature extraction system to perform one or more of the operations described herein. This information may also be used by the display to correctly reproduce or display the decoded media content.

[0055] In step 304, the feature extraction system may identify a set of n features (e.g., feature 130) to be generated based on the decoded media stream. For example, the feature extraction system may identify feature 130 to be generated based on the type of content in the decoded media stream. If the decoded media stream contains video content, then the feature extraction system may identify features corresponding to the video data to be generated. Similarly, if the decoded media stream contains audio content, then the feature extraction system may identify features corresponding to the audio data to be generated. The features to be generated may also be identified based on configuration settings (e.g., settings provided in a configuration file or via an external computing system or configured by the user). Configuration settings may specify one or more features to be generated for various types of media content. In some embodiments, the feature extraction system may identify specific media content in the decoded media content (e.g., a specific program, a specific channel, a specific song, a specific record, etc.) and use the identifier of the media content to identify the features to be generated (e.g., from the configuration). For example, some configurations may perform different operations for different media programs with different attributes (e.g., motion vectors may be generated for sports events, while color and audio analysis may be performed for dramatic television programs). The attributes of media content can include various parameters of the decoded media stream, including the name of the media stream, the type of the media stream (e.g., video, audio, etc.), the subject of the media stream, and any other attributes of the media stream.

[0056] The stored or accessed configuration settings can specify various types of features to be extracted for a particular media content. In some implementations, the feature extraction system can query or transmit requests to a machine learning processor (e.g., machine learning processor 140), which may contain configuration settings for the features to be identified, including one or more formatting requirements of the media content after its generation. The machine learning processor may contain a program or another type of user-configurable settings specifying one or more features to be generated. The feature extraction system can query the machine learning processor to obtain these features and store identifiers of the features to be generated in the memory of the feature extraction system 105 for use by other components when performing this technology.

[0057] Feature extraction systems can identify different features to be generated from video and audio content. For example, for video data, a feature extraction system can identify the types of features to be generated as including brightness histograms, edge detection maps, object detection information (e.g., the presence of one or more objects in a video frame), object classification information (e.g., the classification of one or more objects in a video frame), spectral analysis information, motion vectors, and one or more parameters (e.g., minimum pixel value, maximum pixel value, or average pixel value). Parameters may also include parameters or properties of the video content, including resolution, bitrate, and other information. Features generated for audio content can be identified as including, for example, gain information, band histograms, other histograms of various parameters of the audio content, peak information (e.g., minimum, maximum, range values, etc.), pitch information, tone information, harmonic analysis information, the pitch of a song or the beats per minute in a song, and other audio features. The feature extraction system can list several identified features n.

[0058] In step 306, the feature extraction system can generate the k-th feature using the decoded media stream. In some implementations, the feature extraction system can generate a feature by iterating through a list of features to be generated identified by the feature extraction system according to a counter register k, and then generating the k-th feature in the list. In some implementations, the feature extraction system can extract one or more features in parallel, rather than using a sequential process. To generate features, the feature extraction system can perform one or more processing algorithms on the decoded media stream (e.g., on one or more buffers storing the decoded media content). For example, to generate one or more edge detection maps, the feature extraction system can apply edge detection techniques to one or more frames of the decoded media content. Some examples of edge detection techniques include (but are not limited to) Canny edge detection, Deriche edge detection, differential edge detection, Sobel edge detection, Prewitt edge detection, and Roberts cross edge detection, etc. In some implementations, the feature extraction system may preprocess (e.g., normalization, color adjustment, noise reduction, etc.) the decoded media content before executing processing algorithms to produce the k-th feature. Other algorithms may also be executed to produce additional information. For example, an object detection model (e.g., a neural network) may be performed on frames of video content to detect the presence of one or more objects. In some implementations, several objects (which may be detected as having a specific class or type) can be detected in one or more frames of the video information using such an object detection model.

[0059] Object detection models can use the pixels of a frame as input to perform, for example, in a feedforward neural network or similar artificial intelligence model. Similarly, object classification can be performed on frames (or portions of frames) to classify detected objects. Object classification models can include, for example, CNNs, fully connected neural networks, or other similar artificial intelligence models that can be used to generate object classifications. Feature extraction systems can generate a spectral analysis of a frame about video information, for example, by applying one or more spectral analysis algorithms to the pixels of the frame. Performing spectral analysis on the color channel values ​​of pixels in a frame of video data can involve performing Fourier transforms or fast Fourier transforms on one or more color channels of the frame. The resulting spectral image can indicate the spectral information of the colors in the image, but cannot contain enough information to reconstruct the frame itself (e.g., knowing that a frame contains n pixels of a certain brightness and y pixels of another brightness would not allow us to reconstruct the source image without more information).

[0060] Furthermore, feature extraction systems can perform various algorithms on sequences of frames in video data, for example, to generate features of media content corresponding to time. For instance, a feature extraction system can generate features containing motion vectors of various pixels or objects detected in the media content by performing one or more motion estimation algorithms on the pixels of a sequence of frames of decoded media content. Some examples of motion estimation algorithms include block matching algorithms, phase correlation and frequency domain analysis algorithms, pixel recursion algorithms, and optical flow algorithms. In some implementations, a feature extraction system can generate motion vectors by performing corner detection on each frame in a sequence and subsequently using a statistical function (e.g., Random Sample Consensus (RANSAC)) to compare the positions of the corners over time (e.g., between frames) to estimate motion vectors.

[0061] Features generated for audio content can be identified as including, for example, gain information, band histograms, other histograms of various parameters of the audio content, peak information (e.g., minimum, maximum, range values, etc.), pitch information, tone information, harmonic analysis information, the pitch of a song or the beats per minute of a song, and other audio features. The feature extraction system can also generate features containing various color parameters of frames in a video stream. For example, the feature extraction system can determine the minimum or maximum pixel intensity value of a frame by comparing the color channel intensity values ​​of each pixel in the frame to each other. Similarly, the feature extraction system can calculate the average pixel intensity value by summing the color intensity values ​​of each pixel and dividing the sum by the total number of pixels. The feature extraction system can determine the median pixel intensity value by ranking the pixels by intensity value and then determining the median of the pixels based on said intensity. These operations can be performed for a single color channel or for a combination of color channels. For example, the feature extraction system can generate parameters for the maximum pixel intensity value of the "red" channel or a combination of the "red" and "blue" channels. These parameters can be stored by the feature extraction system as part of a feature.

[0062] The feature extraction system can also generate features of audio content in a content stream (e.g., an audio channel synchronized to video content or one or more audio channels for stereo or surround sound). Some examples of features that can be generated from audio data may include, for example, gain information, pitch information, cadence information, beat information (e.g., using one or more beat detection algorithms), frequency band histograms (e.g., calculated by performing a Fourier transform or FFT algorithm on the audio signal), volume information, peak information (e.g., minimum audio signal value, maximum audio signal value, peak-to-peak difference, etc.), pitch information (e.g., by performing one or more pitch detection algorithms), the pitch of an audio sample (e.g., using a pitch-finding algorithm), or the beats per minute of an audio sample (e.g., using a beat detection algorithm). The feature extraction system can generate features that conform to a format corresponding to the machine learning processor 140. As described above, the feature recognizer 120 can query the machine learning processor 140 or another configuration setting to identify the format in which features should be generated. The feature extraction system can receive this formatted information, which specifies the structure and content of the data structure that makes up the features, and uses the techniques described herein to generate features that conform to an appropriate structure. In this way, the feature extraction system can generate features that are compatible with a variety of customizable machine learning models or processes. The feature extraction system can store the generated features with an appropriate data structure in the memory of the feature extraction system 105.

[0063] In step 308, the feature extraction system determines whether the counter k is equal to the number of features n. To determine whether each of the n features has been generated, the feature extraction system compares the counter register k, which tracks the number of features generated, with the total number of features to generate n. If the counter register k is not equal to (e.g., less than) the number n, the feature extraction system performs step 310 to increment the counter register k and continue generating features. If the counter register k is equal to (e.g., equal to or greater than) the total number of labels in subset n, the feature extraction system performs step 312 to provide the generated features to the machine learning processor.

[0064] In step 310, the feature extraction system increments the counter k. To track the total number of features generated from the decoded media content, the feature extraction system increments the counter register k by 1 to identify the next feature in a set of features to be generated. After incrementing the value of the counter register k, the feature extraction system performs step 306 to generate the next feature in a set of features.

[0065] In step 312, the feature extraction system may provide a set of features to the processor executing the machine learning model. Access to the decoded media stream may be prevented from being accessed by the processor executing the machine learning model. To this end, the feature extraction system may pass features to the machine learning processor as one or more features are generated, or in batches once a predetermined number of features are generated. The feature extraction system may pass features to the machine learning processor via a suitable communication interface. In some embodiments, the feature extraction system may pass features to the machine learning processor by storing a set of features in a memory area accessible to the machine learning processor. For example, the machine learning processor may include one or more memory components. To provide features to the machine learning processor, the feature extraction system may write features to the machine learning processor's memory, which the machine learning processor may then use when executing the machine learning model.

[0066] Therefore, the implementation of the systems and methods discussed in this paper allows for further analysis of video data by machine learning systems without exposing any part of the media data itself to the machine learning system. This enhances the security of media content and improves compliance with SVP requirements and copyright protection. Furthermore, by integrating a feature extraction system into the media processing pipeline, the systems and methods described in this paper extend the functionality of conventional set-top boxes or media processing systems by enabling the integration of configurable processors into the media processing pipeline. Thus, the systems and methods described in this paper provide improvements over conventional media processing techniques.

[0067] Figure 4A and 4B A block diagram depicts a computing device 400 that can be used to implement the computing device described herein. (See diagram for example.) Figure 4A and 4B As shown, each computing device 400 includes a central processing unit 421 and a main memory unit 422. For example... Figure 4A As shown, computing device 400 may include storage device 428, mounting device 416, network interface 418, I / O controller 423, display devices 424a to 424n, keyboard 426, and pointing device 427, such as a mouse. Storage device 428 may include (unlimited) operating system and / or software. Figure 4B As shown, each computing device 400 may also include additional optional elements such as memory port 403, bridge 470, one or more input / output devices 430a to 430n (generally referred to by reference element symbol 430), and cache memory 440 in communication with central processing unit 421.

[0068] Central processing unit 421 is any logic circuit system that responds to and processes instructions fetched from main memory unit 422. In many embodiments, central processing unit 421 is provided by, for example, a microprocessor unit manufactured by Intel Corporation of Mountain View, California; a microprocessor unit manufactured by International Business Machines of White Plains, New York; a microprocessor unit manufactured by Advanced Micro Devices of Sunnyvale, California; or a microprocessor unit manufactured by Advanced RISC Machines (ARM). Computing device 400 may be based on any of these processors or any other processor capable of operating as described herein.

[0069] Main memory unit 422 may be one or more memory chips capable of storing data and allowing direct access from any storage location by microprocessor 421, such as any type or variant of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD). Main memory 422 may be based on any of the memory chips described above or any other available memory chip capable of operating as described herein. Figure 4A In the embodiment shown, the processor 421 communicates with the main memory 422 via the system bus 450 (described in more detail below). Figure 4B An embodiment of computing device 400 is depicted, in which the processor communicates directly with main memory 422 via memory port 403. For example, in Figure 4B In this context, the main memory 422 can be DRDRAM.

[0070] Figure 4B This illustration depicts an embodiment in which the main processor 421 communicates directly with the cache memory 440 via a secondary bus, sometimes referred to as the back-side bus. In other embodiments, the main processor 421 communicates with the cache memory 440 using a system bus 450. The cache memory 440 typically has a faster response time than the main memory 422 and is provided by, for example, SRAM, BSRAM, or EDRAM. Figure 4BIn the embodiment shown, processor 421 communicates with various I / O devices 430 via local system bus 450. Various buses can be used to connect central processing unit 421 to any of the I / O devices 430, such as VESA VL bus, ISA bus, EISA bus, Micro Channel Architecture (MCA) bus, PCI bus, PCI-X bus, PCI-High Speed ​​bus, or NuBus. For an embodiment where the I / O device is a video display 424, processor 421 can use the Advanced Graphics Port (AGP) to communicate with display 424. Figure 4B An implementation of computer 400 is described, wherein the main processor 421 can communicate directly with I / O device 430b, for example, via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communication technologies. Figure 4B An implementation scheme in which local bus and direct communication are mixed is also described: processor 421 communicates with I / O device 430a using local interconnect bus, while communicating directly with I / O device 430b.

[0071] A variety of I / O devices 430a to 430n may exist in the computing device 400. Input devices include keyboards, mice, trackpads, trackballs, microphones, dial pads, touchpads, touch screens, and drawing tablets. Output devices include video displays, speakers, inkjet printers, laser printers, projectors, and dye-to-sublimation printers. I / O devices can be... Figure 4A The I / O controller 423 shown controls the device. The I / O controller can control one or more I / O devices, such as a keyboard 426 and a pointing device 427 (e.g., a mouse or optical pen). Additionally, the I / O devices can provide storage devices and / or mounting media 416 for the computing device 400. In other embodiments, the computing device 400 can provide a USB connection (not shown) to receive a handheld USB storage device, such as a USB flash drive series device manufactured by Twintech Industry, Inc. of Los Alamitos, California.

[0072] Refer again Figure 4AThe computing device 400 may support any suitable installation device 416, such as a disk drive, CD-ROM drive, CD-R / RW drive, DVD-ROM drive, flash drive, tape drive of various formats, USB device, hard disk drive, network interface, or any other device suitable for installing software and programs. The computing device 400 may further include storage devices for storing the operating system and other related software, and for storing application software programs (e.g., any program or software 420 for implementing (e.g., configured and / or designed for) the systems and methods described herein), such as one or more hard disk drives or a redundant array of independent disks. Optionally, any of the installation devices 416 may also be used as storage devices. Furthermore, the operating system and software may run from bootable media.

[0073] Furthermore, the computing device 400 may include a network interface 418 to interface with the network 404 via various connections, including but not limited to standard telephone lines, LAN or WAN links (e.g., 402.11, T1, T3, 56kb, X.25, SNA, DECNET), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over-SONET), wireless connections, or any or all of the above connections. Connections can be established using various communication protocols (e.g., TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 402.11, IEEE 402.11a, IEEE 402.11b, IEEE 402.11g, IEEE 402.11n, IEEE 402.11ac, IEEE 402.11ad, CDMA, GSM, WiMax, and direct asynchronous connections). In one embodiment, computing device 400 communicates with other computing devices 400' via any type and / or form of gateway or tunneling protocol (e.g., Secure Sockets Layer (SSL) or Transport Layer Security (TLS)). Network interface 418 may include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem, or any other device suitable for interfacing computing device 400 to any type of network capable of communicating and performing the operations described herein.

[0074] In some embodiments, computing device 400 may include or be connected to one or more display devices 424a to 424n. Therefore, any of the I / O devices 430a to 430n and / or the I / O controller 423 may include any type and / or form of suitable hardware, software, or a combination of hardware and software to support, enable, or provide computing device 400 with connectivity to and use of display devices 424a to 424n. For example, computing device 400 may include any type and / or form of video adapter, video card, driver, and / or library to interface with, transmit, connect to, or otherwise use display devices 424a to 424n. In one embodiment, a video adapter may include multiple connectors to interface with display devices 424a to 424n. In other embodiments, computing device 400 may include multiple video adapters, each connected to display devices 424a to 424n. In some embodiments, any portion of the operating system of computing device 400 may be configured to use multiple displays 424a to 424n. Those skilled in the art will recognize and understand that the computing device 400 can be configured to have one or more display devices 424a to 424n in various ways and implementations.

[0075] In another embodiment, I / O device 430 may serve as a bridge between system bus 450 and an external communication bus, such as a USB bus, Apple Desktop bus, RS-232 serial connection, SCSI bus, FireWire bus, FireWire 400 bus, Ethernet bus, AppleTalk bus, Gigabit Ethernet bus, Asynchronous Transfer Mode bus, Fibre Channel bus, Serial Attached Small Computer System Interface bus, USB connection, or HDMI bus.

[0076] The embodiments of the subject matter and operation described in this specification can be implemented in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or a combination thereof, embodied in a digital electronic circuit system or on a tangible medium. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, such as one or more components of computer program instructions, encoded on a computer storage medium for execution by or control of a data processing device. Program instructions can be encoded on artificially generated propagated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium can be or be contained in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagated signal, it can contain a source or destination of computer program instructions encoded in artificially generated propagated signals. The computer storage medium can also be or be contained in one or more individual physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0077] The features disclosed herein can be implemented in a smart TV module (or connected TV module, hybrid TV module, etc.), which may include components configured to integrate Internet connectivity with more traditional television program sources (e.g., received via cable, satellite, over-the-air, or other signals). The smart TV module may be physically integrated into a television set or may comprise a separate device, such as a set-top box, Blu-ray or other digital media player, game console, hotel TV system, and other compatible devices. The smart TV module may be configured to allow viewers to watch videos, movies, photos, and other content on the internet, local cable TV channels, satellite TV channels, or stored on a local hard drive. A set-top box (STB) or set-top unit (STU) may include an information electronics device that may contain a tuner and be connected to the television set and external signal sources to convert signals into content, which is then displayed on a television screen or other display device.

[0078] The operations described in the specification can be implemented as operations performed by a data processing device on data stored in one or more computer-readable storage devices or received from an external source.

[0079] The terms “data processing device,” “feature extraction system,” “data processing system,” “client device,” “computing platform,” “computing device,” or “device” encompass all kinds of devices, apparatuses, and machines used for processing data, including (as examples) programmable processors, computers, system-on-a-chip or system-on-multi-chip systems, or combinations thereof. Devices may contain special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, devices may also contain code that creates the execution environment for the computer programs discussed, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. Devices and execution environments can implement various computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0080] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may (but does not need to) correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program under discussion, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on one or more computers located at a single point or distributed across multiple points interconnected by a communication network.

[0081] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input data and producing outputs. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit)), and the device can be implemented as said dedicated logic circuit system.

[0082] As an example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any type of digital computer processors or one or more processors. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The components of a computer include a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to said one or more mass storage devices, or both. However, a computer does not necessarily need to have such devices. Furthermore, a computer may be embedded in another device, such as (for example) a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive). Applicable devices for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including (as examples) semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry systems.

[0083] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), plasma, or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user may include any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user in any form, including acoustic, voice, or tactile input, can be received.

[0084] The embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a client computer having a graphical user interface or web browser through which a user can interact with the embodiments of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), intranets (e.g., the Internet) and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0085] For example, the computing system of feature extraction system 105 may include clients and servers. For instance, feature extraction system 105 may include one or more servers in one or more data centers or server farms. Clients and servers are typically geographically distant and typically interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., HTML pages) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving input from the user). Data generated at the client device (e.g., interactions, calculations, or the results of any other events or calculations) can be received from the client device at the server and vice versa.

[0086] While this invention contains numerous details of specific embodiments, these should not be construed as limiting any invention or claim, but rather as descriptions of features specific to particular embodiments of the systems and methods described herein. Certain features described in this specification within the context of individual embodiments may also be implemented in combination with a single embodiment. Conversely, various features described in the context of individual embodiments may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may be removed from said combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0087] Similarly, although the operations are depicted in a specific order in the figures, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or that all illustrated operations are performed to achieve the desired result. In some cases, the actions cited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result.

[0088] In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the embodiments described above should not be construed as requiring this separation in all embodiments, and it should be understood that the described program components and systems can typically be integrated together in a single software product or packaged into multiple software products. For example, the feature extraction system 105 may be a single module or a logic device with one or more processing modules.

[0089] Several illustrative embodiments and implementation schemes have been described. It is clear that the foregoing is illustrative and non-limiting, and has been presented as examples. Specifically, although many of the examples presented herein involve specific combinations of method actions or system elements, those actions and elements can be combined in other ways to accomplish the same objective. Actions, elements, and features discussed in connection with only one embodiment are not intended to exclude them from similar roles in other embodiments or implementation schemes.

[0090] The phrases and terms used herein are for descriptive purposes and should not be considered restrictive. The use of “comprising,” “including,” “having,” “containing,” “involving,” “characterized by / characterized in that,” and variations thereof, herein means to cover all items listed thereafter, their equivalents and additional items, and alternative embodiments consisting only of those listed thereafter. In one embodiment, the system and method described herein consist of one, every combination of, or all of the described elements, actions, or components.

[0091] Any reference to an embodiment, element, or action of the system or method herein cited in the singular may also include embodiments comprising multiple such elements, and any reference to any embodiment, element, or action herein cited in the plural may also include embodiments comprising only a single element. References in either the singular or plural form are not intended to limit the currently disclosed system or method, its components, actions, or elements to a singular or plural configuration. A reference to any action or element based on any information, action, or element may include that said action or element is at least in part based on an embodiment of that information, action, or element.

[0092] Any embodiment disclosed herein may be combined with any other embodiment, and references to “implementation,” “some embodiments,” “alternative embodiments,” “various embodiments,” “an embodiment,” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment. Such terms used herein do not necessarily all refer to the same embodiment. Any embodiment may be combined with any other embodiment inclusively or exclusively in any manner consistent with the aspects and embodiments disclosed herein.

[0093] A reference to "or" can be understood as inclusive, such that any term described using "or" can refer to any one, more than one, or all of the described items.

[0094] Where a reference numeral follows a technical feature in the drawings, detailed description, or any claim, the reference numeral is included solely for the purpose of enhancing the comprehensibility of the drawings, detailed description, and claims. Therefore, the reference numeral and its presence do not have any limiting effect on the scope of any claim element.

[0095] The systems and methods described herein may be embodied in other specific forms without departing from their characteristics. While the examples provided may be useful for improving the security of media streaming, the systems and methods described herein can be applied to other environments. The foregoing embodiments are illustrative and not limiting of the systems and methods described. The scope of the systems and methods described herein is therefore indicated by the appended claims, rather than by the foregoing description, and variations derived from the equivalent meaning and scope of the claims are included therein.

Claims

1. A method for improving the security of a media stream, the method being performed within a system for improving the security of a media stream, the system including at least a media processing pipeline and hardware components integrated into the media processing pipeline, the method comprising: The hardware components receive a decoded media stream from the media decoding pipeline that receives and decodes an encoded media stream that needs to be protected against unauthorized access; The hardware component identifies a set of features to be generated based on the decoded media stream; The hardware component generates the set of features using the decoded media stream, and stores the set of features in a first memory area accessible by a processor executing a machine learning model, while storing the decoded media stream in a second memory area protected from access by the processor executing the machine learning model, wherein the generated set of features cannot be used to reconstruct the decoded media stream; as well as The hardware component using the first memory area provides the set of features to the processor executing the machine learning model, wherein the hardware component prevents the processor executing the machine learning model from accessing the second memory area storing the decoded media stream.

2. The method of claim 1, wherein the decoded media stream comprises video data, and wherein the set of features comprises a luminance histogram of the video data.

3. The method of claim 1, wherein identifying the set of features further comprises: The type of media in the decoded media stream is determined by the hardware component; and The hardware component identifies the set of features based on the type of media in the decoded media stream.

4. The method of claim 1, wherein identifying the set of features further comprises: The hardware component extracts the raw video data from the decoded media stream; and The hardware component identifies the set of features based on the attributes of the original video data.

5. The method of claim 1, wherein the hardware component is a field-programmable gate array or an application-specific integrated circuit.

6. The method of claim 1, wherein generating the set of features further comprises applying edge detection technology or spectrum analysis technology to the decoded media stream by the hardware components.

7. The method of claim 1, wherein generating the set of features further comprises the hardware component determining one or more of a minimum pixel intensity value, a maximum pixel intensity value, an average pixel intensity value, or a median pixel intensity value based on one or more color channels of the decoded media stream.

8. The method of claim 1, wherein generating the set of features further comprises detecting motion vectors in the decoded media stream by the hardware component.

9. The method of claim 1, wherein the processor of the machine learning model lacks either permission to access the second memory area or a hardware connection.

10. The method of claim 1, further comprising generating an output media stream comprising an output based on the machine learning processor and a graph generated from the decoded media stream by the hardware component; and The output media stream is provided by the hardware components for presentation at the display.

11. A system for improving the security of media streams, comprising: A media decoding pipeline is needed to receive and decode encoded media streams to prevent unauthorized access. and Hardware components integrated into the media processing pipeline, the hardware components being configured to: Receive decoded media streams from the media decoding pipeline; Based on the set of features to be generated by the decoded media stream identification; The set of features is generated using the decoded media stream; The set of features is stored in a first memory area accessible by a processor executing the machine learning model, while the decoded media stream is stored in a second memory area protected from access by the processor executing the machine learning model, wherein the generated set of features cannot be used to reconstruct the decoded media stream; as well as The set of features is provided to the processor executing the machine learning model via the first memory area, wherein the hardware components prevent the processor executing the machine learning model from accessing the second memory area storing the decoded media stream.

12. The system of claim 11, wherein the decoded media stream comprises video data, and wherein the set of features comprises a luminance histogram of the video data.

13. The system of claim 11, wherein, in order to identify the set of features, the hardware components are further configured to: Determine the type of media in the decoded media stream; and The set of features is identified based on the type of media in the decoded media stream.

14. The system of claim 11, wherein, in order to identify the set of features, the hardware components are further configured to: Extract raw video data from the decoded media stream; and The set of features is identified based on the attributes of the original video data.

15. The system of claim 11, wherein the hardware component is a field-programmable gate array or an application-specific integrated circuit.

16. The system of claim 11, wherein, in order to generate the set of features, the hardware components are further configured to apply edge detection technology or spectrum analysis technology to the decoded media stream.

17. The system of claim 11, wherein, in order to generate the set of features, the hardware components are further configured to determine one or more of a minimum pixel intensity value, a maximum pixel intensity value, an average pixel intensity value, or a median pixel intensity value based on one or more color channels of the decoded media stream.

18. The system of claim 11, wherein the processor of the machine learning model lacks either permission to access the second memory area or a hardware connection.

19. A circuit for improving the security of a media stream, comprising: A bitstream decoder is a hardware component integrated into the media processing pipeline that receives encoded bitstreams that need to be protected from unauthorized access and generates decoded media streams that need to be protected from unauthorized access. and A processing component, which is a hardware component integrated into the media processing pipeline, receives the decoded media stream from the bitstream decoder. The processing component is configured to: Based on the set of features to be generated by the decoded media stream identification; The set of features is generated using the decoded media stream; The set of features is stored in a first memory area accessible by a processor executing the machine learning model, while the decoded media stream is stored in a second memory area protected from access by the processor executing the machine learning model, wherein the generated set of features cannot be used to reconstruct the decoded media stream; as well as The set of features is provided to a second processor that executes the machine learning model, wherein the second processor executing the machine learning model is prevented from accessing the second memory area storing the decoded media stream.

20. The circuit of claim 19, wherein the second processor of the machine learning model lacks either permission to access the second memory region or a hardware connection thereof.

Citation Information

Patent Citations

  • Real person verifying method and system based on machine learning

    CN106650555A