Facial Recognition System and Method for Media Playback Devices

CN118413721BActive Publication Date: 2026-08-14AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

此类方法通常缓慢、昂贵,且可能无法使用

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118413721B_ABST
    Figure CN118413721B_ABST
Patent Text Reader

Abstract

This application relates to a facial recognition system and method for a media playback device. One method includes: accessing a video comprising a set of frames; receiving a command for identifying a target displayed in a first frame; generating first target data based on the target; generating a first confidence level between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; outputting the identity data of the second target data in response to the first confidence level being at or above a confidence level threshold; otherwise, generating third target data based on the target displayed in the second frame of the set of frames; generating a second confidence level of similarity between the third target data and the second target data stored in the remote database; and outputting the identity data of the second target data in response to the second confidence level being at or above the confidence level threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein may generally relate to systems and methods for video processing, and more specifically, to facial recognition for media devices. Background Technology

[0002] Media devices (e.g., set-top boxes) can support playback of video content. When watching a video, a user may want to identify the people shown in the video and obtain more information about them. One approach is to utilize cloud-based facial recognition services by uploading the video of interest to the cloud. Such methods are typically slow, expensive, and may not be available. An aspect of this technology includes an edge-based facial recognition process that provides a local processing flow. Summary of the Invention

[0003] An exemplary embodiment includes a method comprising accessing a video comprising a set of frames; receiving a command for identifying a target displayed in a first frame of the set of frames; generating first target data based on the target displayed in the first frame; generating a first confidence level of similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; outputting the identity data of the second target data in response to the first confidence level being at or above a confidence level threshold; and generating third target data based on the target displayed in the second frame of the set of frames in response to the first confidence level being at or above the confidence level threshold; generating a second confidence level of similarity between the third target data and the second target data stored in the remote database; outputting the identity data of the second target data in response to the second confidence level being at or above the confidence level threshold; and determining that the target has not been identified in response to the second confidence level being below the confidence level threshold.

[0004] An exemplary embodiment also includes a system comprising a processor. The processor is configured to perform operations including accessing a video comprising a set of frames; receiving a command for identifying a target displayed in a first frame of the set of frames; generating first target data based on the target displayed in the first frame; generating a first confidence level of similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; outputting the identity data of the second target data in response to the first confidence level being at or above a confidence level threshold; and generating third target data based on the target displayed in the second frame of the set of frames in response to the first confidence level being at or above the confidence level threshold; generating a second confidence level of similarity between the third target data and the second target data stored in the remote database; outputting the identity data of the second target data in response to the second confidence level being at or above the confidence level threshold; and determining that the target has not been identified in response to the second confidence level being below the confidence level threshold.

[0005] An exemplary embodiment further includes a system comprising a processor. The processor is configured to perform operations including accessing a video comprising a set of frames; receiving a command for identifying a target displayed in a first frame of the set of frames; generating first target data based on the target displayed in the first frame; generating a first confidence level of similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; in response to the first confidence level being at or above a confidence level threshold: displaying the identity data of the second target data; storing the second target data in a local database; in response to the first confidence level being below the confidence level threshold: generating third target data based on the target displayed in the second frame of the set of frames; generating a second confidence level of similarity between the third target data and the second target data stored in the remote database; in response to the second confidence level being at or above the confidence level threshold, outputting the identity data of the second target data; and in response to the second confidence level being below the confidence level threshold, determining that the target has not been identified. Attached Figure Description

[0006] Some features of this technology are set forth in the appended claims. However, for illustrative purposes, several exemplary embodiments of this technology are illustrated in the following drawings.

[0007] Figure 1 Describe the exemplary network configuration of the media device according to one or more exemplary implementation schemes.

[0008] Figure 2 Describe an exemplary computing system for a media device based on one or more exemplary implementation schemes.

[0009] Figure 3 A schematic diagram illustrating a machine learning model for facial recognition based on one or more exemplary implementation schemes.

[0010] Figure 4 A schematic diagram illustrating a multichannel manager based on one or more exemplary implementations.

[0011] Figure 5 This describes instance frames of a video based on one or more exemplary implementation schemes.

[0012] Figure 6 A schematic diagram illustrating a process for facial recognition in a video according to one or more exemplary implementation schemes.

[0013] Figure 7 A flowchart illustrating a process for facial recognition in a video according to one or more exemplary implementation schemes.

[0014] The figures depict various implementation schemes for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative implementation schemes of the structures and methods described herein can be employed without departing from the principles described herein.

[0015] However, not all depicted components are usable in all embodiments, and one or more embodiments may include additional components or components different from those shown in the figures. Variations in the arrangement and type of components may be made without departing from the spirit or scope of the claims set forth herein. Additional components, different components, or fewer components may be provided. Detailed Implementation

[0016] The detailed description set forth below is intended as a description of various configurations of the present technology, and not as representing the only configuration in which the present technology can be practiced. The accompanying drawings are incorporated herein and form part of the detailed description. The detailed description contains specific details for the purpose of providing a thorough understanding of the present technology. However, the present technology is not limited to the specific details set forth herein, and can be practiced using one or more other embodiments. In one or more embodiments, structures and components are shown in block diagram form to avoid obscuring the concept of the present technology.

[0017] When watching videos, users may want to identify the people shown in the video and obtain more information about them. For example, in a movie, a user might want to identify a new actor and learn about other movies starring that actor; in a football match, a user might want to identify players and view their statistics; in a news report, a user might want to identify a politician and view their background. One method for identifying people shown in a video and obtaining information about a target person involves transmitting frames (e.g., video frames) showing the target person to a remote server for identification. However, such methods can be slow, expensive, and may not be usable. An aspect of this technology includes an edge-based facial recognition process that provides a local processing flow to address this issue. This aspect may be based on facial detection, tracking, and / or recognition technologies that can be implemented in a system-on-a-chip (SoC).

[0018] Figure 1 This describes an exemplary network configuration 100 for a media device according to one or more exemplary embodiments. The media device 102 may include a set-top box, digital media player, streaming media adapter, or any other device capable of presenting video via video signal 112 and / or receiving commands 114 from user 118. The video may be a file, stream, or any other format accessible locally and / or remotely (e.g., from server 106). Command 114 may be a process performed by the media device 102, such as initiating a facial recognition process, pausing video, identifying an object in the video, requesting more information about an object in the video, and any other video-related commands. The media device 102 may be connected to an electronic display 120 (e.g., a television set) that can receive the video signal 112 that can be used to present a face to user 118.

[0019] A start frame (also referred to herein as the current frame) can be selected via a user interface (e.g., a remote control or microphone) on media device 102, which can generate a start signal that causes media device 102 to begin face recognition. Faces in the start frame can be detected using face detection and / or recognition, and facial data of the detected faces can be generated. If the identity of the detected face cannot be determined based on facial data from a single frame, a bidirectional refinement process can be performed, where neighboring frames are used to generate additional facial data. This technology can automatically determine the number of frames used to complete face detection and recognition and can also index the frames for display, or abandon detection efforts if no face can be detected in the current and neighboring frames. While one use case might be real-time or on-demand video for face recognition, this technology can also be applied in other applications, such as home surveillance and ad verification. For example, a system can identify unknown objects in a user's home by comparing faces captured via a security camera with the faces of family members. As another example, advertisers can verify that their ads have been decoded on a client's media device by comparing faces in decoded video with faces in an advertisement.

[0020] Media device 102 can perform all or most of the video analysis locally (also referred to herein as "edge" or "on-device") to reduce the workload that could be performed remotely, for example, by a cloud server (also referred to herein as "external" or "off-device"). Media device 102 may include one or more machine learning models for face detection, tracking, and / or identification, which can be executed locally. One or more machine learning models can analyze one or more video frames to identify one or more objects. One or more machine learning models can also extract facial data of the identified objects. Media device 102 can cross-reference the facial data with a database of object identities represented by the facial data. Facial data may include images, features (e.g., the location of ears, eyes, mouth, etc.), digital representations (e.g., embeddings, vectors, etc.), and / or any other data related to the face (hereinafter "facial data"). The database may be located on media device 102, server 106, another media device 108, or any other network location via network 104. Transmitting facial data can utilize less bandwidth than video data (e.g., frames), which can improve the efficiency of video analytics.

[0021] Facial data can be stored in local storage (e.g., system memory, flash drive, etc.). When subsequently presenting video to user 118, media device 102 can refer to the local storage to save computational power. Facial data can also or alternatively be stored on a remote server (e.g., server 106). In an example where another media device 108 presents video, that other media device 108 can refer to the remote storage to save computational power.

[0022] Figure 2 This describes a computing system 200 according to one or more exemplary embodiments. The computing system 200 may be a media device 102 and / or a part of a media device 102, such as... Figure 1 As shown in the figure. The computing system 200 may include various types of computer-readable media and interfaces for various other types of computer-readable media. The computing system 200 includes a bus 210, a processing unit 220, a storage device 202, a system memory 204, an input device interface 206, an output device interface 208, a face detector 212, a face data extractor 214, a graphics engine 216, and / or a network interface 218.

[0023] Bus 210 collectively represents all system, peripheral, and chipset buses that communicatively connect various components of computing system 200. In one or more embodiments, bus 210 communicatively connects processing unit 220 to other components of computing system 200. From various memory locations, processing unit 220 retrieves instructions to be executed and data to be processed in order to perform the operations of this disclosure. In various embodiments, processing unit 220 may be a controller and / or a single-core or multi-core processor.

[0024] Bus 210 is also connected to input device interface 206 and output device interface 208. Input device interface 206 enables the system to receive input. For example, input device interface 206 allows a user to convey information and select commands on system 200. Input device interface 206 can be used with input devices such as keyboards, mice, and other user input devices, as well as microphones, cameras, and other sensor devices. Output device interface 208 enables the display of frames (e.g., images) generated by computing system 200. Output devices that can be used with output device interface 208 may include, for example, printers and display devices such as liquid crystal displays (LCDs), light-emitting diode (LED) displays, organic light-emitting diode (OLED) displays, flexible displays, flat panel displays, solid-state displays, projectors, or any other device for outputting information. One or more embodiments may include devices that function as both input and output devices, such as touchscreens.

[0025] Bus 210 also couples system 200 to one or more networks (e.g., network 104) and / or one or more network nodes via network interface 218. Network interface 218 may include one or more interfaces that allow system 200 to be part of a network of computers (e.g., a local area network (LAN), a wide area network (WAN), or a network of networks (the Internet)). Any or all components of system 200 may be used in conjunction with this disclosure.

[0026] The graphics engine 216 may be hardware and / or software for manipulating frames (e.g., video frames). The graphics engine 216 may prepare frames for presentation via the output device interface 208. The graphics engine 216 may crop, scale, or otherwise manipulate frames to meet the input requirements of one or more machine learning models (e.g., face detector 212 and face data extractor 214) utilized in this technology. The graphics engine 216 may also include a video decoder for receiving frames from an encoded image buffer and outputting decoded frames to a decoded image buffer. The graphics engine 216 may also include a post-processing engine for receiving frames from the decoded image buffer and outputting frames to a display image for displaying the frames on an electronic display (e.g., output device interface 208). The post-processing engine may also convert frames from the decoded image buffer to match the input size and / or pixel format of the face detector 212.

[0027] The face detector 212 may be hardware and / or software for recognizing one or more faces in a frame. In one or more embodiments, the face detector 212 may be included in a neural processing unit. Face detection may be synchronous or asynchronous with the display of the frame (e.g., via output device interface 208). The face detector 212 may receive frames as input, which may contain frames of video. The face detector 212 may output one or more bounding boxes that delineate one or more faces presented in the frame. The output may be generated based on one or more machine learning models trained to recognize one or more faces and generate corresponding bounding boxes. The machine learning models may be trained using a dataset containing a set of frames labeled with faces included in the frames. The machine learning models may also, or alternatively, be trained to recognize one or more pixels representing facial features.

[0028] The face data extractor 214 may be hardware and / or software for extracting face data. In one or more embodiments, the face detector 212 may be included in a neural processing unit. The face data extractor 214 may receive frames (or subsets thereof) input to the face detector 212 as input, such as a portion of a frame contained within a bounding box. The face data extractor 214 may output, for example, vectors, embeddings, images, or any other face-related data.

[0029] One challenge in facial recognition is that a face may not be facing the camera. A turned face may only allow for the extraction of partial facial data. The facial data extractor 214 may include algorithms to refine the facial data to obtain more facial data. Refinement may involve referencing neighboring frames to select frames that display the target face more completely (e.g., when the face is turned towards the camera). To select such frames, the facial data extractor 214 may select frames containing the target face based on matching partial facial data. The facial data extractor 214 may compare the partial facial data with facial data from faces in neighboring frames and determine whether faces in neighboring frames match the partial facial data. For example, the target face and faces in neighboring frames may have matching facial regions, such as the base of the chin, the top of the nose, the outer corners of the eyes, and various points around the eyes and mouth.

[0030] Storage device 202 may be a read-write memory device. Storage device 202 may be a non-volatile memory unit that stores instructions and data (e.g., static and dynamic instructions and data) even when computing system 200 is powered off. Storage device 202 may include one or more databases, storage volumes, caches, and other data structures. In one or more embodiments, a mass storage device (e.g., a disk or optical disk and its corresponding disk drive) may be used as storage device 202. In one or more embodiments, a removable storage device (e.g., a floppy disk, flash drive, and its corresponding disk drive) may be used as storage device 202.

[0031] Similar to storage device 202, system memory 204 may be a read-write memory device. However, unlike storage device 202, system memory 204 may be a volatile read-write memory, such as random access memory. System memory 204 may store any instructions and data that one or more processing units 220 may need to perform operations during runtime. System memory 204 may also include one or more buffers, caches, and other forms of data structures. In one or more embodiments, the processes of this disclosure are stored in system memory 204 and / or storage device 202. From these various memory units, one or more processing units 220 retrieve instructions to be executed and data to be processed in order to perform the processes of one or more embodiments.

[0032] The embodiments within the scope of this disclosure may be implemented in whole or in part using tangible computer-readable storage media (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions. The tangible computer-readable storage media may also be non-transitory.

[0033] Computer-readable storage media can be any storage medium that can be read, written, or otherwise accessed by a general-purpose or special-purpose computing device, which includes any processing electronic device and / or processing circuitry capable of executing instructions. For example, but not limited to, computer-readable media can include any volatile semiconductor memory (e.g., system memory 204), such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. Computer-readable media can also include any non-volatile semiconductor memory (e.g., storage device 202), such as ROM, PROM, EPROM, EEPROM, NVRAM, flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and millipede memory.

[0034] Furthermore, the computer-readable storage medium may comprise any non-semiconductor memory, such as optical disc storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In one or more embodiments, the tangible computer-readable storage medium may be directly coupled to a computing device, while in other embodiments, the tangible computer-readable storage medium may be indirectly coupled to a computing device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.

[0035] Instructions can be directly executable or used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or as instructions in a high-level language form that can be compiled to produce executable or non-executable machine code. Furthermore, instructions can be implemented as data or can contain data. Computer executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As will be recognized by those skilled in the art, details including, but not limited to, the number, structure, order, and organization of instructions can vary significantly without altering the underlying logic, functionality, processing, and output.

[0036] While the above discussion primarily concerns microprocessors or multi-core processors that execute software, one or more implementations are executed by one or more integrated circuits (e.g., ASICs or FPGAs). In one or more implementations, such integrated circuits execute instructions stored on the circuit itself.

[0037] Figure 3This illustration shows a machine learning model 300 for face recognition according to one or more exemplary embodiments. Both the face detector 212 and the face data extractor 214 can be machine learning models. The machine learning model may contain a series of neural network layers. For example, machine learning model 300 includes a neural network architecture for an instance face detector. Machine learning model 300 may include a backbone 302, a feature pyramid network 312, and a prediction network 320. Features can be extracted using the backbone network 302, which may contain multiple layers (e.g., layers 304, 306, 308, 310), which may represent a set of convolutions. The extracted features at different scales can then be fused in the feature pyramid network 312, which may contain multiple layers (e.g., layers 314, 316, 318). Finally, the prediction network 320, which may also have multiple layers (e.g., layers 322, 324, 326), can be used to predict bounding box coordinates, class probabilities, etc., based on the fused features from the feature pyramid network 312.

[0038] Figure 4 This diagram illustrates a multichannel manager 400 according to one or more exemplary embodiments. Manager 400 can be configured to accelerate inference of one or more machine learning models (e.g., machine learning model 300). For example, manager 400 can coordinate the execution of processing for machine learning models such that layers of the machine learning models can be scheduled to utilize processor time, while other layers are paused. The partitioning of processor time among machine learning models allows multiple models to be processed simultaneously, thereby dividing the work of each model into discrete units (e.g., per layer or channel). The partitioning of processor time among machine learning models allows manager 400 to prioritize certain machine learning tasks. For example, when media device 102 listens to audio (containing commands) from user 118, the media device can utilize a machine learning model for speech analysis to extract command 114 from the audio captured by the microphone of media device 102. To prevent latency in audio processing and enhance the user experience, manager 400 can prioritize one or more layers of the machine learning model processing user 118's speech, such that command 114 can be determined while other machine learning processes can be temporarily paused.

[0039] Manager 400 may include channel queue 401, layer scheduler 404, and layer engine 408. When processing of the machine learning model begins, a work list 403 may be created in memory (e.g., system memory 204). Work list 403 may contain a sequence of layer execution tasks, where layers may be part of channels belonging to the machine learning model. For example, work list 403 may contain layers 304, 306, 308, and 310 of the backbone channel 302 of machine learning model 300. Layers in work list 403 may be arranged according to the priority of their corresponding machine learning models. Layers in work list 403 may be grouped according to their corresponding machine learning models.

[0040] Layer scheduler 404 may add work list 403 to its queue based on priority levels (e.g., per machine learning model, per channel, or per layer), allowing higher-priority layers to be processed first. Layer engine 408 may then process layers according to the order provided by layer scheduler 404. Layer engine 408 may also write the processing result 412 as input to the next layer, or as output if the processed layer is an output layer. Channel queue 401 tracks layer processing and notifies the corresponding machine learning model (or a subset thereof) when a layer of a machine learning model has been completed.

[0041] For example, suppose machine learning model 414 is configured to process voice commands from a user. Machine learning model 414 may have a pre-determined priority level 402. Machine learning model 300 may be initiated to recognize one or more faces in a frame and may have a priority level 410. Priority level 402 may be a higher priority than priority level 410, thereby improving the responsiveness of the media device and the user experience.

[0042] In this example, machine learning model 300 may be placed in channel queue 401 (e.g., in system memory 204) when it begins execution. Layers 406 and 407 of machine learning model 300 may be added to work list 403 for storage, and indications (e.g., pointers) of layers 406 and 407 may be added to layer scheduler 404 to queue layers 406 and 407 for processing. Layer 406 may be the first in the queue and, along with its associated inputs and weights 409, is passed as input to layer engine 408 for processing. The processing result of layer 406 may be passed to layer 407 to modify layer 407 (e.g., adjust associated inputs and / or weights). Other layers of other machine learning models may be queued among layers 406 and 407. Once layers 406 and 407 have finished processing, machine learning model 300 may be notified and / or removed from channel queue 401.

[0043] Continuing from the previous example, machine learning model 414 can then be added to the queue (as per the previous example). Figure 4 (Indicated by the dotted line in the diagram). Because machine learning model 414 has a higher priority than machine learning model 300, it can be placed at the top of channel queue 401, which allows layers 302 and 312 of machine learning model 414 to be placed at the front of the work list. Then, layer scheduler 404 can interrupt the execution of machine learning model 300 to execute the higher-priority task from machine learning model 414.

[0044] Figure 5This describes frame 500 of a video according to one or more exemplary embodiments. To present the video, it may undergo a display process. This display process may encode, decode, resize, or otherwise manipulate the frame before it is displayed. A facial recognition process may be synchronized with the display process. A user may send a start signal (e.g., command 114) to a media device (e.g., media device 102) such as a video (e.g., file or stream) player or set-top box. This signal may be sent, for example, by pressing a button on a remote control associated with the media device, using a keypad on the remote control for conversation, or via far-field voice on the media device. The media device may pause the video at frame 500 and perform the facial recognition process on frame 500. The facial recognition process may identify faces 502, 504 displayed in scene 506 of frame 500. In one or more embodiments, the facial recognition process may identify at least a portion of a face and utilize neighboring frames (e.g., the 10th, 20th, and 30th) to identify the displayed faces 502, 504 with higher confidence. One or more face blocks for one or more displays can be generated from a portion of one or more frames corresponding to one or more bounding boxes of one or more displays. The displayed faces 502, 504 can be presented to the user by overlaying face blocks 508, 510 of the displayed faces 502, 504 and their corresponding identity data 512, 514 onto the currently displayed frame 500. The user can select a specific face to obtain more identity data about the individual associated with that face.

[0045] In one or more implementations, the facial recognition process can be asynchronous with the display process. When a user initiates facial recognition, the video can continue without pausing. The media device can run an edge-based facial recognition process in the background. The displayed faces 502, 504 can be extracted and displayed on top of frame 500.

[0046] A synchronous or asynchronous face recognition process can consist of two steps. First, the face recognition process can be used to identify faces displayed in a frame. If the detected faces in the frame can be identified with high confidence (e.g., above a confidence level threshold), then the recognition process can end, and a list of displayed faces can be returned. Otherwise, a face recognition refinement process is used to further identify unknown faces by looking at temporally adjacent frames to gather more data about unknown faces.

[0047] Figure 6This diagram illustrates a process 600 for facial recognition in video according to one or more exemplary embodiments. The video may be input to a media device (e.g., media device 102) in an encoded format. Frames of the video may populate an encoded image buffer 602. A video decoder 604 may decode frames from the encoded image buffer 602 to a decoded image buffer 606. In one or more embodiments, the decoded image buffer 606 may be allocated one or more additional frame buffers such that the looping of the decoded frame buffers can be delayed, and thus one or more previous frames (e.g., frames that have been passed to the post-processing engine 608) are held in the decoded image buffer 606. The delay allows processing unit 220 to access frames in the decoded image buffer 606 preceding the current frame without having to re-decode the frames (e.g., accessing neighboring frames preceding the current frame, as shown at box 714). Figure 7 (As described).

[0048] Decoded video frames in the decoded image buffer 606 can be sent to the display buffer 609 and / or adjusted by the post-processing engine 608 according to the input size and / or pixel format of the face detector 212 (step 1). Adjusting the decoded frames may include, but is not limited to, scaling, cropping, color space conversion, HDR to SDR tone mapping, etc. The result of the adjustment is written to the scaled image buffer 610. Then, the face detector 212 can perform face detection on the scaled frames from the scaled image buffer 610 and output a list of detected faces with bounding boxes 614 (step 2). The processing unit 220 can extract face patches 622 of the user-selected face from the scaled frames based on the bounding boxes 614 (step 3). Face patches 622 (e.g., a subset of frames containing the displayed face) can be further adjusted by the graphics engine 216 to match the input size of the face data extractor 214 (step 4). The scaled face patches 626 can be sent to the face data extractor 214 to obtain face data 628 of the user-selected displayed face (step 5). The processing unit 220 can compare the extracted facial data 628 with the facial data in the facial database 620, and select facial data in the facial database 620 that matches the facial data 628, so that the confidence level of the match is higher than the confidence level threshold (step 6).

[0049] In an example where the extracted facial data 628 is matched with selected facial data at a confidence level higher than a threshold (e.g., 75%), the extracted facial data associated with the selected facial data and any associated identity data can be sent to the graphics engine 216 for rendering 632 (step 7). Rendering 632 may include the current frame from the display buffer 609, facial patch 622, and / or identity data 630. For example, identity data 512, 514 may be overlaid on the current frame 500 along with their corresponding facial patches 508, 510, as shown. Figure 5 As shown in the image.

[0050] In instances where the extracted facial data 628 does not match the selected facial data with a confidence level higher than a confidence level threshold, the extracted facial data 628 may not contain sufficient data (e.g., because the face may be angled) to produce a confidence match with any facial data in the facial database 620. Therefore, a bidirectional facial recognition refinement process can be performed to improve the extracted facial data 628 by examining temporally adjacent frames (e.g., frames within a time period before or after the current frame). To improve the extracted facial data 628, processing unit 220 may select frames adjacent to the current frame. Adjacent frames may be in a direction after or before the current frame and contain faces with facial data corresponding to the extracted facial data 628 (e.g., the extracted facial data 628 is a subset of the facial data). Steps 1 to 6 described above can be repeated on the extracted facial data 628 and / or facial data from faces in adjacent frames. In cases where facial data from neighboring frames still does not return a match from the facial database 620, the processing unit 220 may repeat the process of selecting neighboring frames until a frame containing a face with facial data corresponding to the extracted facial data 628 is found. Once a neighboring frame is selected, steps 1 to 6 may be repeated for the neighboring frames until matching facial data with a confidence level higher than the confidence level threshold is found in the facial database 620.

[0051] In one or more implementations, a candidate face list 618 may be generated to create a queue of facial data whose identity will be determined. The candidate face list 618 can be used to determine the identity data of multiple faces displayed in the current frame. For example, a user may want to determine the identity data of two faces in the current frame; in this case, the facial data of both faces may be added to the candidate face list 618 to be identified. If identity data for a facial data is found in the candidate face list 618, then the facial data may be removed from the candidate face list 618; otherwise, the facial data may be skipped (e.g., to generate an error indicating that the face cannot be recognized). The media device may determine the identity of the remaining facial data in the candidate face list 618 (e.g., via steps 1 to 7) until the candidate face list 618 is empty or contains only skipped facial data.

[0052] Figure 7 This document describes a flowchart of a process 700 for facial recognition in a video, according to one or more exemplary embodiments. For illustrative purposes, this document primarily refers to... Figures 1 to 6 Process 700 is described. One or more blocks (or operations) of process 700 may be performed by one or more components of a suitable device (e.g., system 200). Furthermore, for illustrative purposes, the blocks of process 700 are described herein as occurring sequentially or linearly. However, multiple blocks of process 700 may occur in parallel. Additionally, the blocks of process 700 do not need to be performed in the order shown, and / or one or more blocks of process 700 need not be performed and / or may be replaced by other operations.

[0053] In process 700, at box 702, a media device (e.g., media device 102) can access video. The video can be a file, a stream, or any other format. The video can contain a set of frames. One or more frames in the set can contain one or more objects (e.g., faces). The video can be accessed by downloading, streaming, receiving, retrieving, querying, or any other method of acquiring the video.

[0054] At frame 704, the media device may (e.g., from a user) receive a command to identify a target in the first frame (“current frame”) of the video. The target can be any object in the frame, such as a human, animal, object, and / or a portion thereof. For example, the target could be a face (e.g., as shown in the image). Figure 6 (As discussed). The command can be an instruction for identifying one or more targets from one or more targets that can be displayed in a frame. If one or more frames display one or more targets, then one or more targets displayed in one or more frames can be identified by a target detector (e.g., face detector 212).

[0055] At box 706, first target data can be generated based on targets displayed in the first frame. Generating the target data may involve a target detector of the media device (e.g., face data extractor 214) performing target detection on the current frame and outputting one or more detected targets with their corresponding bounding boxes (e.g., bounding box 614) and target blocks (e.g., face blocks 622). Target blocks may contain images of the targets displayed in the frame as defined by the bounding boxes. In one or more embodiments, the target blocks may be adjusted (e.g., scaled) by a graphics engine (e.g., graphics engine 216) to match the input size of the target data extractor. The target blocks (e.g., scaled face blocks 626) may be sent to the target data extractor to convert the target blocks (e.g., images of faces) into target data corresponding to features of the targets displayed in the target blocks. Converting target chunks into target data can include identifying metrics associated with target features, such as the distance between the eyes, the distance from the forehead to the chin, the distance between the nose and mouth, the depth of the eye sockets, the shape of the cheekbones, the contours of the lips and ears, and any other metrics related to the target shown in the target chunks. The identified metrics can be converted into target data, which may include a set of numbers, data points, target embeddings, or other data representing the target.

[0056] In one or more embodiments, the video may undergo preprocessing before target identification. Preprocessing may include filling an encoded image buffer (e.g., encoded image buffer 602) with encoded frames, decoding the frames using a video decoder (e.g., video decoder 604), and preparing frames for display by filling a display buffer (e.g., display buffer 609). The preprocessed frames may be adjusted based on the input size and format of a machine learning model (e.g., face detector 212). Adjusting the frames may include scaling, cropping, color space conversion, HDR-to-SDR tone mapping, etc.

[0057] At box 708, the media device may generate a first confidence level of similarity between first target data and second target data stored in a remote database. The database may be a repository of target data. The target data stored in the database may be associated with the target's identity data. For example, the target data may contain representations of an actor's facial features (e.g., eyes, nose, mouth), and the target data may be associated with the actor's identity data, which may contain the actor's profile, including data such as name, age, and other programs in which the actor has starred.

[0058] A confidence level for the similarity between target data and target data from a database can be generated using methods that include, but are not limited to, machine learning models (e.g., computer vision models). The machine learning model can be trained on training data containing target data stored in the database and known identity data associated with each target data. The trained machine learning model can receive target data as input and produce multiple or a set of values ​​as output representing the confidence level (e.g., percentage) for identifying the input target data from the training data. The confidence level represents the accuracy of the comparison and includes how closely the input target data matches the target data in the training data, where the training data is based on target data from a database. Determining the closeness of the match between the input target data and the target data in the training data can be a function of variables that include the number of matching data points between several sets of target data, the distance between data points in several sets of target data, and any other metrics associated with several sets of target data.

[0059] At box 710, the determined confidence level can be compared to a confidence threshold level. The determined confidence level can be a percentage and can be required to meet or exceed a predetermined percentage (e.g., a threshold level) to be considered a successful match between the target data of the displayed target (e.g., first target data) and the target data selected from the database. If the confidence level is at or above the threshold level, then process 700 can proceed to box 712.

[0060] At box 712, the media device may output identity data associated with the selected target data. The identity data may include data such as name, occupation, age, or any other data related to the target. The identity data may be contained in a database and / or retrieved from a remote source. In one or more embodiments, any data generated in process 700, such as target chunks, bounding boxes, target data, etc., may also be output. Outputting the identity data may include sending the identity data to a graphics engine (e.g., graphics engine 216) to overlay the identity data onto the current frame to display the identity data to the user. For example, when a user pauses the video and issues a command to identify the target displayed in the paused frame, the media device may execute process 700 so that the user sees the frame and the identity data (e.g., name) of the target next to the target on the frame.

[0061] In one or more embodiments, the identity data of the displayed target and the target data may be stored in a remote or local storage device. Storing the identity data can be used to cache data to improve the efficiency of subsequent target identification performed by the media device or another media device. Thus, the same or different media devices can query the corresponding identity data from the remote or local storage device containing the video data and target data. For example, if a first user guides a first media device to identify a character in a program at a specific time, the character's identity data, the program, and the time can be cached in the remote storage device. When a second user guides a second media device to identify the same character in the same program at a similar time, the second media device can query the remote storage device for the target data of the character in the program at the time of the program to obtain the character's identity data.

[0062] In one or more embodiments, the media device may generate a candidate target list (e.g., candidate face list 618) for identifying multiple targets in the current frame. One or more sets of target data corresponding to one or more targets to be identified in the current frame may be added to the candidate target list. If identity data corresponding to the target data is found or cannot be found after analyzing one or more (e.g., up to a threshold number) neighboring frames, the target data may be removed from the candidate target list. Process 700 may be repeated for one or more sets of target data in the candidate target list until the candidate target list is empty or until the candidate target list contains one or more sets of target data after repeated execution of process 700. In one or more embodiments, the media device may output an error (e.g., a message declaring the target "unknown") indicating that a set of target data could not be identified, although said set of target data remains in the candidate target list.

[0063] Returning to box 710, if the determined confidence level is lower than the confidence level threshold, the process can proceed to box 714, where the target data can be refined by repeating process 700 on frames different from the current frame in an attempt to generate a higher confidence level.

[0064] In one or more embodiments of block 714, the different frames may be one or more neighboring frames used to refine the target data of the displayed target. Neighboring frames are included in frames preceding or following the current frame and may be included in a decoded image buffer (e.g., decoded image buffer 606). Accessing one or more neighboring frames can reveal additional target data of the displayed target, which can improve the confidence level between the target data of the displayed target and target data in a database. For example, a target in the current frame may not be entirely in the frame, and therefore only a portion of the target can be identified in the frame. In the case where only a portion of the target is in the current frame, only a portion of the target data can be extracted, which can reduce the confidence level of the accuracy of the comparison. Neighboring frames may include targets in different locations, allowing more targets to be captured in the frame, thereby revealing more targets and allowing the extraction of more target data.

[0065] To select neighboring frames, the media device can identify one or more targets in neighboring frames, similar to box 702. The media device can identify a corresponding target in a neighboring frame that may correspond to the target displayed in the current frame. If the target data of a target contains target data other than that of the target displayed in the current frame, then the target may be the corresponding target. That is, when the media device determines that there is at least a partial match between the target data of the displayed target and the target data of the corresponding target, the target data of the target in a neighboring frame can be refined by adding the target data of the corresponding target to the target data of the displayed target. If a target in a neighboring frame does not correspond to the displayed target, then a new neighboring frame can be selected until a target corresponding to the displayed target is identified. Box 704 can be performed on the refined target data from neighboring frames (e.g., target data from the corresponding target). For example, the selected target data from a target database can be compared with the refined target data, a confidence level of accuracy between the selected target data and the refined target data can be determined, and if the confidence level is at or above a confidence level threshold, then identity data corresponding to the selected target data can be output. If the refined target data is still insufficient to generate a confidence level higher than the confidence level threshold, the media device may continue to refine the target data of the displayed target (e.g., up to the threshold number of frames) or determine that the target has not yet been identified (e.g., when the threshold number of frames has been analyzed).

[0066] In one or more embodiments, the media device may analyze a scene (e.g., scene 506) to determine the direction for selecting neighboring frames (e.g., behind or in front). If the scene has changed between the current frame and neighboring frames, then a target from the current frame is unlikely to be in the second frame. To determine whether the scene has changed, the media device may first select neighboring frames in an initial direction from the current frame (e.g., behind or in front). The media device may identify one or more features of the scene from the current frame and the scene from the neighboring frames to determine whether the scene has changed from the current frame to the neighboring frame. For example, scene features may include color, brightness, characters, and other visual characteristics. If the scene has changed, then the media device may terminate the search process in the initial direction (e.g., in front) and search in another direction (e.g., behind). If there is no scene change in the neighboring frames, then the media device may continue to select neighboring frames for refining the target data from the current frame.

[0067] Those skilled in the art will understand that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability between hardware and software, the various illustrative blocks, modules, elements, components, methods, and algorithms have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the overall system. Skilled technicians can implement the described functionality in different ways for each specific application. Various components and blocks can be arranged in different ways (e.g., in different orders or in different ways), all of which remain within the scope of this art.

[0068] It should be understood that any specific order or hierarchy of blocks in the disclosed process is a specification of the instance method. Based on design preferences, it should be understood that the specific order or hierarchy of blocks in the process may be rearranged, or all specified blocks may be executed. Any block may be executed concurrently. In one or more embodiments, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0069] As used in this specification and any claim of this application, the terms "base station," "receiver," "computer," "server," "processor," and "memory" refer to electronic or other technical devices. These terms exclude persons or groups of people. For the purposes of this specification, the term "display" means displaying on an electronic device.

[0070] As used herein, the phrase "at least one" preceding a series of items, where the terms "and" or "or" are used to separate any one of the items, modifies the list as a whole, not each member of the list (i.e., each item). The phrase "at least one" does not require selection of at least one of each of the listed items; rather, the phrase allows for the meaning of at least one of any of the items, and / or at least one of any combination of items, and / or at least one of each of the items. For example, the phrases "at least one of A, B, and C" or "at least one of A, B, or C" respectively refer to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0071] The predicates “configured to,” “operable to,” and “programmed to” do not imply any specific tangible or intangible modification to the subject, but are intended to be used interchangeably. In one or more embodiments, a processor configured to monitor and control operation or components may also mean a processor programmed to monitor and control operation or operable to monitor and control operation. Similarly, a processor configured to execute code can be interpreted as a processor programmed to execute code or operable to execute code.

[0072] Phrases such as aspect, said aspect, on the other hand, some aspects, one or more aspects, implementation, said implementation, another implementation, some implementations, one or more implementations, embodiment, said embodiment, another embodiment, some embodiments, one or more embodiments, configuration, said configuration, another configuration, some configurations, one or more configurations, the technology, the disclosure, the present disclosure, and other variations and similarities are for convenience only and do not imply that the disclosure associated with such phrases is essential to the technology, nor does it imply that such disclosure applies to all configurations of the technology. The disclosure associated with such phrases may apply to all configurations or one or more configurations. The disclosure associated with such phrases may provide one or more instances. For example, the phrase "aspect" or "some aspects" may refer to one or more aspects, and vice versa, and this also applies to the other phrases mentioned above.

[0073] The term “exemplary” is used herein to mean “serving as an example, illustration, or description.” Any embodiment described herein as “exemplary” or “example” is not necessarily to be construed as being more preferred or advantageous than other embodiments. Furthermore, where the terms “comprising,” “having,” or similar terms are used in the specification or claims, such terms are intended to be interpreted in a manner similar to how the phrase “comprising” is interpreted as a transition word in the claims.

[0074] All structural and functional equivalents of elements throughout the various aspects described herein that are known or subsequently known to a person skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be made public, whether or not such disclosure is expressly stated in the claims. No claim element should be interpreted in accordance with 35 U.S.SC §112(f) unless the element is expressly stated using the phrase “component for…” or, in the case of a method claim, using the phrase “step for…”.

[0075] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects presented herein, but should be consistent with the full scope of the language of the claims, wherein, unless expressly stated otherwise, an element referred to in the singular does not mean “one and only one,” but rather “one or more.” Unless expressly stated otherwise, the term “some” refers to one or more. A masculine (e.g., his) pronoun includes a feminine (e.g., her) pronoun, and vice versa. Titles and subtitles (if any) are provided for convenience only and do not limit this disclosure.

Claims

1. A facial recognition method for a media playback device, comprising: Accessing a video consisting of a set of frames; Receive a command to identify a target displayed in the first frame of the set of frames; First target data is generated based on the target displayed in the first frame; A first confidence level is generated to determine the similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; In response to the first confidence level being at or above the confidence level threshold, the identity data of the second target data is output; The output step includes: The identity data that displays the second target data; and Store the second target data in a local database; and In response to the first confidence level being lower than the confidence level threshold: A third target data is generated based on the target displayed in a second frame of the set of frames, wherein the second frame is a first neighboring frame in a first direction from the first frame; A second confidence level is generated to determine the similarity between the third target data and the second target data stored in the remote database; In response to the second confidence level being at or above the confidence level threshold, the identity data of the second target data is output; and In response to the second confidence level being lower than the confidence level threshold, it is determined that the target has not been identified.

2. The method of claim 1, further comprising responding to receiving a subsequent command for identifying the target in a third frame of the set of frames: A fourth target data is generated based on the target displayed in the third frame; A third confidence level is generated to determine the similarity between the fourth target data and the second target data stored in the local database; and In response to the third confidence level being at or above the confidence level threshold, the identity data of the second target data is output.

3. The method according to claim 1, further comprising: Generate a candidate target list including the target data and the second target displayed in the first frame; In response to the first confidence level being at or above the confidence level threshold, the target data is removed from the candidate target list; and In response to the first confidence level being lower than the confidence level threshold, the target data in the candidate target list is skipped; and In response to the candidate target list including remaining target data, the access, receiving, generating, and output steps are repeated for the remaining target data in the candidate target list.

4. The method of claim 1, wherein in response to determining that the scene of the first neighboring frame does not match the scene of the frame, the second frame is a third neighboring frame in a second direction from the first frame.

5. A facial recognition system for a media playback device, comprising: A processor configured to perform operations including: Accessing a video consisting of a set of frames; Receive a command to identify a target displayed in the first frame of the set of frames; First target data is generated based on the target displayed in the first frame; A first confidence level is generated to determine the similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; In response to the first confidence level being at or above the confidence level threshold, the identity data of the second target data is output; The output operations mentioned above include: The identity data that displays the second target data; and Store the second target data in a local database; and In response to the first confidence level being lower than the confidence level threshold: A third target data is generated based on the target displayed in a second frame of the set of frames, wherein the second frame is a first neighboring frame in a first direction from the first frame; A second confidence level is generated to determine the similarity between the third target data and the second target data stored in the remote database; In response to the second confidence level being at or above the confidence level threshold, the identity data of the second target data is output; and In response to the second confidence level being lower than the confidence level threshold, it is determined that the target has not been identified.

6. The system of claim 5, wherein the operation further comprises responding to receiving a subsequent command for identifying the target in a third frame of the set of frames: A fourth target data is generated based on the target displayed in the third frame; A third confidence level is generated to determine the similarity between the fourth target data and the second target data stored in the local database; and In response to the third confidence level being at or above the confidence level threshold, the identity data of the second target data is output.

7. The system of claim 5, wherein the operation further comprises: Generate a candidate target list including the target data and the second target displayed in the first frame; In response to the first confidence level being at or above the confidence level threshold, the target data is removed from the candidate target list; and In response to the first confidence level being lower than the confidence level threshold, the target data in the candidate target list is skipped; and In response to the candidate target list including remaining target data, the access, receiving, generating, and output steps are repeated for the remaining target data in the candidate target list.

8. The system of claim 5, wherein in response to determining that the scene of the first neighboring frame does not match the scene of the frame, the second frame is a third neighboring frame in a second direction from the first frame.

9. A facial recognition system for a media playback device, comprising: A processor configured to perform operations including: Accessing a video consisting of a set of frames; Receive a command to identify a target displayed in the first frame of the set of frames; First target data is generated based on the target displayed in the first frame; A first confidence level is generated to determine the similarity between the first target data and second target data stored in a remote database, wherein the second target data includes identity data; In response to the first confidence level being at or above the confidence level threshold: Display the identity data of the second target data; The second target data is stored in a local database; In response to the first confidence level being lower than the confidence level threshold: A third target data is generated based on the target displayed in a second frame of the set of frames, wherein the second frame is a first neighboring frame in a first direction from the first frame; A second confidence level is generated to determine the similarity between the third target data and the second target data stored in the remote database; In response to the second confidence level being at or above the confidence level threshold, the identity data of the second target data is output; and In response to the second confidence level being lower than the confidence level threshold, it is determined that the target has not been identified.

10. The system of claim 9, wherein the operation further comprises responding to receiving a subsequent command for identifying the target in a third frame of the set of frames: A fourth target data is generated based on the target displayed in the third frame; A third confidence level is generated to determine the similarity between the fourth target data and the second target data stored in the local database; and In response to the third confidence level being at or above the confidence level threshold, the identity data of the second target data is output.

11. The system of claim 9, wherein the operation further comprises: Generate a candidate target list including the target data and the second target displayed in the first frame; In response to the first confidence level being at or above the confidence level threshold, the target data is removed from the candidate target list; and In response to the first confidence level being lower than the confidence level threshold, the target data in the candidate target list is skipped; and In response to the candidate target list including remaining target data, the access, receiving, generating, and output steps are repeated for the remaining target data in the candidate target list.

12. The system of claim 9, wherein in response to determining that the scene of the first neighboring frame does not match the scene of the frame, the second frame is a third neighboring frame in the set of frames that precedes or follows the first frame.

Citation Information

Patent Citations

  • Face recognition in video content

    CN102542249A