Dynamic visual sensor for visual audio processing

By fusing EDS with an RGB camera and an audio sensor, and using neural networks for face tracking, the challenges of face tracking and speech recognition in noisy environments are solved, achieving efficient and low-latency face tracking and recognition.

CN115485748BActive Publication Date: 2026-03-24SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-21
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In noisy environments, existing technologies struggle to effectively track mouth positions and rapid facial movements during face tracking and speech recognition, requiring high-speed cameras that result in high bandwidth and high power consumption.

Method used

By fusing an event-driven sensor (EDS) with an RGB camera and an audio sensor, changes in facial lighting are detected by the EDS and combined with RGB and audio data, and a neural network is used for facial tracking and recognition.

Benefits of technology

It improves the robustness and accuracy of facial tracking, reduces latency and energy consumption, and adapts to fast movement and complex lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115485748B_ABST
    Figure CN115485748B_ABST
Patent Text Reader

Abstract

To track certain difficult facial features such as mouth corners and teeth during speech, the camera sensor system (212 / 306 / 308 / 318 / 320) generates RGB / IR images and the system also uses light intensity change signals from the event-driven sensor (EDS) (212 / 306 / 318), as well as signals from the microphone (328) for speech analysis. In this way, the camera sensor system achieves improved performance tracking at lower bandwidth and power consumption (equivalent to using very high speed cameras).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates generally to technologically innovative unconventional solutions that must stem from computer technology and produce concrete technical improvements. BACKGROUND

[0002] Certain parts of the face, such as the dark regions inside the mouth, teeth, and rapid movements of the facial structure during speech, pose a challenge to accurate tracking when performing face tracking in speech processes to identify a speaker in a noisy environment, or to detect fake videos, or for speech recognition to resolve ambiguities, or for animation or other purposes. SUMMARY

[0003] The technical challenge posed by the above is that, for better operation, a high-speed camera can be required to reduce latency and improve tracking performance, which requires increasing the camera data frame rate, but such higher frame rates require higher bandwidth and processing, thus consuming relatively large power and generating heat.

[0004] To address the challenges mentioned herein, a camera sensor system is provided that not only includes a sensor unit with light intensity photodiodes under color and an infrared filter if needed to capture RGB and IR images, but also an event-driven sensor (EDS) sensing unit that detects motion by EDS principles. EDS uses changes in light intensity sensed by one or more camera pixels as an indication of motion. EDS has high dynamic range (HDR), no motion blur, and low latency compared to RGB cameras. When EDS information is fused with RGB camera information and audio information, tracking becomes more robust. EDS information can be more relied on in conditions of rapid motion (e.g., of the mouth) or HDR, where camera images are more relied on in conditions of slow motion and fine details (color, texture). This fusion can also be applied to face tracking, eye tracking, and emotion recognition.

[0005] Current principles use raw event data from EDS fused with RGB camera and audio data, then input to a classifier. The classifier is trained using a training set of all three inputs from audio / camera / event data, in some implementations using a recurrent neural network with convolutional layers.

[0006] Accordingly, an assembly includes at least one camera unit configured to generate a red-green-blue (RGB) image of a face. The assembly also includes at least one event-driven sensor (EDS) configured to output a signal representative of a change in illumination intensity of the face. At least one microphone is configured to output a signal representative of speech. Further, the assembly includes at least one processor configured with executable instructions to receive the signals from the camera unit, the EDS, and the microphone. The instructions are executable to execute at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, at least one of an emotion prediction, a tracking of at least one portion of the face.

[0007] In some examples, the camera unit is configured to generate an infrared (IR) image.

[0008] In example implementations, the camera unit, the processor, and the EDS can be disposed on a single chip.

[0009] In non-limiting implementations, the tracked facial portion can be one or both eyes, in particular one or both pupils, and can be limited to the pupils or can include other facial features. In other implementations, the portion includes a mouth corner and can be limited to the mouth corner and / or the interior of the mouth including the teeth, or can also include other facial features.

[0010] In another aspect, a system includes at least one camera unit configured to generate a red-green-blue (RGB) image and / or an infrared (IR) image of a person. The system also includes at least one microphone and at least one event-driven sensor (EDS) configured to output a signal representative of the person. The system further includes at least one processor programmed with instructions to process the output of the microphone using a short-time Fourier transform (STFT) and to process the output of the STFT using at least one audio processing convolutional neural network (CNN). The instructions are executable to process at least features in the image from the camera unit using at least one vision processing CNN. Further, the instructions are executable to process the representative of the output signal from the EDS using at least one event processing CNN. The instructions in the system can be executable to fuse the outputs of the CNNs in a fully connected neural network layer to generate one or more of an emotion prediction for the person, a tracking of at least one portion of the face of the person, at least one virtual reality (VR) image of the person, and an identity of the person.

[0011] In one example of the latter aspect, the processor can be configured with instructions to process the outputs of the CNNs using a recurrent neural network (RNN) and to process the output of the RNN using a fully connected neural network layer to generate a tracking of the mouth of the person.

[0012] In another aspect, a method includes receiving signals from at least one camera unit, receiving signals from at least one event-driven sensor (EDS), and receiving signals from at least one microphone. The method includes executing at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, at least one of: an emotion prediction, a tracking of at least a portion of a face, an identity of a person, a generation of a virtual reality (VR) image of the person.

[0013] The details of the application, both as to its structure and operation, can be best understood with reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a block diagram of an example system including examples in accordance with the principles of the application;

[0015] Figure 2 A simplified sensor data flow is shown;

[0016] Figure 3 A sensor related to a face of a tracked person is shown;

[0017] Figure 4 An example system is shown in block diagram format;

[0018] Figure 5 Data flow for RGB input, audio input, and EDS input for emotion recognition or speaker recognition is shown;

[0019] Figure 6 An alternative classifier architecture for speech recognition is shown;

[0020] Figure 7 Face tracking information from a camera on a head-mounted display (HMD) or head-mounted camera (HMC) is shown;

[0021] Figure 8 Example tracking logic is shown in example flowchart format;

[0022] Figure 9 Example training logic is shown in flowchart format;

[0023] Figure 10 An example HMC is shown; and

[0024] Figure 11 Facial features imaged by Figure 10 an HMC are shown. DETAILED DESCRIPTION

[0025] This disclosure generally relates to a computer ecosystem, including various aspects of consumer electronics (CE) device networks and standalone computer simulation systems, such as, but not limited to, computer simulation networks such as computer gaming networks. The systems described herein may include server and client components connected via a network, enabling data exchange between the client and server components. The client components may include one or more computing devices, including those such as Sony… Game consoles, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (such as smart TVs and internet-enabled TVs), portable computers (such as laptops and tablets), and other mobile devices (including smartphones and additional examples discussed below) made by Microsoft, Nintendo, or other manufacturers are permitted. These client devices can operate in a variety of operating environments. For example, some client computers may use operating systems such as Linux, Microsoft's operating system, or Unix, or an operating system made by Apple or Google. These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted on internet servers discussed below. Furthermore, the operating environment according to the principles of the present invention can be used to execute one or more computer game programs.

[0026] The server and / or gateway may include one or more processors that execute instructions to configure the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server may connect via a local intranet or virtual private network. The server or controller may be a game console (such as Sony). Instantiation of personal computers, etc.

[0027] Information can be exchanged between clients and servers over a network. For this purpose and for security, servers and / or clients may include firewalls, load balancers, temporary storage devices, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form a device that implements methods for providing secure communities (such as online social networking sites) to network members.

[0028] As used herein, an instruction is a computer-implemented step for processing information in a system. Instructions can be implemented in software, firmware, or hardware and include any type of programming step performed by components of the system.

[0029] The processor can be any conventional general purpose single- or multi-chip processors that, in response to computer readable instructions or software can manipulate the transitory signals. Such a processor can include a storage for code, such as in the register, and transitory signals.

[0030] The software modules described herein by flowcharts and user interfaces can include various subroutines, programs, etc. Without limitation, the logic performed by a particular module can be distributed to other software modules and / or combined with a single module and / or made available in a shared library, etc.

[0031] The principles of the application described herein can be implemented as hardware, software, firmware, or combinations thereof; as such, the illustrative components, blocks, modules, circuits, and steps are arranged or implemented by functional means to achieve the stated ends.

[0032] In addition to the functionality described above, logic block, modules, and circuits described below can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA) or other programmable logic device, such as a

[0033] The functions and methods described below, when implemented in software, can be written in an appropriate language such as, but not limited to, Java, C# or C++, and can be stored on or transmitted over a computer-readable storage medium such as a Random Access Memory (RAM), a Read-Only Memory (ROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other optical storage devices such as Digital Versatile Discs (DVDs), a magnetic storage device such as a diskette or magnetic hard drive, or other magnetic storage devices including removable thumb drives, etc. Connections can establish computer-readable media. Such connections can include, for example, wired hard-wired cables including fiber optic and coaxial wires and digital subscriber line (DSL) and twisted pair wires. Such connections can include wireless communication connections, including infrared and radio.

[0034] Components included in one implementation can be used in other implementations in any appropriate combination. For example, any of the various components described herein and / or depicted in the Figures can be combined, interchanged or excluded from other implementations.

[0035] A system "having at least one of A, B, and C" (likewise, "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes systems that have only A, only B, only C, A and B, A and C, B and C, and / or A and B and C, etc.

[0036] Referring now to the drawings in detail Figure 1 An example system 10 is shown that can include one or more of the example devices mentioned above and described further below in accordance with the principles of the application. A first one of the example devices included in system 10 is a consumer electronics (CE) device such as an audio video device (AVD) 12 such as, but not limited to, an Internet-enabled TV with a TV tuner (equivalently, a set-top box that controls a TV). Alternatively, however, AVD 12 can be an appliance or a home item, e.g., a computerized Internet-enabled refrigerator, washer, or dryer. Alternatively, AVD 12 can also be a computerized Internet-enabled ("smart") phone, tablet computer, notebook computer, computerized Internet-enabled wearable device such as, e.g., a computerized Internet-enabled watch, computerized Internet-enabled bracelet, other computerized Internet-enabled device, computerized Internet-enabled music player, computerized Internet-enabled earpiece, computerized Internet-enabled implantable device such as an implantable skin device, etc. Regardless, it is to be understood that AVD 12 is configured to take the principles of the application (e.g., to communicate with other CE devices to take the principles of the application, to execute the logic described herein, and to perform any other functions and / or operations described herein).

[0037] Thus, to implement such principles, AVD 12 can be implemented by Figure 1Some or all of the illustrated components are established. For example, the AVD 12 can include one or more displays 14, which can be implemented by high definition or ultra-high definition "4K" or higher flat screens, and can support touch for receiving user input signals via touch on the display. The AVD 12 can include one or more speakers 16 for outputting audio in accordance with the principles of the application, and at least one additional input device 18, such as an audio receiver / microphone, for example, for inputting audible commands to the AVD 12, for example, to control the AVD 12. The example AVD 12 can also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, a LAN, etc., under the control of one or more processors 24. A graphics processor 24A can also be included. Thus, the interface 20 can be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It will be understood that the processor 24 controls the AVD 12 to employ the principles of the application, including the other elements of the AVD 12 described herein, such as controlling the display 14 to present images on the display and to receive input from the display. Further, it is noted that the network interface 20 can be, for example, a wired or wireless modem or router or other appropriate interface, such as a wireless telephone transceiver, or a Wi-Fi transceiver as mentioned above, etc.

[0038] In addition to the foregoing, the AVD 12 can also include one or more input ports 26, such as a high definition multimedia interface (HDMI) port or a USB port, for example, for physically connecting (e.g., using a wired connection) to another CE device, and / or a headphone port for connecting headphones to the AVD 12 for presenting audio from the AVD 12 to a user through the headphones. For example, the input port 26 can be connected, via a wire or wirelessly, to a wired or satellite source 26a of audio video content. Thus, the source 26a can be, for example, a separate or integrated set top box or satellite receiver. Alternatively, the source 26a can be a game console or disk player containing content that can be user's favorites for channel allocation purposes described further below. When implemented as a game console, the source 26a can include some or all of the components described below with respect to the CE device 44.

[0039] The AVD 12 can also include one or more computer memories 28, such as disk-based storage or solid state storage, that are not transient signals, in some cases embodied as a separate device in the housing of the AVD, or as a personal video recording device (PVR) or video disk player for playing back AV programs inside or outside the housing of the AVD, or as a removable memory media. Also in some embodiments, the AVD 12 can include a position or location receiver, such as but not limited to a cell phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic position information, for example, from at least one satellite or cell tower, and to provide that information to the processor 24 and / or to determine, in conjunction with the processor 24, an altitude at which the AVD 12 is located. However, it will be understood that another suitable position receiver other than a cell phone receiver, a GPS receiver, and / or an altimeter can be used to determine, for example, a position of the AVD 12 in all three dimensions, for example, in accordance with the principles of the application.

[0040] Continuing the description of the AVD 12, in some embodiments, the AVD 12 can include one or more cameras 32, which can be, for example, thermal imaging cameras, digital cameras such as webcams, and / or cameras integrated into the AVD 12 and capable of being controlled by the processor 24 to collect pictures / images and / or video, in accordance with the principles of the application. A Bluetooth transceiver 34 and other near field communication (NFC) elements 36 can also be included on the AVD 12 for communicating with other devices using Bluetooth and / or NFC technology, respectively. An example NFC element can be a radio frequency identification (RFID) element.

[0041] In addition, the AVD 12 can include one or more auxiliary sensors 37 (e.g., motion sensors such as accelerometers, gyroscopes, odometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, gesture sensors (e.g., for sensing gesture commands, etc.)) that provide input to the processor 24. The AVD 12 can include an over-the-air TV broadcast port 38 for receiving OTA TV broadcasts that provide input to the processor 24. In addition to the foregoing, it is noted that the AVD 12 can also include an infrared (IR) emitter and / or IR receiver and / or IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) can be provided for powering the AVD 12.

[0042] Still referring to Figure 1In addition to the AVD 12, the system 10 can include one or more other CE device types. In one example, a first CE device 44 can be used to send computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or through a server described below, while a second CE device 46 can include similar components as the first CE device 44. In the illustrated example, the second CE device 46 can be configured as an AR headset worn by a player 47, as shown. In the illustrated example, only two CE devices 44, 46 are shown, it being understood that fewer or more devices can be used.

[0043] In the illustrated example, all three devices 12, 44, 46 are assumed to be members of, for example, an in-home entertainment network, or at least to be present in proximity to one another in a location such as a house, for purposes of illustrating the principles of the application. However, the principles of the application are not limited to the particular location shown by the dashed line 48 unless explicitly required by the claims.

[0044] The example, non-limiting first CE device 44 can be established by any of the devices described above (e.g., a portable wireless laptop or notebook computer or game controller), and thus can have one or more of the components described below. The first CE device 44 can be a remote control (RC) for issuing AV play and pause commands to the AVD 12, for example, or it can be a more complex device such as a tablet computer, a game controller that communicates with the AVD 12 and / or a game console via a wired or wireless link, a personal computer, a wireless telephone, etc. The second CE device 46 can be implemented by a head-mounted display (HMD) or a head-mounted camera (HMC).

[0045] Accordingly, the first CE device 44 can include one or more displays 50 that can support touch for receiving user input signals via touch on the display. The first CE device 44 can include one or more speakers 52 for outputting audio in accordance with the principles of the application, and at least one additional input device 54, such as an audio receiver / microphone, for example, for inputting audible commands to the first CE device 44, for example, to control the device 44. The exemplary first CE device 44 can also include one or more network interfaces 56 for communicating over the network 22 under the control of one or more CE device processors 58. A graphics processor 58A can also be included. Thus, the interface 56 can be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, including a mesh network interface. It should be understood that the processor 58 controls the first CE device 44 to implement the principles of the application, including the other elements of the first CE device 44 described herein, such as controlling the display 50 to present images on the display and to receive input from the display. Further, it should be noted that the network interface 56 can be, for example, a wired or wireless modem or router or other appropriate interface, such as a wireless telephone transceiver, or a Wi-Fi transceiver as mentioned above, and the like.

[0046] In addition to the foregoing, the first CE device 44 can also include one or more input ports 60, such as an HDMI port or a USB port, for example, for physically connecting, e.g., using a wired connection, to another CE device and / or a headphone port for connecting headphones to the first CE device 44 for presenting audio from the first CE device 44 to a user through the headphones. The first CE device 44 can also include one or more tangible computer readable storage media 62, such as a disk-based storage or a solid state storage. Also in some embodiments, the first CE device 44 can include a location or positioning receiver, such as but not limited to a cell phone and / or GPS receiver and / or altimeter 64, configured to receive geographic position information from at least one satellite and / or cell tower, for example, using triangulation, and to provide that information to the CE device processor 58 and / or to determine, in conjunction with the CE device processor 58, an altitude at which the first CE device 44 is located. However, it should be understood that another suitable location receiver other than a cell phone and / or GPS receiver and / or altimeter can be used to determine, in accordance with the principles of the application, a positioning of the first CE device 44 in, for example, all three dimensions.

[0047] Continuing with the description of the first CE device 44, in some embodiments, the first CE device 44 can include one or more cameras 66, which can be, for example, thermal imaging cameras, digital cameras such as webcams, and / or cameras integrated into the first CE device 44 and capable of being controlled by the CE device processor 58 to collect pictures / images and / or video, in accordance with the principles of the present application. Also included on the first CE device 44 can be a Bluetooth transceiver 68 and other near field communication (NFC) elements 70 for communicating with other devices using Bluetooth and / or NFC technology, respectively. An example NFC element can be a radio frequency identification (RFID) element.

[0048] In addition, the first CE device 44 can include one or more auxiliary sensors 72 (e.g., motion sensors such as accelerometers, gyroscopes, odometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, gesture sensors (e.g., for sensing gesture commands), etc.) that provide input to the CE device processor 58. The first CE device 44 can include still other sensors that provide input to the CE device processor 58, such as, for example, one or more climate sensors 74 (e.g., barometers, humidity sensors, wind sensors, light sensors, temperature sensors, etc.) and / or one or more biometric sensors 76. In addition to the foregoing, it is noted that, in some embodiments, the first CE device 44 can also include an infrared (IR) emitter and / or IR receiver and / or IR transceiver 78, such as an IR data association (IRDA) device. A battery (not shown) can be provided for powering the first CE device 44. The CE device 44 can communicate with the AVD 12 through any of the aforementioned communication modes and related components.

[0049] The second CE device 46 can include some or all of the components shown for the CE device 44. Either or both CE devices can be powered by one or more batteries.

[0050] Reference is now made to the aforementioned at least one server 80, which includes at least one server processor 82, at least one tangible computer readable storage medium 84 (such as a disk-based storage or solid state storage), and at least one network interface 86 that, under control of the server processor 82, permits communication with other devices over the network 22, and in fact can facilitate communication between servers and client devices in accordance with the principles of the present application. It is noted that the network interface 86 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other appropriate interface such as a wireless telephony transceiver. Figure 1

[0051] ​Accordingly, in some embodiments, the server 80 can be an Internet server or a field of servers, and can include and perform "cloud" functionality, such that the devices of the system 10 can access a "cloud" environment via the server 80 in example embodiments such as a networked gaming application. Alternatively, the server 80 can be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown in FIG. 1. Figure 1

[0052] The methods herein can be implemented as software instructions executed by a processor, a suitably configured application-specific integrated circuit (ASIC) or field programmable gate array (FPGA) module, or any other convenient means as will be appreciated by the person skilled in the art. Where employed, the software instructions can be embodied in a non-transitory device such as a CD ROM or flash drive. The software code instructions can alternatively be embodied in a transitory arrangement such as a radio or optical signal, or via download over the internet.

[0053] Figure 2 Visual input 200 from, for example, an RGB camera used for mouth tracking 202 is shown, from which visual features 204 are extracted, including the lips 206 and the area within the mouth 208. The features 204 are augmented by key points from event signals 210 from a dynamic vision sensor (DVS) 212, which images the same person as imaged by the RGB camera and is implemented according to event-driven sensor (EDS) principles. The event signals 210 can be used to produce event-based optical flow 214, which can be used for low-latency output 216, to augment the output 218 that can be used for virtual animation of the person imaged at 200, for example.

[0054] Reference can be made to U.S. Patent No. 7,728,269 and to the "Dynamic Vision Platform" from iniVation AG of Zurich, Switzerland, which combines monochrome intensity and DVS sensor cameras, at https: / / inivation.com / dvp, both of which are incorporated herein by reference, in implementing these sensors.

[0055] An EDS consistent with the present disclosure provides an output indicative of a change in light intensity sensed by at least one pixel of a light sensing array. For example, if the light sensed by a pixel is decreasing, the output of the EDS can be -1; if it is increasing, the output of the EDS can be +1. An output binary signal of 0 can indicate no change in light intensity below a certain threshold.

[0056] Figure 3 Further shown is a virtual reality (VR) and / or augmented reality (AR) head mounted display (HMD) 300, which can be worn by the person 302, can incorporate Figure 1 ​Any of the components shown and described herein. In the illustrated example, HMD 300 can include a display 304, such as a partially transparent AR display or an opaque VR display. HMD 300 can also include left and right eye DVS cameras 306 and left and right eye RGB cameras 308 to image eyes 310, including pupils 312, of a person 302, as well as to image the person's eyebrows 314 and nose 316.

[0057] In addition, HMD 300 can include a mouth imaging DVS 318 and RGB camera 320 oriented to image a mouth 322 of person 302, including tongue 324 and teeth 326. In addition, HMD 300 can include at least one microphone 328 to detect speech of person 302.

[0058] As indicated by representation 330, the above-described sensor output signals, which can be time-stamped, output a representation 330 of the face of person 302. Display luminance information can be used to normalize the size of the imaged pupil 312.

[0059] Figure 4 An example embodiment is shown using Figure 3 The illustrated sensors. In the illustrated example, audio sensor 302, RGB sensor 304, and DVS 306 send signals to a processor, such as can be implemented by an artificial intelligence (AI) chip 400 executing an algorithm such as one or more neural networks (NNs) 402. The outputs of sensors 306 / 308 / 318 / 320 are fused together and processed by NN 402, and output labels such as emotion or speech are sent to a computer game or other application 404. Note that audio signals from audio sensor 328 can first be processed by an audio digital signal processor (DSP) before being input to NN 402, and similarly, outputs from DVS 306 can be appropriately processed before being input to NN 402. Non-machine learning algorithms can also be executed.

[0060] In the illustrated example, both the cameras and EDS, as well as processor 400, are implemented on a single chip 406, which can include local memory for storing images including images. Processing of components can be performed by a single digital signal processor (DSP). In any case, processor 400 outputs labels to one or more external applications 404, such as VR object generation algorithms, etc.

[0061] Figure 5 Further details are shown consistent with Figure 1 to Figure 4 Figure 5 ​In this embodiment, the audio input 500 from the microphone described herein is sent to a Short-Term Fourier Transform (STFT) 502 to be converted into the frequency domain. The STFT 502 outputs the signal to one or more audio processing convolutional neural networks (CNNs) 504.

[0062] Visual input 506 from the RGB camera described herein is sent to image recognition engine 508 to extract facial features, such as any of the features mentioned above. The output of engine 508 is sent to one or more visual processing CNNs 510.

[0063] The EDS input from the DVS 512 (such as any DVS described herein) is sent to the frame generator 514 to generate low-latency, high-data-rate virtual frames, which are then sent to the image filter 516 for filtering. The output 517 of the filter 516 is sent to one or more event processing CNNs 518.

[0064] like Figure 5 As shown, the outputs of CNNs 504, 510, and 518 are fed into fully connected layer 520, which, together with the CNN, forms a classifier. The output of the classifier can be used to detect the person being imaged (e.g., Figure 3 The emotions of the person in question (302) are used to identify people through speaker recognition.

[0065] The "fully connected layer" (520) is part of the network where all neurons connect to all neurons in the next layer. The classifier includes a CNN for feature vector extraction. The "fully connected layer" ultimately provides the output label. Essentially, the fully connected input layer receives the output of the CNN and "flattens" them into a single vector, which is then fed into the next stage. The fully connected layer applies weights to the features to predict the label and gives a probability for each label.

[0066] Figure 6 An alternative architecture is shown, which uses Figure 5 The components described herein up to CNNs 504, 510, and 518 for tracking the mouth of a human subject 600, in addition to a recurrent neural network 602 that may include one or more Long Short-Term Memory (LSTM) layers 604 receiving the output of the CNN and then outputting the signal to a fully connected layer 520. The RNN 602 provides temporal information encoding.

[0067] Figure 6 The architecture integrates raw event data from DVS with RGB camera and audio data as input to a classifier to overcome motion blur, intraoral regions that are difficult to image with RGB cameras, varying lighting conditions, teeth that are difficult to image, and rapid mouth movements.

[0068] Please note, Figure 5The architecture described is particularly suitable for, but not only for, human emotion recognition, where processing of ½ frame images with short time period audio is sufficient to recognize emotions. On the other hand, Figure 6 An RNN 602 is added to impart time memory to the classifier, which is useful for audiovisual speech recognition. Thus, Figure 6 The same functionality can be provided as Figure 5 with the addition of an RNN for temporal information.

[0069] Figure 7 A CNN layer implementation of Figure 5 and Figure 6 is provided for further illustration of imaging of the eye region. The eye image 700 is produced by one or both of the HMD 300 shown in Figure 3 or the HMC described elsewhere herein. From the image, eyelid and eyebrow features 702 are extracted, as are pupil segmentation 704 (normalized against display luminance) and gaze direction 706. Use of a DVS facilitates high dynamic range, which can detect structures in the mouth that are difficult to detect, and can quickly detect eye movement, lip / tongue / teeth movement, etc. These can be important attributes in estimating the emotion of a person. The DVS image facilitates segmentation of the eye image to detect eye vibrations and segmentation of the mouth image to detect lip / tongue / teeth and their movement.

[0070] Figure 8 An example overall logic is shown in example flowchart format. At block 800, audio is received from a person being imaged. At block 802, RGB images are received from any of the cameras described herein, and at block 804, event information is received from any of the DVS described herein. All three are input at block 806 to any of the classifiers described herein, which output information useful in detecting emotions, mouth movement, eye movement, etc. at block 808 for rendering a virtual image of the person or for other purposes including cropping the person's emotional background.

[0071] Figure 9 It is shown that the classifiers herein can receive a ground truth training set of audio, RGB images, and event signals at block 900. The ground truth training set can include audio / visual clips and EDS data, paired with an output of labeled emotions or speech text output. Corresponding ground truth mouth tracking, emotion classification, eye tracking, etc. are also provided at block 902. From the training set, the classifiers learn the correct output from real data. The classifiers can be trained using all three inputs (from audio / camera / event data), with the classifiers including an RNN and CNN in example implementations.

[0072] Figure 10 and Figure 11The HMC 1000 is shown as being worn by a person 1002 with the inward-facing imaging assembly 1004 oriented toward the person's 1002 face and spaced in front of them so that their eyes are not obstructed from seeing the real world. The assembly 1004 produces RGB and DVS signals representing various parts of the person's 1002 face, including her mouth 1006 and tongue 1008 and teeth 1010, nose 1012, eyebrows 1014, and eyes 1016 including pupils 1018.

[0073] As shown, the imaging assembly 1004 can include one or more DVS imagers 1100, one or more RGB cameras 1102, and one or more microphones 1104, consistent with the principles described herein. Figure 11 The signals from the various sensors in the imaging assembly 1004 are time-stamped so that they can be correlated with each other in time to produce an output 1106 representing the person's 1002 face (and sounds).

[0074] It will be appreciated that, although the principles of the application have been described in reference to some example implementations, these implementations are not intended to be limiting, and that various alternatives can be utilized in practicing the subject matter claimed herein.

Claims

1. A component for visual audio processing, comprising: At least one camera unit, the at least one camera unit being configured to generate an image of a face; At least one event-driven sensor (EDS) is configured to output a signal representing a change in illumination intensity for at least a portion of the face, the EDS providing a first output in response to a decrease in the intensity of sensed light, a second output in response to an increase in the intensity of sensed light, and a third output in response to light intensity falling below a threshold. At least one microphone, the at least one microphone being configured to output a signal representing speech; as well as At least one processor, the at least one processor being configured with executable instructions to: Receive signals from the at least one EDS representing changes in illumination intensity of at least a portion of the face, signals from the at least one microphone representing speech, and signals from the at least one camera unit representing at least a portion of the face; At least one neural network is executed to generate, based on signals representing at least a portion of the face from the at least one camera unit, signals representing speech from the at least one microphone, and signals representing changes in illumination intensity of at least a portion of the face from the at least one EDS, one of the following: emotion prediction, tracking of at least a portion of the face, or both.

2. The component of claim 1, wherein the at least one camera unit is configured to generate an infrared (IR) image.

3. The component of claim 1, wherein the at least one camera unit, the at least one processor, and the at least one EDS are disposed on a single chip.

4. The component of claim 1, wherein the instructions are capable of performing: The at least one neural network is executed to generate emotion prediction based on signals representing at least a portion of the face from the at least one camera unit, signals representing speech from the at least one microphone, and signals representing changes in illumination intensity of at least a portion of the face from the at least one EDS.

5. The component of claim 1, wherein the instructions are capable of executing to: The at least one neural network is executed to generate tracking of at least a portion of the face based on signals representing at least a portion of the face from the at least one camera unit, signals representing speech from the at least one microphone, and signals representing changes in illumination intensity of at least a portion of the face from the at least one EDS.

6. The component of claim 5, wherein the portion includes at least one eye pupil.

7. The component of claim 5, wherein the portion comprises at least one of the following: a corner of the mouth or the interior of the mouth including teeth.

8. The component of claim 1, wherein the instructions are capable of performing: The signal representing at least a portion of the face from the at least one camera unit, the signal representing speech from the at least one microphone, and the signal representing changes in illumination intensity of at least a portion of the face from the at least one EDS are fused together. The at least one neural network is executed based on the fused signals.

9. A system for visual audio processing, comprising: At least one camera unit, the at least one camera unit being configured to generate red-green-blue (RGB) images and / or infrared (IR) images of a person; At least one microphone; At least one event-driven sensor EDS, the at least one EDS being configured to output a signal representing at least a portion of the person’s face; and At least one processor, said at least one processor being programmed with instructions to: The output of the at least one microphone is processed using a Short-Term Fourier Transform (STFT). The output of the STFT is processed using at least one audio processing convolutional neural network (CNN); At least one visual processing CNN is used to process features from at least one image from the at least one camera unit; The representation of the output signal from the at least one EDS is processed using at least one event-processing CNN; and The outputs of the CNN are fused into a fully connected neural network layer to generate: The prediction of human emotions, Tracking of at least a portion of the person's face, At least one virtual reality (VR) image of the person, The identity of the person mentioned; or All four.

10. The system of claim 9, wherein the at least one processor is configured with instructions to: The output of the CNN is processed using a recurrent neural network (RNN); and The output of the RNN is processed using the fully connected neural network layer to generate the person's mouth tracking.

11. The system of claim 9, wherein the at least one processor is configured with instructions to: Generate a prediction of the person's emotions.

12. The system of claim 9, wherein the at least one processor is configured with instructions to: Generate tracking of at least a portion of the person's face.

13. The system of claim 9, wherein the at least one processor is configured with instructions to: Generate at least one virtual reality (VR) image of the person.

14. The system of claim 9, wherein the at least one processor is configured with instructions to: Generate the identity of the person.

15. A method for visual audio processing, comprising: Receive signals representing images of a face from at least one camera unit; Receive a signal representing speech from at least one microphone; Receive a signal from at least one event-driven sensor (EDS) representing a change in illumination intensity for at least a portion of the face; Execute at least one neural network to generate, based on a signal representing a change in illumination intensity of at least a portion of the face from the at least one EDS, a signal representing speech from the at least one microphone, and an image signal representing at least a portion of the face from the at least one camera unit: emotion prediction, tracking of at least a portion of the face, human identity, a virtual reality (VR) image of the human, or all four of the above. as well as Presenting at least one of the following on at least one display: emotion prediction, tracking of at least a portion of a face, human identity, or generating a virtual reality (VR) image of the person. The EDS is configured to provide a first output in response to a decrease in the intensity of sensed light, a second output in response to an increase in the intensity of sensed light, and a third output in response to a light intensity below a threshold.

16. The method of claim 15, further comprising: At least one neural network is executed to generate emotion prediction based on signals representing illumination intensity changes of at least a portion of the face from the at least one EDS, signals representing speech from the at least one microphone, and signals representing images of at least a portion of the face from the at least one camera unit.

17. The method of claim 15, further comprising: At least one neural network is executed to generate tracking of at least a portion of the face based on signals representing illumination intensity changes of at least a portion of the face from the at least one EDS, signals representing speech from the at least one microphone, and signals representing images of at least a portion of the face from the at least one camera unit.

18. The method of claim 15, further comprising: At least one neural network is executed to generate the identity of the person based on signals representing at least a portion of the face from the at least one camera unit, signals representing speech from the at least one microphone, and signals representing changes in illumination intensity of at least a portion of the face from the at least one EDS.

19. The method of claim 15, further comprising: At least one neural network is executed to generate a virtual reality (VR) image of the person based on signals representing at least a portion of the face from the at least one camera unit, signals representing speech from the at least one microphone, and signals representing changes in illumination intensity of at least a portion of the face from the at least one EDS.

20. The method of claim 17, wherein the portion of the face is the corner of the mouth and the interior of the mouth.

Citation Information

Patent Citations

  • Photoarray for detecting time-dependent image data

    US7728269B2

  • Method and apparatus of image representation and processing for dynamic vision sensor

    US20170278221A1

  • Vehicle manipulation using convolutional image processing

    US20180189581A1