Vehicle, network node and method

The vehicle-based system efficiently captures and updates 3D avatar data using onboard cameras and machine learning, addressing the cumbersome studio requirements of existing methods and providing high-fidelity, cost-effective 3D avatar generation.

WO2025172420A1PCT designated stage Publication Date: 2025-08-21SONY GROUP CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/053825
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods for generating high-quality 3D avatars, such as photogrammetry, require a cumbersome and costly studio setup, making them inaccessible to ordinary consumers.

Method used

A vehicle equipped with a set of cameras and circuitry for capturing and updating 3D avatar data of users within the vehicle, allowing for efficient and time-saving image capture and data updating, even during movement, using various camera types and machine learning algorithms for feature extraction and comparison.

Benefits of technology

Enables high-quality 3D avatar generation with high fidelity to the user's features, allowing for real-time updates and accurate representation of facial and body movements, reducing the need for studio setups and associated costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025053825_21082025_PF_FP_ABST
    Figure EP2025053825_21082025_PF_FP_ABST
Patent Text Reader

Abstract

A vehicle comprises a set of cameras for capturing images of a user of the vehicle for updating a 3D avatar of the user and circuitry for updating 3D avatar data, the 3D avatar data representing an avatar of a user. The circuitry is configured to obtain 3D avatar data of the user, capture a set of images of the user located within the vehicle with the set of cameras, and update the 3D avatar data of the user based on the captured set of images.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VEHICLE, NETWORK NODE AND METHOD

[0002] TECHNICAL FIELD

[0003] The present disclosure generally pertains to a vehicle, a network node and a method.

[0004] TECHNICAL BACKGROUND

[0005] Using a three-dimensional (3D) avatar is becoming increasingly prevalent for human interaction via technological means. For example, during long-distance interaction via PC, smartphone, tablet or the like, instead of presenting a person’s real face to their interaction partner, for example, by showing a video or an image of the person, an avatar may instead be used. Oftentimes, the avatar should represent the person accurately. For that purpose, an accurate 3D avatar of the person or the person’s face can be generated.

[0006] Although there exist techniques for generating a 3D avatar, it is generally desirable to improve existing techniques.

[0007] SUMMARY

[0008] According to a first aspect the present disclosure provides a vehicle comprising: a set of cameras for capturing images of a user of the vehicle for updating a 3D avatar of the user; and circuitry for updating 3D avatar data, the 3D avatar data representing an avatar of a user, configured to: obtain 3D avatar data of the user; capture a set of images of the user located within the vehicle with the set of cameras, and update the 3D avatar data of the user based on the captured set of images.

[0009] According to a second aspect the present disclosure provides a network node for updating 3D avatar data of a user of a vehicle configured to: obtain 3D avatar data of the user; obtain from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and update the 3D avatar data of the user based on the acquired set of images. According to a third aspect the present disclosure provides a method for updating a 3D avatar data of a user of a vehicle comprising: obtaining 3D avatar data of the user; obtaining from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and updating the 3D avatar of the user based on the acquired set of images.

[0010] Further aspects are set forth in the dependent claims, the drawings and the following description.

[0011] BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Embodiments are explained by way of example with respect to the accompanying drawings, in which:

[0013] Fig. 1 schematically illustrates an embodiment of a vehicle for updating 3D avatar data of a user;

[0014] Fig. 2 schematically illustrates an embodiment of a vehicle for updating 3D avatar data of a user with a user located in the vehicle;

[0015] Fig. 3 schematically illustrates an embodiment for updating 3D avatar data;

[0016] Fig. 4 schematically illustrates an embodiment for updating 3D avatar data including feature extraction and feature comparison;

[0017] Fig. 5 schematically illustrates an embodiment for updating 3D avatar data including feature extraction, feature comparison and face recognition;

[0018] Fig. 6 schematically illustrates an embodiment for updating 3D avatar data including a synthetic voice model of the 3D avatar data;

[0019] Fig. 7 schematically illustrates an embodiment of a network node for updating 3D avatar data;

[0020] Fig. 8 schematically illustrates an embodiment for updating 3D avatar data including feature extraction, feature comparison and face recognition;

[0021] Fig. 9 schematically illustrates an embodiment of circuitry that implements updating of 3D avatar data;

[0022] Fig. 10 schematically illustrates an embodiment of a method for updating 3D avatar data;

[0023] Fig. 11 schematically illustrates an embodiment of a method for updating 3D avatar data including feature extraction and feature comparison; Fig. 12 schematically illustrates an embodiment of a method for updating 3D avatar data including face recognition.

[0024] DETAILED DESCRIPTION OF EMBODIMENTS

[0025] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.

[0026] It has been recognized that capturing more than 10 input images for the generation of a 3D avatar, for example, a 3D face avatar, using algorithms such as photogrammetry or the like is a problem. That is, for example a studio environment may be set up for capturing the images of a particular person for generating their 3D avatar. Thus, the particular person’s likeness is captured in a clearly defined environment. A sophisticated studio or a light stage, typically used for capturing images of humans, in particular, human faces, for purposes of generating a realistic 3D avatar, for example, a realistic 3D face avatar, include multiple cameras and strobe lights that are synchronized to the cameras. It has been recognized that setting up such a studio environment is a cumbersome process, especially when targeting to capture approximately 100 images or more for a high-quality 3D (face) avatar. In this vein, an ordinary consumer may not go through with the described process, in particular, as it may be costly and time consuming, and therefore the described process may be exclusive to particular groups, for example, artists, actors, photo models or the like.

[0027] Hence some embodiments pertain to a vehicle comprising: a set of cameras for capturing images of a user of the vehicle for updating a 3D avatar of the user; and circuitry for updating 3D avatar data, the 3D avatar data representing an avatar of a user, configured to: obtain 3D avatar data of the user; capture a set of images of the user located within the vehicle with the set of cameras; and update the 3D avatar data of the user based on the captured set of images.

[0028] The vehicle may be any kind of vehicle, e g., any kind of automotive vehicle, such as a car, a truck, a bus, etc., or a train or a plane, or the like. The camera may be any type of camera, for example, an RGB camera, a time of flight (ToF) camera, such as, an indirect ToF (iToF) or a direct time of flight (dTof) camera, an infrared (ZR) camera, an event-based camera or the like.

[0029] The set of cameras may be one or more camera. The camera may be a wide-angle camera. The camera may be a video camera. The camera may follow the movement of the user.

[0030] The camera may be pivotable in any direction, may be movable along a horizontal or vertical axis, and / or may be movable along a sliding bar or the like. In case of multiple cameras, the cameras may be multiple different types of cameras.

[0031] The captured set of images may be one or more images for generating a high-quality 3D avatar. The image(s) may include a grayscale image, an RGB image, an intensity image, a depth image, a confidence image, an IR image or the like. The image may have low or high resolution. The set of images may also be a sequence of images, such as a video.

[0032] The images may capture different perspectives of the user, that is, the images may capture multiple different sides of the user. That is, the images(s) may include one or more front images, side images (left and / or right side), back images and / or top images of the user or of one or more body parts of the user, such as, of the head or upper body of a user.

[0033] Furthermore, one, multiple or all sides of any body part of a user, in particular, one, multiple or all sides of a user’s upper body including the head, more particularly, one, multiple or all sides of a user’s head or face, may be captured in the set of images. That is, the head including the face, the shoulders, the torso, the arms and / or the hands of a user, independent of the body parts being clothed / covered or unclothed, may be captured in the set of images. In particular, multiple images of a user’s face may be imaged. That is, a user’s face from multiple different angles may be imaged.

[0034] The user is a user of the vehicle, that is, the user may be an occupant of the vehicle. Therefore, the user is located in the vehicle, for example as a driver or a passenger. The user may be sitting in the vehicle, or the user may be standing in the vehicle, for example, in case of a bus. If the user is sitting in the vehicle the back of the torso may or may not be accessible to the camera as the user may be leaning back in the seat. However, the back of the head may be accessible, i.e., imageable, by the set of cameras. Also, the back of the head or any other feature, e.g., any feature not imaged or not accessible for imaging by the set of cameras, may be hallucinated based on the available set of images and / or based on predefined stored features. For example, a learning model may be used for hallucinating such features. The learning model may, for example, be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, or the like.

[0035] The vehicle may be moving while the user is located within the vehicle and while the set of images of the user are captured. For example, the user may be driving the vehicle while their (set of) images are captured. The user may be moving while being occupied in the vehicle, for example, head movements, facial expressions or movements of the limbs and torso may be possible.

[0036] Thus, the set of images of the user can be captured when the user is occupied or bound to the vehicle, for example, while driving or being driven to a location. Hence, instead of spending time for capturing a set of images appropriate for updating a 3D avatar, i.e., also appropriate for generating a 3D avatar, in particular, images appropriate for generating or updating a realistic face avatar of the user, which is inherently a time-consuming process, in a studio, time may be saved by capturing the set-of images while the user is in the vehicle. Therefore, during a time that a user is bound to the vehicle and therefore remains in the same location for a time, the set of images can be taken, thereby saving time. The updating of the 3D avatar data may be conducted essentially immediately after the capturing of the images, for example, also while the user may still be located in the vehicle. Alternatively, the updating of the 3D avatar data may be conducted while the vehicle is in a high-energy mode, for example, when the battery is full or when the energy level is above a predetermined threshold. The predetermined threshold may be a learnt threshold. The learnt threshold may be learnt based on a learning model, which may include any feature of a learning model described in this specification. The set of cameras may be integrated in the vehicle and / or mounted on the vehicle (either inside or outside the vehicle). The set of cameras may be located at the front mirror of the vehicle, on the inside of the roof of the vehicle, at the back of a vehicle seat pointing towards a backseat for imaging the front of a backseat passenger or pointing towards the front seat for imaging the backside of a front seat passenger or driver, on the side mirrors pointing into the vehicle, on the side, e.g., the inside of the doors and windows, or any other position. The field of view or sensing area of one or more cameras of set of cameras may be movable, that is the camera of the set of cameras may be pivotable, for example, horizontally and / or vertically. The camera may be movable horizontally and / or vertically in a translation motion, for example, sliding along a camera rail or camera bracket.

[0037] The circuitry may include one or more processors, logical circuits, memory (read only memory, random memory, etc., storage, e.g., hard disk, compact disc, flash drive, etc.), an interface for (wireless, e.g., Bluetooth, infrared) communication via a network (local area network, wireless network, internet). Moreover, it may include input means (mouse, keyboard, microphone, camera etc.), output means (loudspeakers, display (e.g., liquid crystal, (organic) light emitting diode, etc.)), and sensors for sensing audio data (microphone), still image or video image data (image sensor, camera sensor, video sensor, etc.). The sensors may be the set of cameras included in the vehicle. The memory may store the captured images, 3D avatar data and also updated 3D avatar data. Also, the circuitry may receive, for example, via the interface, 3D avatar data transmitted from an outside source, for example, from a server, or for example, a network node as describe below. Thus, the vehicle may be connected as a network node to other network nodes via a network. The circuitry may be configured to transmit, for example via the interface, the captured set of images and / or the 3D avatar data to outside sources, for example, to a server, or to a network node as described below.

[0038] The 3D avatar data of the user represents an avatar of the user. In other words, an avatar of a user, for example a 3D avatar, may be based on the 3D avatar data. That is, the avatar of the user may be based on all or only some of the 3D avatar data.

[0039] The circuitry may be further configured to generate an avatar, such as a 2D avatar or a 3D avatar, based on the 3D avatar data. In this way, when the 3D avatar data is updated, also an existing generated avatar may be updated.

[0040] The avatar may be generated in real-time. For example, the avatar may be generated upon obtaining the 3D avatar data and / or upon obtaining the updated 3D avatar data. In this way, the generated avatar may also be updated in real-time. Alternatively, the avatar may be generated based on stored (updated) 3D avatar data. That is, the updated 3D avatar data may be stored, for example on a server or on another device, and the stored updated 3D avatar data may be obtained from the storage to generate the avatar. Generating the avatar may be based on the available computing resources or energy.

[0041] The avatar of a user may be a visualization, for example, a 2D or 3D visualization, representing the user, for example during calls, e.g., video calls, games, or the like. For example, the avatar may be a modelled or rendered avatar.

[0042] The avatar of the user may be a human or a human-like avatar. The avatar of the user may represent one or more features of the user. In particular, the avatar of the user may be a face avatar of the user, for example, a highly realistic face avatar. The human-like avatar, in particular, the face of the avatar, may be a highly accurate rendering of the face of the user. In this way, the generated face avatar will be highly accurate. For example, a face avatar that is immediately recognizable by humans as a person of their acquaintance or a face avatar that is essentially unrecognizable by humans from a user’s image captured by a camera. Alternatively, the avatar of the user may correspond only to one or only a few of the user’s features, for example, facial features. That is, the avatar of the user may be an animal avatar, a plant avatar or an object avatar instead of a human avatar or a human face avatar. The avatar of the user may also be configurable, for example, based on user input. For example, a user may decide to be represented as a human male 3D avatar and then change this to be represented as a human female 3D avatar.

[0043] The non-human avatar may be anthropomorphic. For example, the hair color of the user may be represented in the fur color of an animal avatar or the shape of a user’s face may inform the shape of an anthropomorphic object avatar etc. Thus, the user may adapt the configurable avatar based on the 3D avatar data to be, for example, represented as an anthropomorphic train avatar.

[0044] If the avatar of the user is a human avatar, it may be an essentially accurate representation of the user, including most or all the information of the 3D avatar data. For example, the 3D avatar data may include data on facial features, such as color, position, shape and composition of the face or head or its components, e.g., eyes, nose, mouth, ears, hair etc., corresponding to the facial features of the user and the same features may be represented in the human avatar of the user. However, also other body parts of the user, for example the upper body, the torso, the shoulders, the arms and hands, may be represented as features of the 3D avatar data, which may then be used for generating an avatar of the user. The avatar features may, therefore, correspond to or represent, the features of the user.

[0045] The circuitry may be further configured to: extract a feature from the captured set of images of the user; and update the 3D avatar data of the user based on the extracted feature.

[0046] For example, extracting the feature may be conducted on an image of the captured set of images. The feature extracted from the captured set of images may include one or more, i.e., multiple, features. Also, more than one feature may be extracted, and one or more features may be extracted from one or more images. The one or more features may be extracted from one image of the set of images. The one or more features may be extracted from multiple images of the set of images. For example, one feature may be extracted from multiple images, e.g., if the images of the set of images of the user depict the same feature of the user, for example, from different angles. The feature may be a facial feature, for example, the color, shape, position and composition of the face and its parts, e g., of the eye, mouth, nose, ears, skin, facial hair etc. Also, other head features, such as hair features (shape (of the hair cut), color, position and composition (straight, curly)) or other body features, such as features of the shoulders, torso, arms and hands may be extracted.

[0047] The extracted feature may include a variable feature, such as, a (facial) hair feature, wherein, for example, beard growth may vary for a person over time, or a skin feature, as the skin may wrinkle over time, or it may become tanned based on sun exposure. Other variable features may pertain to clothing, accessories or other covering the user may be wearing, such as, glasses, hats, j ewelry or the like.

[0048] The extracted feature may include a static feature, for example pertaining to the relative position or shape of facial features, which are typically also used in biometric pictures for IDs or passports and the like.

[0049] The updating of the 3D avatar data may, therefore, be based only on a variable feature or more than one variable feature extracted from the set of images.

[0050] The circuitry may be further configured to update the 3D avatar data of the user based on comparing the extracted feature with a corresponding feature included within the 3D avatar data of the user.

[0051] For example, the 3D avatar data may include data on facial features of the 3D avatar and the feature extracted from the set of images (e.g., from an image of the set of images) may also be facial features. Then the facial features extracted from the image are compared to the corresponding facial features of the 3D avatar data. Based on this comparison the 3D avatar data, that is, the facial features of the 3D avatar data, may be updated.

[0052] The corresponding feature may include one or more features. For example, if multiple features are extracted, such as multiple different features, then the corresponding feature may also include multiple features, such as the multiple different features corresponding to the multiple extracted features.

[0053] The circuitry may be further configured to update the 3D avatar data of the user based on a difference between the extracted feature and the corresponding feature included within the 3D avatar data of the user.

[0054] The difference may include multiple differences. For example, there may be multiple difference when comparing one feature with its corresponding feature or there may be multiple differences, because multiple features are compared and there may be a difference in each compared feature pair or there may be multiple differences in each compared feature pair of multiple compared feature pairs or any combination of such.

[0055] From the image(s), which may be a video, of the user one or more features, e.g., facial features, may be extracted.

[0056] Feature extraction may occur based on typical image processing techniques, such as, typical feature extraction techniques known in the art, for example, based on computer vision techniques, such as, feature description, meanshift, triangulation and tracking.

[0057] Also, a learning model may be used for feature extraction. The learning model may, for example, be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, or the like. The learning model may include any feature of a learning model described in this specification.

[0058] The learning model may be trained in advance, for example, based on user data (e g., images of users) or based on image data of other persons.

[0059] The feature extraction may include a feature detection and a feature classification, for example, based on body part categories, e.g., head, face, eyes, arms, legs, etc., which may be hierarchical, e.g., upper body, head, face, eyes, left eye etc., and / or based on feature type, e.g., variable feature or static feature, and / or based on general features, such as body part of the user or accessories etc., or the like.

[0060] The features extracted from the image(s) may then be compared to the data of the 3D avatar data representing the corresponding features. For example, a variable facial feature, such as a beard feature, extracted from the set of images may be compared to the corresponding feature data of the 3D avatar data, which may be the beard data.

[0061] The circuitry may be further configured to update the 3D avatar data of the user by updating, with the extracted feature, the corresponding feature of the 3D avatar data based on the difference, between the extracted feature and the corresponding feature of the 3D avatar data, exceeding a predetermined or learnt threshold of difference. The learnt threshold may be based on a learning model which may include any feature of a learning model described in this specification. Also, the predetermined threshold may be a learnt threshold.

[0062] For example, if the beard of the user has been growing since the last update of the 3D avatar, which is reflected in the 3D avatar data, there might be a difference in beard length between the beard variable feature extracted from the image of the user and the beard data of the 3D avatar data representing the 3D avatar of the user. If the difference is big enough, for example, if the beard of the user has been grown 10 cm since the last update of the 3D avatar, that is, the 3D avatar data, the predetermined or learnt threshold of difference may be exceeded. In that case, the beard data of the 3D avatar data may be updated to the new beard length. In other words, the 3D avatar data is updated to reflect the extracted feature of the image. Thus, a 3D avatar of the user may be updated in line with the updated 3D avatar data to represent the updated feature. In this vein, on the basis of the 3D avatar data a 3D avatar of the user may be updated. In other words, the user’s 3D avatar face may now include a beard that is longer (corresponding to 10 cm) than before the update.

[0063] The predetermined or learnt threshold may be configurable, for example, user configurable. That is, the user may decide how sensitive the updating of the 3D avatar is to changes of the user, i.e., to variations in the user’s features.

[0064] The circuitry may be further configured to update the 3D avatar data by leaving the 3D avatar data unchanged if the difference, between the extracted feature and the corresponding feature of the 3D avatar data, is below the predetermined or learnt threshold of difference.

[0065] For example, if the beard growth of the user is minimal as compared to the representation of the beard of the user in the 3D avatar data, for example, corresponding to the beard represented in the 3D avatar of the user, the difference between the beard feature extracted from the image(s) of the user and the corresponding beard feature data of the 3D avatar data may be smaller than the predetermined or learnt threshold. Therefore, the 3D avatar data, in particular the beard feature data, is left unchanged.

[0066] The circuitry may be further configured to obtain the 3D avatar data of the user based on face recognition of the user. Face recognition may be based on known facial recognition techniques of the art, such as, computer vision techniques. Furthermore, face recognition may be based on one or more images, for example, a video, of the user’s face. Face recognition may further initialize the 3D avatar data updating. That is based on the recognized face, for example a learning model, for updating the 3D avatar data may be initialized. The learning model may have any feature of a learning model described in this specification.

[0067] The 3D avatar data updating may be based on a learning model. The learning model may have any feature of a learning model described in this specification. The learning model may, for example, be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, or the like. The learning model may be based on an encoder-decoder neural network. The learning model may be used on the set of images of the user for updating the 3D avatar data. The learning model may be used to track the movements of the body parts of the user, such as the facial expression, head movement, limb movements, arm movements, hand movements etc.

[0068] The learning model may be trained in advance, for example, based on images of a user’s face or based on image data of other persons’ faces. The learning data may be trained on one or multiple images of one or multiple, for example thousand or more, persons.

[0069] Also, a learning model may be used for face recognition. The learning model may have any feature of a learning model described in this specification. The learning model may, for example, be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, or the like.

[0070] The learning model may be trained in advance, for example, based on images of a user’s face or based on image data of other persons’ faces.

[0071] Also, one learning model may be used for updating the 3D avatar data including feature extraction from the images of the user and / or including feature comparison between the extracted features from the images and the features of the 3D avatar data and / or including face recognition. The learning model may have any feature of a learning model described in this specification.

[0072] Before the updating of the 3D avatar data is started, the circuitry may fist perform face recognition. Then, based on the recognized face, the updating may occur, for example, the capturing of the set of images for updating purposes may occur. In other words, one or more images of the user may be captured for face recognition. The images for face recognition, may, be used for updating the 3D avatar data, e.g., features extracted for updating the 3D avatar may be extracted from the image(s) captures for face recognition. Also, separate image(s) for face recognition may be captured. For example, a low resolution frontal face image may be used for face recognition and high resolution frontal face image(s) may be used for feature extraction for updating the 3D avatar data.

[0073] The circuitry may be further configured to perform face recognition based on the captured set of images. Thus, one or more of the same images used for face recognition may be used for updating the 3D avatar data. Also, the face recognition may occur after or during the updating of the 3D avatar data. The circuitry may be further configured to perform face recognition based on the extracted feature from the captured set of images.

[0074] The circuitry may be further configured to perform face recognition based on a static feature of the user, which may be an extracted feature from the image. That is, the extracted feature on which the face recognition is based may be a static feature. By contrast, the updating of the 3D avatar may occur based on a variable feature. That is the extracted feature from the image of the user may be a variable feature.

[0075] The vehicle may further comprise a microphone and the circuitry may be further configured to update a synthetic voice model of the 3D avatar data of the user based on an audio signal of the user captured with the microphone.

[0076] That is, updating the synthetic voice model may be performed in line with updating of the 3D avatar data, as described above including, for example, feature extraction and feature comparison, wherein instead of the captured images, the audio signal of the user is used. Therefore, the microphone, which may be one or more microphones located within the vehicle, may capture the voice of the user as an audio signal, i.e., the audio signal may correspond to the voice of the user.

[0077] Furthermore, a feature, which may include one or more features, may be extracted from the audio signal, which may occur based on a learning model. For example, the learning model may be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, e g., transformer model, or the like. The learning model may be trained in advance, for example based on voice samples of the user.

[0078] The extracted feature(s) may be based on variable or static features. The updating of the synthetic voice model of the 3D avatar data may be based on variable features.

[0079] Furthermore, the extracted feature(s) may be compared to the synthetic voice model, that is, for example, the corresponding features of the synthetic voice model. Therefore, based on a difference between the audio signal of the user or the extracted features of the audio signal and the synthetic voice model of the 3D avatar data or the features of the synthetic voice model, the synthetic voice model of the 3D avatar data is updated. Therefore, the synthetic voice model is updated in line with the audio signal, that is, to reflect the extracted feature(s) of the audio signal representing the current voice of the speaker, i.e., the user located in the vehicle. The updating may occur based on the difference exceeding a predetermined or a learnt threshold. The predetermined threshold may be a learnt threshold. The learnt threshold may be learnt based on a learning model, which may include any feature of a learning model described in this specification.

[0080] Also, the circuitry may be configured to perform voice recognition, for example, for initializing the 3D avatar data updating or to obtain the 3D avatar data of the user (e.g., from a databank), for example, to obtain the synthetic voice model of the 3D avatar data of the user. In other words, 3D avatar data of the user may be obtained based on the voice recognition. Voice recognition may be based on a learning model. For example, the learning model may be based on a machine learning algorithm, such as a support vector machine (SVM) or random forest, an artificial neural network, e.g., a convolutional neural network (CNN) or a multimodal Al model, e g., transformer model, or the like. The learning model may be trained in advance, for example, based on voice samples of the user.

[0081] During the time the user is bound to the vehicle a constant sequence of images, which may be a video, may be captured of the user. The set of cameras, which may be one camera, may, therefore, capture a video of the user. In this way, also the movement(s) of the user, e.g., head movement, facial expressions, etc. may be captured by the images and used for repeatedly updating the 3D avatar data.

[0082] The movement of the user may be tracked based on the set of images. The movement of the user may also be analyzed based on the set of images. The updating of the 3D avatar data may be based on the movement of the user. For example, the learning model described above may be used for updating the 3D avatar data according to the movement of the user. The learning model may include any feature described in this specification concerning learning models. For example, the updating of the 3D avatar data may be based on the tracked and / or based on the analyzed movements, which may also occur via the learning model.

[0083] The tracked movements of the user may include movements of the body, such as torso and limb movement (e.g., gestures or the like), head movements, eye movements, facial expressions, also micro-expressions (e g., movements of the wrinkles) etc. Movements may include unconscious movements (e.g., facial movements) and even unusual or pathological movements (e.g., of the eye or the limbs). The tracking may be based on movement-related feature capture, such as movement or position of wrinkles or unusual and pathological movements (e g., movement ticks). Such movements may be user specific. Thus, the tracking may be user specific and / or may be based on user specific capture of movement-related features Therefore, the learning model for tracking the movement and / or updating the 3D avatar data based on the movement may be updated based on the user specific tracking of movement-related features. For example, while tracking the user movements over time the learning model may be updated (e.g., rigged) based on the user specific movements, such as movements occurring repeatedly (e.g., movement ticks, wrinkle movement etc.).

[0084] In some embodiments, the set of images may include a sequence of images of the user. The sequence of images may be a video.

[0085] In some embodiments updating the 3D avatar data may include repeated updating of the 3D avatar data based on the set of images, for example based on the sequence of images.

[0086] In some embodiments the extracted feature may include multiple features as described above and updating the 3D avatar data may be based on comparing the extracted features with corresponding features included within the 3D avatar data. Thus, the circuitry may be further configured to extract multiple features, e.g., from the video, wherein updating the 3D avatar data may be based on comparing the extracted features, e.g., from the video, with corresponding features included within the 3D avatar data.

[0087] In some embodiments, the sequence of images, for example the video, may capture a movement of the user and the extracted feature(s) may include a movement of the user.

[0088] In some embodiments, updating the 3D avatar data may include representing the movement of the user within the 3D avatar data. Updating the 3D avatar data of the user may occur repeatedly in a way that the updated 3D avatar data follows the natural movement of the user.

[0089] In turn, a repeatedly updated 3D avatar may be generated which represents the movements, e.g., facial expressions, head movements, hand movements, etc., of the user accurately.

[0090] That is, the generated avatar and the updated 3D data on which basis the avatar may be generated, may also represent the motion of the user, e.g., of the driver or the passenger. Any feature of the user or avatar described in this specification may be captured in motion and represented by the 3D avatar data, and also the generated avatar. For example, the movement of the head, facial expressions, the torso and / or limbs, of the driver or passenger may be captured and represented by the 3D avatar data and reflected in the generated avatar.

[0091] This has the distinct advantage of a very high accuracy of the generated avatar, for example, the head or face avatar, resulting in high fidelity to the actual user’ head or face.

[0092] The constant image capture or video filing of the user within the vehicle allows essentially realtime updates of the 3D avatar data and of the generated avatar. Also, any of the learning models based on image data described above may be trained based on a sequence of images capturing the motion of the subject of the images, i.e., the user or other people. Thus, the learning model may be used for repeatedly updating the 3D avatar data, for example, for updating the 3D avatar data for movement of the user.

[0093] The circuitry may be further configured to upload the 3D avatar data to a server for sharing the 3D avatar data with other users. For example, the 3D avatar data may be shared and uploaded to an online platform, e.g., a social media platform, accessible to multiple people. In this way the 3D avatar data may be shared with a digital artist or an artificial intelligence (Al) model operator. The Al model operator may then use the 3D avatar data to generate an avatar or generate various different types of Al generated models. In turn, a user updating their 3D avatar data may receive monetary compensation for sharing their 3D avatar data.

[0094] Some embodiments pertain to a network node for updating 3D avatar data of a user of a vehicle, the 3D avatar data representing an avatar of the user, configured to: obtain 3D avatar data of the user; obtain from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and update the 3D avatar data of the user based on the acquired set of images.

[0095] A network node may be, for example, a server or an electronic device, such as a terminal computer, a mobile device, such as a smartphone, tablet or the like. The network node may also be a vehicle, for example, a vehicle as described above. The network node may be connected via a network to one or more other network nodes.

[0096] The vehicle from which the network node obtains the set of images of the user, may be a vehicle as described above or it may be a vehicle as described above, but without the circuitry configured to update the 3D avatar data. Instead, the vehicle may send the set of images captured by its camera(s) to the network node for updating the 3D avatar. Thus, time can be saved by using the time a user is occupied in the vehicle for capturing a set of images. At the same time the vehicle may only send the captured set of images to the network node if the vehicle is in a high-energy mode and / or when there is a high data transmission rate available. That is the vehicle may wait to send the captured set of images data to the network node if the transmission rate is above a predetermined threshold, for example, if the vehicle is at a predetermined distance to a base station (of a mobile telecommunication system) or the signal received from a base station is above a predetermined threshold or if the vehicle is at a predetermined distance to the network node. The predetermined threshold may be a learnt threshold. The learnt threshold may be learnt based on a learning model, which may include any feature of a learning model described in this specification. Alternatively, the user may transfer the captured set of images from the vehicle to a mobile external memory, for example, a flash drive. Afterwards, the user may transfer the captured set of images from the external memory to the network node, for example, to a terminal computer. Also, the vehicle may send the captured set of images via a network node to another network node, for example, from the vehicle to a smartphone to a server.

[0097] The network node may include circuitry as explained above in regard to the vehicle including circuitry for updating 3D avatar data. The network node may also implement any one or more of all the processes described above. For example, if the updated 3D avatar data is stored in the network node or in another device, server or the like, the network node may obtain the 3D avatar data and / or generate the avatar based on the energy-mode or the data transmission rate. For example, when the network node is in a predetermined (or learnt) high-energy mode or when there is a predetermined (or learnt) high data transmission rate, e.g., based on a predetermined or learnt threshold, the network node may generate the updated 3D avatar.

[0098] Also, any processing of the network node, such as capturing the set of images, data transmission and reception (e.g., of the set of images as explained above with regard to the vehicle, or of the 3D avatar data, i.e., obtaining the 3D avatar data, or of the updated 3D avatar data, or the like), updating the 3D avatar data and / or generating an avatar may be based on the energy -mode and / or the transmission rate. Thresholding, based on predetermined thresholds, may be used to determine the sufficiency of the energy of the network node and transmission rate of the network. The predetermined threshold may be a learnt threshold. The learnt threshold may be learnt based on a learning model, which may include any feature of a learning model described in this specification.

[0099] The 3D avatar data may be obtained by the network node from the memory (of the circuitry) included in the network node, or from circuitry, for example, the memory, included in another network node.

[0100] The network node, for example, the server or the terminal device, for example the home computer of the user, may therefore, receive the captured images of the user and update the 3D avatar data based on the images. The network node may receive the captured images and perform the updating long after the 3D images have been captured by the vehicle. The network may be any kind of network, for example a local area network, wireless network, internet etc.

[0101] The network node for updating the 3D avatar data of the user of the vehicle may be further configured to save the set of images and / or the updated 3D avatar data within a user profile authorized and accessible by the user. Furthermore, the user profile may be data protected. In other words, the images and 3D avatar data may be data protected and connected to the user profile. Thus, the images and / or 3D avatar data may be transmitted from network node to network node without losing the connection to the user.

[0102] Some embodiments pertain to a method for updating a 3D avatar data of a user of a vehicle, the 3D avatar data representing an avatar of the user, comprising: obtaining 3D avatar data of the user; obtaining from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and updating the 3D avatar of the user based on the acquired set of images.

[0103] The method may also include any one or more of the features described in this specification, for example concerning the vehicle, the network node or the circuitry. Also, the vehicle from which the set of images of the user are obtained may be a vehicle as described above, or it may be a vehicle as described above but without the circuitry for performing 3D avatar data updating or it may be a network node as described above.

[0104] Thus, the method may further include extracting a feature from the captured set of images of the user, wherein updating the 3D avatar data of the user may be based on the extracted feature.

[0105] The method may further include updating the 3D avatar data of the user based on comparing the extracted feature with a corresponding feature included within the 3D avatar data of the user.

[0106] The method may further include updating the 3D avatar data of the user based on a difference between the extracted feature and the corresponding feature included within the 3D avatar data of the user.

[0107] The method may further include updating the 3D avatar data of the user by updating, with the extracted feature, the corresponding feature of the 3D avatar data based on the difference, between the extracted feature and the corresponding feature of the 3D avatar data, exceeding a predetermined or learnt threshold of difference. The predetermined threshold may be a learnt threshold. The learnt threshold may be learnt based on a learning model, which may include any feature of a learning model described in this specification.

[0108] The method may further include updating the 3D avatar data by leaving the 3D avatar data unchanged if the difference, between the extracted feature and the corresponding feature of the 3D avatar data, is below the predetermined or learnt threshold of difference.

[0109] The method may further include obtaining the 3D avatar data of the user based on face recognition of the user.

[0110] The method may further include performing face recognition based on the captured set of images.

[0111] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.

[0112] Returning to Fig. 1, an embodiment of a vehicle for updating 3D avatar data of a user is schematically illustrated. Vehicle 1 includes cameras 2, a rearview mirror 3, seats 4, windows 5 a backside 6, and microphones 7. The cameras 2 include a RGB camera 2a and a ToF camera 2b located at the backside and pointing towards the inside of the vehicle through the backside window as well as a RGB camera 2a and a ToF camera 2b located at the inside roof pointing at the seats 4. A camera array 2c is located at the front of the vehicle 1 on the inner roof. A microphone 7 is located on the camera array panel. The cameras of camera array 2c may point in different directions inside the vehicle. The cameras 2 may be pivotable. The cameras 2 may pivot depending on the imaging target, which may be a user located within the vehicle. The cameras 2 may capture images of a user located within the vehicle. Microphones 7 may capture the voice of the occupants of the vehicle. One microphone 7 is located close to the front seats at the panel of camera array 7 and another is located close to the backseats. Other locations for microphones 7 within the vehicle 1 may be possible. The user may be sitting on any of the seats 4, for example, in the driver’s seat or in a passenger seat at the front or in the back. The microphone 7 may capture audio signals of the voice of the occupants of the vehicle, which may be used for updating a synthetic voice model of the 3D avatar data as described in Fig. 6. The vehicle may include circuitry, for example circuitry 100 of Fig. 9 for updating a 3D avatar data of a user as described in more detail in Figs. 2 to 7. Also, vehicle 1 may be network node 43 as described in Fig. 7 and it may include circuitry for sending the captured images to other network nodes for updating 3D avatar data. Cameras 2, 2a, 2b and 2c may be implemented as cameras 106 of Fig. 9 and microphones 7 may be implemented as microphone 107 of Fig. 9.

[0113] Fig. 2 schematically illustrates an embodiment of a vehicle for updating 3D avatar data of a user with a user located in the vehicle. Vehicle 1 includes seats 4, cameras 2 including a wide-angle roof camera 2d and two RFB cameras located on the inner left and inner right side of the vehicle 1. A user 8 is located in vehicle 1. User 8 is sitting in the backseat of vehicle 1. A driver (not visible) may be driving the vehicle 1. The user sits within the imaging area 9. Imaging area 9 represents the area imageable by the cameras 2a, 2d. That is, all objects and persons located within the imaging area 9 can be captured by the cameras 2a, 2d as images. Thus, the head including the face, the right hand and arm as well as the left arm and hand (not visible), the upper body including the shoulders and torso as well as part of the upper legs of user 8 are captured by the cameras 2a, 2d. In particular, the top and sideview of the aforementioned body parts of user 8 are imaged and the front is imaged by camera 2d, which may be movable horizontally and / or vertically, for example, which may move vertically up- and down along the body of the user 8. The vehicle may include circuitry, for example circuitry 100 of Fig. 9 for updating a 3D avatar data of a user as described in more detail in Figs. 2 to 7. Also, vehicle 1 may be network node 44 as described in Fig. 7 and it may include circuitry for sending the captured images to other network nodes for updating 3D avatar data. The cameras 2a, 2d may be implemented as cameras 106 of Fig. 9.

[0114] Fig. 3 schematically illustrates an embodiment for updating 3D avatar data. A set of cameras 2 of a vehicle, for example, cameras 2, 2a-d of vehicle 1 of Figs. 1 and 2, capture a set of images 10 of a user (e.g., 8, Fig. 2) located within the vehicle. The set of images 10 of the user as well as 3D avatar data 13 of the user (8, Fig. 2) are used as input for the process 14 of updating 3D avatar data, which results in updated 3D avatar data 15 of the user. Hence, 3D avatar data 13 of the user is obtained for 3D avatar data updating 14.

[0115] A circuitry, for example, circuitry 100 of Fig. 9, may execute the process of updating 14 the 3D avatar data and obtaining the set of images 10 and 3D avatar data 13. In particular, the updating 14 may be implemented in one or more processors of the circuitry, e.g., processors such as the CPU 101 of Fig. 9. Alternatively, or additionally, updating 14 may be based on a machine learning model and implemented in an artificial intelligence processor, for example in Al processor 110 of Fig. 9. The circuitry (e g., 100 of Fig. 9) may be included in vehicle 1 (Figs. 1 to 2) and / or it may be implemented in a network node, for example, network node 41, 42, 43 or 44 of Fig. 7. The cameras may be implemented as cameras 106 of Fig. 9. 3D avatar data 13 may be obtained from storage within the circuitry, for example, storage 102 of Fig. 9. 3D avatar data 13 may be transmitted via interface, for example Bluetooth 104 or WLAN 105 of Fig. 9, from an external source, for example, a network node as described in Fig. 7, or via a drive, in case, the 3D avatar data may be stored in an external storage medium inserted in the drive. If the vehicle including the set of cameras is a network node as described in Fig. 7 it may send the set of images 10 of the user, that is, the image data, to a network node that includes circuitry for performing 3D avatar data updating. The image data may be transmitted via WLAN 105 or Bluetooth 104 of the circuitry 100 of Fig. 9

[0116] Fig. 4 schematically illustrates an embodiment for updating 3D avatar data including feature extraction and feature comparison. A set of cameras 2 of a vehicle, for example, cameras 2, 2a-d of vehicle 1 of Figs. 1 and 2, capture a set of images 10 of a user (8, Fig. 2) located within the vehicle. The set of images 10 of the user are used as input for the feature extraction 11, which results in one or more extracted features from at least one image. Thus, the extracted features may be extracted from multiple images. The at least one extracted feature from at least one image as well as 3D avatar data 13 are used as input for feature comparison 12. Hence, 3D avatar data 13 of the user is obtained for feature comparison 12. For feature comparison 12 an extracted feature from the image of the user is compared to its corresponding feature from the 3D avatar data 13. More than one extracted feature may be compared to their corresponding feature in the 3D avatar data 13. If the compared features exceed a predetermined threshold of difference, the 3D avatar data 13, that is, the feature of the 3D avatar data 13 on which basis the comparison is made, is updated in a way to reflect the extracted feature from the set of images 10, which results in updated 3D avatar data 15. If multiple features are compared and differences above a predetermined threshold are found 3D avatar data 13 is updated for multiple different features.

[0117] A circuitry, for example, circuitry 100 of Fig. 9, may execute feature comparison, feature extraction and updating the 3D avatar data and obtaining the set of images 10 and 3D avatar data 13. In particular, feature extraction 11, feature comparison 12 and / or updating 14 may be implemented in one or more processors of the circuitry 100, e.g., processors such as the CPU 101 of Fig. 9. Feature extraction 11 and / or feature comparison 12 and / or updating 14 may be based on one or more a machine learning models and implemented in an artificial intelligence processor, for example in Al processor 110 of Fig. 9. The circuitry 100 may be included in vehicle 1 (Figs. 1 to 2) and / or it may be implemented in a network node, for example, network node 41, 42, 43 or 44 of Fig. 7. The cameras may be implemented as cameras 106 of Fig. 9. The set of images 10 captured by the vehicle cameras and / or 3D avatar data 13 may be obtained from storage within the circuitry, for example, storage 102 of Fig. 9. 3D avatar data 13 may be transmitted via interface, for example Bluetooth 104 or WLAN 105 of Fig. 9, from an external source, for example, a network node as described in Fig. 7, or via a drive, in case, the 3D avatar data may be stored in an external storage medium inserted in the drive. If the vehicle including the set of cameras is a network node as described in Fig. 7 it may send the set of images 10 of the user, that is, the image data, to a network node that includes circuitry for performing 3D avatar data updating including feature extraction 11 and feature comparison 12. The image data may be transmitted via WLAN 105 or Bluetooth 104 of the circuitry 100 of Fig. 9.

[0118] Fig. 5 schematically illustrates an embodiment for updating 3D avatar data including feature extraction, feature comparison and face recognition. A set of cameras 2 of a vehicle, for example, cameras 2, 2a-d of vehicle 1 of Figs. 1 and 2, capture a set of images 10 of a user (8, Fig. 2) located within the vehicle. The images 10 include images of the face of the user. The set of images 10 of the user are used as input for the feature extraction 11, which results in one or more extracted features from at least one image. Thus, the extracted features may be extracted from multiple images. The extracted features include facial features. Based on the extracted facial features face recognition 11 is performed. That is, the face of the user of the vehicle is recognized. Based on the recognized face, 3D avatar data 13 of the user is obtained, for example from a database of 3D avatar data 13 of multiple users. The at least one extracted feature from at least one image as well as the 3D avatar data 13 are used as input for feature comparison 12. For feature comparison 12 an extracted feature from the image of the user is compared to its corresponding feature from the 3D avatar data 13. More than one extracted feature may be compared to their corresponding feature in the 3D avatar data 13. If the compared features exceed a predetermined threshold of difference, the 3D avatar data 13, that is, the feature of the 3D avatar data 13 on which basis the comparison is made, is updated in a way to reflect the extracted feature from the set of images 10, which results in updated 3D avatar data 15. If multiple features are compared and differences above a predetermined threshold are found the 3D avatar data is updated for multiple different features.

[0119] A circuitry, for example, circuitry 100 of Fig. 9, may execute feature comparison, face recognition, feature extraction and updating the 3D avatar data and obtaining the set of images 10 and 3D avatar data 13. In particular, feature extraction 11 and / or feature comparison 12 and / or face recognition 16 and / or updating 14 may be implemented in one or more processors of the circuitry 100, e.g., processors such as the CPU 101 of Fig. 9. Feature extraction 11 and / or feature comparison 12 and / or face recognition 16 and / or updating 14 may be based on one or more machine learning models and implemented in an artificial intelligence processor, for example in Al processor 110 of Fig. 9. The circuitry 100 may be included in vehicle 1 (Figs. 1 to 2) and / or it may be implemented in a network node, for example, network node 41, 42, 43 or 44 of Fig. 7. The cameras may be implemented as cameras 106 of Fig. 9. The set of images 10 captured by the vehicle cameras and / or 3D avatar data 13 may be obtained from storage within the circuitry, for example, storage 102 of Fig. 9. 3D avatar data 13 may be transmitted via interface, for example Bluetooth 104 or WLAN 105 of Fig. 9, from an external source, for example, a network node as described in Fig. 7, or via a drive, in case, the 3D avatar data may be stored in an external storage medium inserted in the drive. If the vehicle including the set of cameras is a network node as described in Fig. 7 it may send the set of images 10 of the user, that is, the image data, to a network node that includes circuitry for performing feature extraction, face recognition, feature comparison and 3D avatar data updating. The image data may be transmitted via WLAN 105 or Bluetooth 104 of the circuitry 100 of Fig. 9.

[0120] Fig. 6 schematically illustrates an embodiment for updating 3D avatar data including a synthetic voice model of the 3D avatar data. A set of cameras 2 of a vehicle, for example, cameras 2, 2a-d of vehicle 1 of Figs. 1 and 2, capture a set of images 10 of a user (8, Fig. 2) located within the vehicle. A set of microphones 7, which may be one microphone, capture the voice of the user as an audio signal 17. The set of images 10 and the audio signal 17 correspond to the same user. The set of images 10 of the user and the audio signal 17 of the user, that is, of the user’s voice, are used as input for the feature extraction 11, which results in one or more extracted features from at least one image and one or more extracted features from the audio signal. The extracted features from the set of images may be extracted from multiple images. The at least one extracted feature from at least one image and the extracted feature(s) from the audio signal as well as 3D avatar data 13 including a synthetic voice model of the user are used as input for feature comparison 12. Hence, 3D avatar data 13 of the user including a synthetic voice model of the user is obtained for feature comparison 12.

[0121] For feature comparison 12 an extracted feature from the image of the user is compared to its corresponding feature from the 3D avatar data 13 and an extracted feature or multiple extracted features from the audio signal are compared to the corresponding synthetic voice model of the 3D avatar data. Also, more than one extracted feature from the image may be compared to their corresponding feature in the 3D avatar data 13. If the compared features exceed a predetermined threshold of difference, the 3D avatar data 13, that is, the feature of the 3D avatar data 13 on which basis the comparison is made, is updated in a way to reflect the extracted feature from the set of images 10 or from the audio signal 17, which results in updated 3D avatar data 15, which may include an updated synthetic voice model. That is, the extracted feature(s) from the audio signal and the feature(s) of the synthetic voice model of the 3D avatar data are compared 12. In this vein, if the difference between the extracted feature(s) from the audio signal and the synthetic voice model exceed a predetermined difference, the synthetic voice model of the 3D avatar data 13 is updated 14 to reflect the extracted features of the audio signal 17, i.e., corresponding to the current voice of the user, which results in an updated synthetic voice model and therefore updated 3D avatar data 15 of the user.

[0122] If multiple features are compared and differences above a predetermined threshold are found 3D avatar data 13 is updated for multiple different features.

[0123] A circuitry, for example, circuitry 100 of Fig. 9, may execute feature extraction, feature comparison and updating the 3D avatar data as well as obtaining the set of images 10, the audio signal 17 and 3D avatar data 13. In particular, feature extraction 11, feature comparison 12 and / or updating 14 may be implemented in one or more processors of the circuitry 100, e.g., processors such as the CPU 101 of Fig. 9. Feature extraction 11, feature comparison 12 and / or updating 14 may be based on one or more machine learning models and implemented in an artificial intelligence processor, for example in Al processor 110 of Fig. 9. The circuitry 100 may be included in vehicle 1 (Figs. 1 to 2) and / or it may be implemented in a network node, for example, network node 41, 42, 43 or 44 of Fig. 7. The cameras may be implemented as cameras 106 of Fig. 9. The microphone 7 may be implemented as microphone 107 of Fig. 9. The set of images 10 captured by the vehicle set of cameras, audio signal 17 and / or 3D avatar data 13 may be obtained from storage within the circuitry, for example, storage 102 of Fig. 9. 3D avatar data 13 may be transmitted via interface, for example Bluetooth 104 or WLAN 105 of Fig. 9, from an external source, for example, a network node as described in Fig. 7, or via a drive, in case, the 3D avatar data may be stored in an external storage medium inserted in the drive. If the vehicle including the set of cameras and the set of microphones is a network node 44 as described in Fig. 7 it may send the set of images 10 of the user, that is, the image data, and the audio signal 17 of the user, to a network node that includes circuitry for performing feature extraction, feature comparison and 3D avatar data updating. The image data may be transmitted via WLAN 105 or Bluetooth 104 of the circuitry 100 of Fig. 9.

[0124] Fig. 7 schematically illustrates an embodiment of a network node for updating 3D avatar data. The network 40 is coupled to multiple network nodes. A server 41, a smartphone 42 and a home terminal computer 43 constitute a network node. But also, a vehicle 44 constitutes a network node. The vehicle 44 as a network node may be vehicle 1 of Figs. 1 or 2. That is, the set of images of the user located in the vehicle 44 are captured by cameras 2 (Figs. 1, 2) included in vehicle 44. Updating of the 3D avatar data of the user may be performed in the vehicle 44 as described in Figs. 1 to 6. In that case, the vehicle 44 may obtain 3D avatar data from any one of the network nodes 41 to 43 and may send updated 3D avatar data to any one of the network nodes 41 to 43. However, updating may also occur in any one of the other network nodes 41 to 43 or in another network node 44, that is a vehicle 44 capable of updating 3D avatar data. In that case, the set of images, that is, the image data, may be transferred via network 40 to either, the server 41 or the smartphone 42 or the terminal computer 43 or another vehicle 44 for purposes of updating 3D avatar data. Any one of the network nodes 41 to 44 may have circuitry for updating 3D avatar data as described in Figs. 3 to 6. Thus, based on the obtained set of images (i.e., the image data), the network nodes may perform 3D avatar data updating as described in Fig. 3 to 6, which may include obtaining the 3D avatar data (e.g., 13 of Figs. 3 to 6) and which may include feature extraction (12, Figs. 4-6) and feature comparison (Figs. 4-6) and / or face recognition (16, Fig. 5). Furthermore, vehicle 44 may capture an audio signal (e.g., 17 of Fig. 6) of the user via microphone (7, Figs. 1, 6) for updating the synthetic voice model as described in Fig. 6. The vehicle 44 may also transfer the audio signal to any one of the network nodes for updating the synthetic voice model of the 3D avatar data as described in Fig. 6. The capture of the images and / or the capture of the audio signal may occur in vehicle 44 while the user is occupied with being located in the vehicle 44, for example, occupied with driving or being driven. Sending or receiving data, e.g., captured set of images, audio signal, 3D avatar data and / or updated 3D avatar data may occur via network 40, which may be a local area network, wireless network, the internet etc., or may be performed by the user, who might transfer the data via a portable storage medium, which may for example, be inserted in a drive of vehicle 44 and / or the terminal computer 43, the smartphone 42 or a similar device. Similarly, a transmission via Bluetooth, NFC or the like between network nodes is possible. Transmitting the captured set of images, audio signal, and / or updated 3D avatar data from the vehicle 44 may occur at a predetermined time after the capture of the images and / or audio signal, for example, at a time when the transferrable data transmission rate exceeds a predetermined threshold. Also, data transmission of any data may occur in multiple steps, that is, through multiple network nodes For example, from vehicle 44 to smartphone 42 to server 41 or the like.

[0125] Fig. 8 schematically illustrates an embodiment for updating 3D avatar data including feature extraction, feature comparison and face recognition. The user can access a user profile 47, for example, via smartphone 42. The user profile is stored in the server 41. The user profile 47 includes a user ID 46 for identifying the user’s corresponding user profile 47 and a database 45. The database 45 may store the captured set of images of the user (10, Fig. 3 to 6), the captured audio signal (17, Fig. 6) of the user, the 3D avatar data (13, Figs. 3 to 6) as well as the updated 3D avatar data (15, Figs. 3 to 6). Furthermore, extracted features from the set of images and / or from the audio signal may also be stored in database 45. For example, if a vehicle, such as vehicle 1 of Figs. 1 and 2 or vehicle 44 of Fig. 7, captures a set of images (10, Fig. 3 to 6), the set of images (10, Fig. 3 to 6) may be stored in database 45 of user profile 47 accessible to the user (8 of Fig. 2). Similarly, if a vehicle, for example, vehicle 1 of Figs. 1 or vehicle 44 of Fig. 7, captures an audio signal (17, Fig. 6) of the user, the audio signal may be stored in database 45 of user profile 47 accessible to the user. Also, the updated 3D avatar data (15, Figs. 3 to 6) may be stored in database 45 after updating (14, Figs. 3 to 6). Thus, the 3D avatar data may be transferred to database 45 of user profile 47 after completion of the updating process, for example, as described in Figs 3 to. 7. Similarly, the 3D avatar data (13, Figs. 3 to 6) may be obtained from database 45 of user profile 47. For example, if face recognition 16 is performed for obtaining 3D avatar data 13 for updating 14 as described in Fig. 5, the user ID 46 may be used to access user profile 47 including database 45 with 3D avatar data 13 corresponding to the recognized face.

[0126] User profile 47 may also be accessible through any computer, for example, a terminal computer, a mobile device, such as a tablet, smartphone etc. or a vehicle, such as vehicle 1, 44 or the like. Furthermore, the user profile may be stored in a storage of any of the network nodes 41 to 44 of Fig. 7 and the user profile 47 may be accessible through any network node 41 to 44 by the user.

[0127] Fig. 9 schematically illustrates an embodiment of circuitry that implements updating of 3D avatar data. The circuitry 100 may be implemented in a vehicle, for example, vehicle 1 of Figs. 1 and 2 or 44 of Fig. 7. The circuitry 100 includes a CPU 101 as processor. Additionally, or alternatively, other computation hardware, such as GPU, TPU, DSP etc. may be used. The circuitry 100 further includes camera(s) 206, microphone(s) 107 and loudspeaker(s) 108 that are connected to the processor 101. The processor 101 may for example implement 3D avatar updating, feature extraction, feature comparison and / or face recognition according to the processes described with regard to Figs. 3 to 6. The microphone 107 may be configured to receive any kind of audio signal and may be microphone 7 of vehicle 1 of Figs. 1 and 6. The camera 106 may be one or more cameras, such as an RGB camera, and IR camera, a ToF camera, for example, an iToF or dTof, an event-based camera or the like and may be camera 2, 2a-2d of Figs. 1 or 2. The camera 106 may be implemented to capture the set of images 10 as described in Figs. 3 to 6. The circuitry 100 further includes a user interface 109 that is connected to the processor 101. This user interface 109 acts as a man-machine interface and enables a dialogue between a user (e.g., user 8 of Fig. 2) and the circuitry 100. For example, a user may make configurations to the system using this user interface 109.

[0128] The circuitry 100 further includes a Bluetooth interface 104, and a WLAN interface 105. These units 104, 105 act as I / O interfaces for data communication with external devices. An ethemet interface may also be possible. For example, additional loudspeakers, microphones, and cameras, e.g., a ToF camera, RGB camera or an event-based camera with WLAN or Bluetooth connection may be coupled to the processor 101 via these interfaces 104 and 105. That is cameras 2, 2a to 2d of Figs. 1 to 2 may be coupled to the processor 101 via the interfaces 104 and 105.

[0129] The circuitry 100 further includes a data storage 102 and a data memory 103 (here a RAM). The data memory 103 is arranged to temporarily store or cache data or computer instructions for processing by the processor 101. The data storage 102 is arranged as a long-term storage, e.g., of the captured set of images (e.g. 10 of Figs. 3 to 6), of the captured audio signal (e.g., 10 of Fig. 6), of the 3D avatar data (e g., 13 of Figs. 3 to 6) and / or the updated 3D avatar data (e.g., 15 of Figs. 3 to 6), which may be obtained via the processor 101 that may implement updating 14 of the 3D avatar data including face recognition 16 and / or feature extraction 11 and feature comparison 12.

[0130] The connection between the processor 101 and the camera 106 may include a camera serial interface (CSI). The CSI is an interface between a camera 106 and a host processor 101. Thus, control signals and data from the processor 101 to the camera 106 as well as from the camera 106 to the processor 101 may be sent.

[0131] Furthermore, the circuitry 101 includes an artificial intelligence (Al) processor 110. The Al processor 110 may include a graphics processing unit (GPU) and / or a tensor processing unit 20 (TPU). The Al processor 110 may be configured to execute an Al model (e.g., an artificial neural network), for example, the machine learning model for updating the 3D avatar data, for feature extraction, feature comparison and / or face recognition of Figs. 3 to 6.

[0132] Fig. 10 schematically illustrates an embodiment of a method for updating 3D avatar data. At 51 3D avatar data of a user is obtained. At 52 a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle is obtained. At 53 the 3D avatar data of the user is updated based on the acquired images. The vehicle may be vehicle 1 of Figs. 1 and 2 or vehicle 44 of Fig. 7. Thus, the set of cameras may be cameras 2, 2a-2d, 106 of Figs. 1, 2 and 9. Fig. 11 schematically illustrates an embodiment of a method for updating 3D avatar data including feature extraction and feature comparison. At 51 3D avatar data of a user is obtained. At 52 a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle is obtained. At 54 a feature is extracted from the captured set of images of the user. At 55 the extracted feature is compared with a corresponding feature included within the 3D avatar data of the user. At 56 the corresponding feature of the 3D avatar data is updated with the extracted feature (e.g., from the image) based on a difference between the extracted feature and the corresponding feature of the 3D avatar data. For example, based on the difference exceeding a predetermined or learnt threshold of difference. The vehicle may be vehicle 1 of Figs. 1 and 2 or vehicle 44 of Fig. 7. Thus, the set of cameras may be cameras 2, 2a-2d, 106 of Figs. 1, 2 and 9.

[0133] Fig. 12 schematically illustrates an embodiment of a method for updating 3D avatar data including face recognition. At 52 a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle is obtained. At 54 a feature from the captured set of images of the user is extracted. For example, a feature from an image of the captured set of images is extracted. At 57 face recognition is performed based on the extracted feature. At 58 3D avatar data of the user is obtained based on face recognition. At 55 the extracted feature (e.g., from the image) is compared with a corresponding feature included within the 3D avatar data of the user. At 56 the feature of the 3D avatar data is updated with the corresponding extracted feature (e.g., from the image) based on a difference between the extracted feature and the corresponding feature of the 3D avatar data. For example, based on the difference exceeding a predetermined threshold of difference. The vehicle may be vehicle 1 of Figs. 1 and 2 or vehicle 44 of Fig. 7. Thus, the set of cameras may be cameras 2, 2a-2d, 106 of Figs. 1, 2 and 9.

[0134] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example, the ordering of 51 and 52 in the embodiment of Fig. 10 may be exchanged. Also, the ordering of 51, 52, 54 and 55 in the embodiment of Fig. 11 may be exchanged, for example, for the ordering of 52, 51, 54 and 55 or 52, 54, 55 and 51. Further, also the ordering of 52, 54 and 57 in the embodiment of Fig. 12 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person.

[0135] The methods as described above can also be implemented as a computer program causing a computer and / or a processor, such as processor 101 discussed above, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non- transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.

[0136] All entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.

[0137] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.

[0138] Also, some embodiments described in this specification may pertain to the adoption of a vehicle environment for capturing a vast number of photos of, for example, a person’s face when sitting inside the vehicle by making use of the vehicle’s built-in cameras, thereby resembling and replacing a photo studio. The particular conditions of the drivers and passengers being positioned in the same place during a drive, for example, a long drive, is, thusly, exploited in the capturing process. Having this vast number of captured photos available may allow the use of the photos for generating or updating a super-realistic 3D face avatar. However, it may also allow the use of the photos as input for training a generative Al system.

[0139] Note that the present technology can also be configured as described below.

[0140] [1] A vehicle comprising: a set of cameras (2) for capturing images of a user (8) of the vehicle for updating a 3D avatar of the user (8); and circuitry (100) for updating 3D avatar data (13), the 3D avatar data (13) representing an avatar of a user (8), configured to: obtain 3D avatar data (13) of the user (8); capture a set of images (10) of the user (8) located within the vehicle with the set of cameras (2); and update (14) the 3D avatar data (13) of the user (8) based on the captured set of images (10).

[0141] [2] The vehicle of [1], wherein the circuitry (100) is further configured to: extract (11) a feature from the captured set of images (10) of the user (8); and update (14) the 3D avatar data (13) of the user (8) based on the extracted feature.

[0142] [3] The vehicle of [2], wherein the circuitry (100) is further configured to update (14) the 3D avatar data (13) of the user (8) based on comparing (12) the extracted feature (10) with a corresponding feature included within the 3D avatar data (13) of the user (8).

[0143] [4] The vehicle of [2] or [3], wherein the circuitry (100) is further configured to update (14) the 3D avatar data (13) of the user (8) based on a difference between the extracted feature (10) and the corresponding feature included within the 3D avatar data (13) of the user (8).

[0144] [5] The vehicle of any one of [2] to [4], wherein the circuitry (100) is further configured to update the 3D avatar data (13) of the user (8) by updating (14), with the extracted feature (10), the corresponding feature of the 3D avatar data (13) based on the difference, between the extracted feature (10) and the corresponding feature of the 3D avatar data (13), exceeding a predetermined or leant threshold of difference.

[0145] [6] The vehicle of any one of [2] to [5], wherein the circuitry (100) is further configured to update (14) the 3D avatar data (13) by leaving the 3D avatar data (13) unchanged if the difference between the extracted feature (10) and the corresponding feature of the 3D avatar data (13) is below the predetermined or leant threshold of difference.

[0146] [7] The vehicle of any one of [2] to [6], wherein the circuitry (100) is further configured to obtain the 3D avatar data (13) of the user (8) based on face recognition (16) of the user (8).

[0147] [8] The vehicle any one of [1] to [7], wherein the circuitry (100) is further configured to perform face recognition (16) based on the captured set of images (10).

[0148] [9] The vehicle of [2] to [8], wherein the circuitry (100) is further configured to perform face recognition (16) based on the extracted feature from the captured set of images (10).

[0149]

[0010] The vehicle of any one of [1] to [9], further comprising a microphone (7), wherein the circuitry (100) is further configured to update (14) a synthetic voice model of the 3D avatar data (13) of the user (8) based on an audio signal (17) of the user (8) captured with the microphone

[0150]

[0011] The vehicle of any one of [1] to

[0010] , wherein the set of images includes a sequence of images of the user.

[0151]

[0012] The vehicle of

[0011] , wherein the sequence of images is a video.

[0013] The vehicle of any one of [1] to

[0012] , wherein updating the 3D avatar data includes repeated updating of the 3D avatar data based on the sequence of images.

[0152]

[0014] The vehicle of any one of [1] to

[0013] , wherein the extracted feature includes multiple features and wherein updating the 3D avatar data is based on comparing the extracted features with corresponding features included within the 3D avatar data.

[0153]

[0015] The vehicle of any one of [1] to

[0014] , wherein the video captures a movement of the user and wherein the extracted features include a movement of the user.

[0154]

[0016] The vehicle of any one of [1] to

[0015] , wherein updating the 3D avatar data includes representing the movement of the user within the 3D avatar data.

[0155]

[0017] The vehicle of any one of [1] to

[0016] , wherein the circuitry if further configured to upload the 3D avatar data to a server for sharing the 3D avatar data with other users.

[0156]

[0018] A network node for updating 3D avatar data (13) of a user (8) of a vehicle (1, 44), the 3D avatar data (13) representing an avatar of the user (8), configured to: obtain 3D avatar data (13) of the user (8); obtain from the vehicle (1, 44) a set of images (10) of the user (8) located within the vehicle (1, 44) and captured by a set of cameras (2) of the vehicle (1, 44); and update the 3D avatar data (13) of the user (8) based on the acquired set of images (10).

[0157]

[0019] The network node of

[0013] may be further configured to save the set of images (10) and / or the updated 3D avatar data (13, 15) within a user profile (47) authorized and accessible by the user (8).

[0158]

[0020] A method for updating a 3D avatar data (13) of a user (8) of a vehicle (1, 44), the 3D avatar data (13) representing an avatar of the user, comprising: obtaining 3D avatar data (13) of the user (8); obtaining from the vehicle (1, 44) a set of images (10) of the user (8) located within the vehicle (1, 44) and captured by a set of cameras (2) of the vehicle (1, 44); and updating (14) the 3D avatar data (13) of the user (8) based on the acquired set of images (10).

[0159]

[0021] The method of

[0020] , further comprising extracting (11) a feature from the captured set of images (10) of the user (8), wherein updating (14) the 3D avatar data (13) of the user (8) is based on the extracted feature from the image (10).

[0022] The method of

[0020] or

[0021] , wherein updating the 3D avatar data (13) of the user (8) is based on comparing (12) the extracted feature (10) with a corresponding feature included within the 3D avatar data (13) of the user (8).

[0160]

[0023] The method of any one of

[0020] to

[0022] , wherein updating (14) the 3D avatar data (13) of the user (8) is based on a difference between the extracted feature (10) and the corresponding feature included within the 3D avatar data (13) of the user (8).

[0161]

[0024] The method of any one of

[0020] to

[0023] , wherein updating (14) the 3D avatar data (13) of the user (8) includes updating (14), with the extracted feature (10), the corresponding feature of the 3D avatar data (13) based on the difference, between the extracted feature (10) and the corresponding feature of the 3D avatar data (13), exceeding a predetermined or learnt threshold of difference.

[0162]

[0025] The method of any one of

[0020] to

[0024] , wherein updating (14) the 3D avatar data (13) includes leaving the 3D avatar data (13) unchanged if the difference between the extracted feature (10) and the corresponding feature of the 3D avatar data (13) is below the predetermined or learnt threshold of difference.

[0163]

[0026] The method of any one of

[0020] to

[0025] , wherein the method further comprises obtaining the 3D avatar data (13) of the user (8) based on face recognition (16) of the user (8).

[0164]

[0027] The method of any one of

[0020] to

[0026] , wherein the method further comprises performing face recognition (16) based on the captured set of images (10).

[0165]

[0028] The method of any one of

[0020] to

[0027] , wherein the method further comprises performing face recognition (16) based on the extracted feature from the captured set of images (10).

[0166]

[0029] The method of any one of

[0020] to

[0028] , wherein the method further comprises updating a synthetic voice model of the 3D avatar data (13) of the user (8) based on an audio signal (17) of the user (8) captured with a microphone (7) of the vehicle (1, 44) and obtained from the vehicle (1, 44).

[0167]

[0030] The method of any one of

[0020] to

[0029] , wherein the set of images includes a sequence of images of the user.

[0168]

[0031] The method of

[0030] , wherein the sequence of images is a video.

[0169]

[0032] The method of any one of

[0020] to

[0031] , wherein updating the 3D avatar data includes repeated updating of the 3D avatar data based on the sequence of images.

[0033] The method of any one of

[0020] to

[0032] , wherein the extracted feature includes multiple features and wherein updating the 3D avatar data is based on comparing the extracted features with corresponding features included within the 3D avatar data.

[0170]

[0034] The method of any one of

[0020] to

[0033] , wherein the video captures a movement of the user and wherein the extracted features include a movement of the user.

[0171]

[0035] The method of any one of

[0020] to

[0034] , wherein updating the 3D avatar data includes representing the movement of the user within the 3D avatar data.

[0172]

[0036] The method of any one of

[0020] to

[0035] , wherein the method further includes uploading the 3D avatar data to a server for sharing the 3D avatar data with other users.

[0037] A computer program comprising program code causing a computer to perform the method according to any one of

[0020] to

[0036] , when being carried out on a computer.

[0173]

[0038] A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor (101), causes the method according to any one of

[0020] to

[0036] to be performed.

Claims

CLAIMS1. A vehicle comprising: a set of cameras for capturing images of a user of the vehicle for updating a 3D avatar of the user; and circuitry for updating 3D avatar data, the 3D avatar data representing an avatar of a user, configured to: obtain 3D avatar data of the user; capture a set of images of the user located within the vehicle with the set of cameras; and update the 3D avatar data of the user based on the captured set of images.

2. The vehicle of claim 1, wherein the circuitry is further configured to: extract a feature from the captured set of images of the user; and update the 3D avatar data of the user based on the extracted feature.

3. The vehicle of claim 2, wherein the circuitry is further configured to update the 3D avatar data of the user based on comparing the extracted feature with a corresponding feature included within the 3D avatar data of the user.

4. The vehicle of claim 3, wherein the circuitry is further configured to update the 3D avatar data of the user based on a difference between the extracted feature and the corresponding feature included within the 3D avatar data of the user.

5. The vehicle of claim 4, wherein the circuitry is further configured to update the 3D avatar data of the user by updating, with the extracted feature, the corresponding feature of the 3D avatar data based on the difference, between the extracted feature and the corresponding feature of the 3D avatar data, exceeding a predetermined or learnt threshold of difference.

6. The vehicle of claim 5, wherein the circuitry is further configured to update the 3D avatar data by leaving the 3D avatar data unchanged if the difference between the extracted feature and the corresponding feature of the 3D avatar data is below the predetermined or learnt threshold of difference.

7. The vehicle of claim 2, wherein the circuitry is further configured to obtain the 3D avatar data of the user based on face recognition of the user.

8. The vehicle of claim 7, wherein the circuitry is further configured to perform face recognition based on the captured set of images.

9. The vehicle of claim 8, wherein the circuitry is further configured to perform face recognition based on the extracted feature from the captured set of images.

10. The vehicle of claim 1, further comprising a microphone, wherein the circuitry is further configured to update a synthetic voice model of the 3D avatar data of the user based on an audio signal of the user captured with the microphone.

11. The vehicle of claim 2, wherein the set of images includes a sequence of images of the user.

12. The vehicle of claim 11, wherein the sequence of images is a video.

13. The vehicle of claim 11, wherein updating the 3D avatar data includes repeated updating of the 3D avatar data based on the sequence of images.

14. The vehicle of claim 12, wherein the extracted feature includes multiple features and wherein updating the 3D avatar data is based on comparing the extracted features with corresponding features included within the 3D avatar data.

15. The vehicle of claim 12, wherein the video captures a movement of the user and wherein the extracted features include a movement of the user.

16. The vehicle of claim 15, wherein updating the 3D avatar data includes representing the movement of the user within the 3D avatar data.

17. The vehicle of claim 1, wherein the circuitry if further configured to upload the 3D avatar data to a server for sharing the 3D avatar data with other users.

18. A network node for updating 3D avatar data of a user of a vehicle, the 3D avatar data representing an avatar of the user, configured to: obtain 3D avatar data of the user; obtain from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and update the 3D avatar data of the user based on the acquired set of images.

19. The network node of claim 18 further configured to save the set of images and / or the updated 3D avatar data within a user profile authorized and accessible by the user.

20. A method for updating a 3D avatar data of a user of a vehicle, the 3D avatar data representing an avatar of the user, comprising: obtaining 3D avatar data of the user; obtaining from the vehicle a set of images of the user located within the vehicle and captured by a set of cameras of the vehicle; and updating the 3D avatar data of the user based on the acquired set of images.

21. The method of claim 20, further comprising extracting a feature from the captured set of images of the user, wherein updating the 3D avatar data of the user is based on the extracted feature.

22. The method of claim 21, wherein updating the 3D avatar data of the user is based on comparing the extracted feature with a corresponding feature included within the 3D avatar data of the user.

23. The method of claim 22, wherein updating the 3D avatar data of the user is based on a difference between the extracted feature and the corresponding feature included within the 3D avatar data of the user.

24. The method of claim 23, wherein updating the 3D avatar data of the user includes updating, with the extracted feature, the corresponding feature of the 3D avatar data based on the difference, between the extracted feature and the corresponding feature of the 3D avatar data, exceeding a predetermined or learnt threshold of difference.

25. The method of claim 24, wherein updating the 3D avatar data includes leaving the 3D avatar data unchanged if the difference, between the extracted feature and the corresponding feature of the 3D avatar data, is below the predetermined or learnt threshold of difference.

26. The method of claim 20, wherein the method further comprises obtaining the 3D avatar data of the user based on face recognition of the user.

27. The method of claim 26, wherein the method further comprises performing face recognition based on the captured set of images.

Citation Information

Patent Citations

  • Method and system for monitoring driving behaviors

    US20180239975A1

  • Shared environment for a remote user and vehicle occupants

    US20200066055A1

  • Method for Presenting Face In Video Call, Video Call Apparatus, and Vehicle

    US20220224860A1