Generating textual data

EP4802527A1Pending Publication Date: 2026-09-09TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023800786
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

Current technologies face challenges in generating large, high-quality datasets for deep learning models, particularly for health monitoring applications, due to limitations in image anonymization and the need for secure data collection.

Method used

A device and method are proposed that collect sensitive data, such as facial and body features, using a smart mirror equipped with cameras and sensors, and convert this data into textual descriptions using a multi-modal language model. The system ensures data security by deleting original images and sending only anonymized textual data to a server.

Benefits of technology

This approach allows for the secure collection and sharing of user data, providing a large, diverse dataset for training health monitoring models. The system improves the generalizability of models across different ethnicities and reduces the impact of varying illumination conditions, leading to more accurate health monitoring and diagnostic capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023080383_08052025_PF_FP_ABST
    Figure EP2023080383_08052025_PF_FP_ABST
Patent Text Reader

Abstract

According to an aspect, there is provided a computer-implemented method for operating a device (102). The method in the device (102) comprises: obtaining (401) one or more images (112, 114) of a user (110) of the device (102); using (403) a language model (140) to generate textual data (118; 120; 122) describing the user (110) shown in the one or more images (112, 114); and sending the generated textual data (118; 120; 122) to a server.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Generating Textual Data

[0002] Technical Field

[0003] This disclosure relates to generating textual data describing a user of a device.

[0004] Background

[0005] Early recognition of illness is critical for timely initiation of treatment. Several diseases can manifest physical symptoms on a patient's face, and astute clinicians might pick up on these subtle changes as part of their assessment.

[0006] For example, Cushing’s Syndrome occurs when a patient’s body is exposed to high levels of the hormone cortisol for a long period of time. Some of the symptoms - round, reddish, full face - often referred to as a “moon face” - and purple or pink stretch marks on the skin. Hypothyroidism can lead to facial puffiness, particularly around the eyes, and Parkinson's Disease has symptoms that include a mask-like expression and decreased frequency of blinking.

[0007] Clinical gestalt is a diagnostic tool that uses analysis of facial features of patients to help the detection of diseases. The value of clinical gestalt as a diagnostic tool has been studied for acute coronary syndrome, heart failure, pneumonia, Covid 19 and sepsis. The clinical gestalt provided and registered by doctors was comparable to clinical scores in “decision” and “exclusion” of patients presenting to the emergency department with certain symptoms. According to Third International Consensus for advocates of sepsis and septic shock (Sepsis-3), clinicians are recommended, in addition to the systemic inflammatory response syndrome (SIRS) criteria, to use clinical gestalt in screening, treating, and risk stratifying patients with infection (see “Deep Learning for Identification of Acute Illness and Facial Cues of Illness” by Castela Forte et al, Font. Med., 26 July 2021 Sec. Translation Medicine Volume 8 - 2021 , https: / / doi.org / 10.3389 / fmed.2021.661309).

[0008] Recently, deep learning models that were trained with a variety of clinical measurements were demonstrated to predict or detect early signs of diseases.

[0009] In “Deep Learning for Identification of Acute Illness and Facial Cues of Illness”, the authors developed a deep learning algorithm to distinguish between healthy and acutely ill individuals based on an analysis of their facial features. They used a synthetic dataset created by applying makeup to healthy individuals to simulate the facial features characteristic of acute illness.

[0010] The synthetic dataset was created on the photographs of twenty-six volunteers, made by a smartphone camera. Two photographs of each participant were selected and included in the study: one without makeup to represent the "healthy" control condition, and the other to represent the "acutely ill" condition. A standardised environment with a grey background and Light Emitting Diode (LED) light was used. White balance of the complete set of photographs was standardised by a professional photographer using a well-known photo editing tool.

[0011] So-called “smart mirror” devices can be used for monitoring a person’s facial features, skin condition, and other physical attributes. US Patent Application US 2018 / 253840 proposes a "Smart Mirror" as an Artificial Intelligence (Al) system that uses image analysis and pattern recognition technologies, which combines a visual display and a Machine Learning (ML) system. Among multiple features, the proposed mirror can monitor, with different sensors, a user’s health- related data such as vital signs, sleep quality, physical activity, body measurements, as well as stress or other mental health states based on user behaviour and interaction. Another patent application relating to a smart mirror (Korean Patent Application KR 20220052439) proposes a non-contact measurement of the user’s temperature. US 2020 / 342987 discloses a system for information exchange that comprises a mirror configured to capture the facial information of a user and display inference information about the user. An on-mirror computation device is configured to receive and process facial information of the user and produce inference information about the user. The mirror is coupled with a cloud-based back end and together they are arranged in a federated learning framework, with a tensor and / or model weights being sent between the mirror and the cloud-based server.

[0012] The performance of Al systems, particularly those based on deep learning methods, relies heavily on the quantity and quality of the available training data. In the context of image and video analysis, the richness of the dataset can significantly influence the accuracy and effectiveness of Al models. However, with the development and continuous improvement of machine learning and intelligent video technologies, there has been a parallel development in image anonymisation technologies.

[0013] In the early stages of image anonymisation, methods such as masking, obfuscation, or pixelation were typically used to hide sensitive information in images. However, they permanently alter the image, potentially losing important information. Typically, they require manual identification of the regions to be anonymised, which can be impractical for large datasets which would be suitable as testing or training data in machine learning.

[0014] The patent application WO 2023 / 060918 “Image anonymization method based on guidance of semantic and pose graphs” discloses a method to anonymise the whole image while maintaining the original semantic information of the image and the pose information of the person(s). In this method, human faces, bodies and background are completely replaced by abstract semantic maps. WO 2023 / 060918 suggests that the semantics, character posture and object motion information of the images can still provide a large amount of training data for Al models. Summary

[0015] There is still a conflict between the quantity and quality of the available training data for Al models, and the need for image anonymisation to enhance security. Despite the clinical gestalt or facial feature analysis increasingly used for building deep learning models, the development is slowed down by lack of large enough datasets and the quality of the images.

[0016] WO 2023 / 060918 doesn’t solve the above-mentioned conflict because it can provide usable training data only for the development and optimisation of humanoid detection, motion detection and other Al algorithm models that do not have high requirements on human faces.

[0017] “Deep Learning for Identification of Acute Illness and Facial Cues of Illness” has several limitations: a small training dataset; the simulated illness features might not adequately represent the spectrum of acute illness presentations, the study population was predominantly Caucasian, limiting the model's generalisability to other ethnicities; lack of real-world testing, and the real patient photographs would come with challenges like differing lighting conditions; ethical and security concerns.

[0018] According to US 2018 / 253840 the smart mirror collects a large amount of personal data; but this could lead to significant security concerns and potential misuse of data. US 2018 / 253840 does not focus on the facial analysis algorithm to distinguish healthy and acutely ill individuals based on facial features.

[0019] Therefore, a number of problems with generating suitable data from images of a subject remain.

[0020] Certain aspects of the disclosure and their embodiments may provide solutions to these or other challenges.

[0021] A device and method are proposed that are able to collect a large set of sensitive data suitable for testing or training Al models, while securing a user’s data. The sensitive data can include data relating to a user’s facial features, voice, and body features, that can be related to, but is not restricted to, health and illness monitoring.

[0022] One embodiment of the proposed device is a surface (e.g. a reflective surface such as a mirror), that can be used in the home, in public or in a healthcare setting. The device can be portable or wall-mounted / wall-mountable. The device (e.g. surface, mirror, etc.) can comprise one or more camera(s) and optionally other sensors, some form of communication technology such as Wi-Fi or Bluetooth, an illumination / lighting system for illuminating the user, and one or more ML engine(s) / processor(s). The camera(s) can obtain image(s) of the user in front of the mirror / device. The device uses the camera(s) to gather personal data, and extracts textual data from the images such that individuals cannot be recognised from the textual data output by the device.

[0023] The proposed method can use one or more ML / AI algorithm(s) to analyse images obtained by the one or more cameras and optional sensors. The ML / AI algorithms can be or include object detection models, deep learning models, language models, multi-modal language models, Large Language Models (LLMs), and / or multi-modal Large Language Models (LLMs). Computer vision algorithms can be used to identify key facial features, such as the eyes, nose, mouth, skin area, eyebrows, mouth, and eyes corner, and further finer details like irises, skin tone, variations in skin texture, pigmentation, circles under the eyes, etc. The device uses an embedded image-to-text Al algorithm / system to translate the collected data to textual content. Audio captured by the camera(s) or other microphones can be recorded / stored, particularly audio relating to coughing, breathing, etc., and a multi-modal LLM can translate information represented in that audio capture, and the images, to textual data.

[0024] Thus, the input to the system / method can be a set of facial and / or body images of one or more users, recorded / obtained by the camera(s) that are embedded in, or otherwise part of, the device. The output of the system / method can be a detailed translation of the image(s) into textual description content including or related to the user’s face, body, skin, eyes, hair, etc.

[0025] The original image(s) from the camera(s) will be permanently destroyed / deleted at the device after the textual content is created, processed, and saved, and without the image(s) being sent to the server. This will enable the security of the user’s data to be fully protected. The textual content can then be sent to a server, for example a server that manages a global Al training database to be used in training a ML / AI model to detect health conditions.

[0026] Certain embodiments may provide one or more of the following technical advantage(s). The techniques described herein allow a unique data set to be gathered in a secure way. Sensitive data can be collected and shared about the user while respecting the security of the user’s data. The data set provides real-world data for health monitoring. Embodiments provide that illness features will adequately represent the broad spectrum of acute illness, and there will not be any limitation on a subsequently trained model's generalisability or applicability to other ethnicities, which means that the performance and accuracy of any monitoring or diagnostic system working based on such data will be significantly improved. Certain embodiments can reduce the effect of illumination in the data. The techniques, and / or a system using a model trained based on the collected data set, can also enable the offload of diagnostic decisions from hospitals and / or experienced clinicians. For certain types of illness, it is possible to detect early warning signs without needing to have regular hospital visits.

[0027] According to a first aspect, there is provided a computer-implemented method for operating a device. The method in the device comprises obtaining one or more images of a user of the device; using a language model to generate textual data describing the user shown in the one or more images; and sending the generated textual data to a server.

[0028] According to a second aspect, there is provided a device that is configured to obtain one or more images of a user of the device; use a language model to generate textual data describing the user shown in the one or more images; and send the generated textual data to a server.

[0029] According to a third aspect, there is provided a device that comprises a processor and a memory, said memory containing instructions executable by said processor whereby said device is operative to obtain one or more images of a user of the device; use a language model to generate textual data describing the user shown in the one or more images; and send the generated textual data to a server.

[0030] According to a fourth aspect, there is provided a computer program product comprising a computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform the method according to the first aspect or any embodiment thereof.

[0031] According to a fifth aspect, there is provided a health monitoring system that comprises: a device according to the second or third aspects or any embodiments thereof; and a server configured or operative to receive the textual data from the device.

[0032] Brief Description of the Drawings

[0033] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings, in which:

[0034] Fig. 1 illustrates a device in the form a mirror according to an embodiment, and examples of inputs and outputs of the image-to-text conversion according to embodiments of the techniques described herein;

[0035] Fig. 2 is a simplified block diagram of a device in accordance with one or more embodiments;

[0036] Fig. 3 is a block diagram illustrating an exemplary textual data generator; and

[0037] Fig. 4 is a flow chart illustrating a method in accordance with some embodiments.

[0038] Detailed Description

[0039] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.

[0040] Fig. 1 illustrates a device 102 in the form of a surface such as a mirror according to an embodiment, along with examples of inputs 104 and outputs 106 of a textual data generator 108. The textual data generator 108 is, or comprises, a language model, such as a Large Language Model (LLM), and in particular embodiments is a multi-modal language model (e.g. a multi-modal LLM) that is able to generate textual data from non-textual inputs, such as video, images, audio signals, etc., or a combination thereof. As discussed in more detail below with reference to Fig. 3, the textual data generator 108 can include modules or functions that are capable of processing non-textual inputs, e.g. to encode and / or ‘perceive’ the content of the non-textual inputs, to enable the language model to generate textual data for those inputs. Alternatively, for example for a multi-modal language model, these modules or functions can be integrated or otherwise part of the language model. In the following explanation, the terms “language model” and “LLM” are generally used interchangeably. The textual data generated by the language model is sent to a server. The server may use the textual data for any suitable purpose, such as for training and / or testing a ML model, establishing a rule-based system or expert system, and / or for monitoring and / or visualisation directly by a healthcare professional, for example. The ML model, rule-based system or expert system may be for any particular purpose, e.g. for diagnosing a health condition, disease, illness, etc., but it will be appreciated that the purpose of the ML model or system has no bearing on the method of collection and generation of the textual data as described herein.

[0041] As noted, the language model can be a LLM. Currently-available models are often several gigabytes in size. For instance, OpenAI's GPT-3 model with 175 billion parameters is around 350 GB. However, lightweight LLMs with very good performance have recently emerged, that can run inference using hardware that is cheap and not too bulky (<600 USD). One example of a lightweight LLM is Stanford University's Alpaca (7B), where with only 52K instruction-following demonstrations they were close to OpenAI’s models.

[0042] For the techniques described herein, a multimodal GPT-based model with <100K images and corresponding reports from medical experts may be needed. The resulting network could then be distilled into a smaller model. GPT-4 is multi-modal and has 1.76 trillion parameters. If a distilled network is assumed that has a similar proportion to that of OpenAI’s Davinci vs Alpaca, a network size of roughly 70B may be achieved. There are already versions such MiniGPT-4, as described in “Minigpt-4: Enhancing Vision-Language Understanding with Advanced Large Language Models” by Zhu et al., https: / / arxiv.org / pdf / 2304.10592.pdf. Use of guard rail models can be considered for medical purposes, to align with predefined rules and constraints, for example to avoid non-medical advice.

[0043] The device 102 can be a type of smart mirror, which can be used in a person’s home, or placed in hospitals / clinics and / or other public spaces.

[0044] While the embodiments are described herein with reference to the user being a person, it will be appreciated that the user 110 may instead be an animal, such as a pet, for example a dog or a cat.

[0045] Fig. 2 is a simplified block diagram of the device 102 in accordance with one or more embodiments. The device 102 comprises processing circuitry (or logic) 201. It will be appreciated that the device 102 may comprise one or more virtual machines running different software and / or processes. The processing circuitry 201 controls the operation of the device 102 to implement the methods described herein, and in particular uses the language model / LLM to generate textual data for the user 110 from the input images 112 / 114. The processing circuitry 201 can comprise one or more processors, processing units, multi-core processors or modules that are configured or programmed to control the device 102 in the manner described herein. In particular implementations, the processing circuitry 201 can comprise a plurality of software and / or hardware modules that are each configured to perform, or are for performing, individual or multiple steps of the method described herein in relation to the device 102.

[0046] The device 102 also comprises a communications interface 202. The communications interface 202 is for use in enabling communications with other devices, user smartphones, computers, remote servers, etc. The communications interface 202 can be configured to transmit to and / or receive from other devices, smartphones, computers or servers requests, acknowledgements, information, data, signals, or similar. For example, the communications interface 202 can be used to transmit the text output 106 to a server that handles and / or stores the data for use in training a ML / AI model. The communications interface 202 can use any suitable communication technology / ies, for example WiFi, Bluetooth, etc.

[0047] The device 102 may comprise a memory 203. In some embodiments, the memory 203 can be configured to store program code that can be executed by the processing circuitry 201 to perform the methods described herein in relation to the device 102. Alternatively or in addition, the memory 203 can be configured to store any input data 104 (e.g. images, image data, sensor measurements), data required by the language model / LLM 108 (including the language model / LLM itself), and any output data 106 (e.g. textual description of any part(s) of the input data 104). The processing circuitry 201 may be configured to control the memory 203 to store such information therein.

[0048] The device 102 comprises, or is associated with, one or more image sensor(s) 207 (e.g. camera(s)) that are used to obtain / generate one or more images of a user 110 that is in front of the mirror device 102. The one or more image sensor(s) 207 may be integral to the device 102, or they may be separate from the device 102 and provide the obtained images to the device 102 / processing circuitry 201 for processing. The one or more image sensor(s) 207 may record a video stream comprising a series of images. The one or more image sensor(s) 207 can include a Red Green Blue (RGB) image sensor 207 for obtaining images based on visible light. In addition or alternatively, the one or more image sensor(s) 207 can include an Infra-Red (IR) image sensor 207, for obtaining images using IR light. The one or more image sensor(s) 207 should provide images of a sufficient quality (e.g. high resolution, high frame rate, focus, etc.) to enable effective and accurate facial feature extraction. Some embodiments can include multiple image sensors 207 arranged to obtain images of the user 110 from multiple angles, enabling three-dimensional (3D) information / images of the user 110 to be obtained.

[0049] Fig. 2 shows one or more light source(s) 209 that are used to illuminate the user 110 when image(s) are being obtained by the image sensor(s) 207. The presence and / or use of the one or more light source(s) 209 is optional. Preferably the one or more light source(s) 209 are arranged or controllable to evenly illuminate the user 110, e.g. to minimise shadows, and allow colour- accurate / colour-consistent images to be obtained. Thus, the light source(s) 209 should generate suitable light, and can be, for example, natural daylight lamps, LED lights (e.g. in the form of light strip lights located along the sides, top, and / or bottom edges of the mirror 102), vanity lights mounted on either side of the mirror 102 at eye level to eliminate shadows on the face, backlighting that provides diffuse illumination from behind the mirror surface for soft and uniform light distribution, etc. The light source(s) 209 may be part of the device 102, or they may be separate from the device 102. In either case, the processing circuitry 201 of the device 102 may be able to control the operation of the light source(s) 209.

[0050] Fig. 2 shows one or more optional ‘other’ sensors 211 that are used to obtain measurements of one or more characteristics of the user 110 and / or the environment around the user 110. Any particular sensor 211 may be part of the device 102, or may be separate from the device 102. In either case, the processing circuitry 201 of the device 102 is able to receive measurements or a measurement signal from the sensor(s) 211 .

[0051] Another / further sensor 211 can be a microphone that is used to record sounds associated with the user 110, such the sound of a user’s cough and / or breathing. Microphones can also be used to capture other audio input, allowing users 110 to interact with the mirror 102 using voice commands, etc.

[0052] Another / further sensor 211 can be a thermometer, e.g. an IR thermometer, that is used to measure the body temperature of the user 110. The thermometer may be able to measure the temperature remotely (e.g. without the user 110 having to touch the device 102), or may require the user 110 to touch a temperature sensing part of the device 102 with their hand or other body part.

[0053] Another / further sensor 211 can be a weight sensor, e.g. weighing scales, smart scales, etc., that is used to measure the weight of the user. Another / further sensor 211 can be a blood pressure monitor for measuring the blood pressure of the user.

[0054] Another / further sensor 211 can be a sensor that can obtain heart rate information such as heart rate and / or heart rate variability of the user.

[0055] Another / further sensor 211 can be a sensorthat can measure the breathing rate of the user.

[0056] Another / further sensor 211 can be one or more light sensors for measuring the intensity of light in the environment around the device 102 (which is also referred to as the ambient light or ambient light level). A measurement of the ambient light can be used by the device 102 to: adjust the brightness of the light generated by the light source(s) 209 to provide an appropriate illumination level; adjust the light sensitivity of the image sensor(s) 207; and / or adjust the brightness of any display components of the device 102.

[0057] Another / further sensor 211 can be one or more proximity sensors that can detect the presence of a user 110 in front of the mirror 102 and activate the mirror display and / or adjust its brightness when the user 110 approaches or moves away from the mirror 102.

[0058] Another / further sensor 211 can be one or more motion sensors, such as IR and / or ultrasonic sensors that can detect movement in the vicinity of the mirror 102. The measurement signals from the motion sensor(s) can be used to trigger specific actions, such as turning on the display of the mirror 102, or activating certain features when motion is detected.

[0059] Another / further sensor 211 can be one or more touch sensors which allows users to interact with the mirror’s control interface and / or control various functions.

[0060] As noted above, the device 102 can be or comprise, a surface. The surface can be a surface that the user may pass by when performing their day to day routine, such as shop windows, lift walls, public toilets, bathrooms, etc. The surface may be a reflective surface, e.g. a mirror. In some embodiments, the mirror comprises a traditional mirrored surface for reflecting light. The mirror may also include a separate display screen for displaying information to the user 110, or display elements may be integrated into the reflective surface. In other embodiments, instead of a reflective surface the mirror can comprise a display screen that is used to display a video stream of the user 110 that is obtained by the one or more image sensor(s) 207.

[0061] The device 102 can be part of a system, such as a health monitoring system, that also comprises a server for receiving textual data from the device 102. As noted above, the server may use the textual data for any suitable purpose, such as for training and / or testing a ML model, establishing a rule-based system or expert system, and / or for monitoring and / or visualisation directly by a healthcare professional. The health monitoring system can generate health monitoring outputs using the ML model. For example, depending on the type of ML model used, a health monitoring output can be any of: an indication of a health condition / illness / disease of the user 110; an indication of whether the user 110 should seek medical advice from a healthcare professional; advice on exercises / activities / diet suitable for any detected health conditions of the user 110; etc.

[0062] Fig. 1 shows several exemplary inputs 104 relating to the user 110 obtained by the device 102. In this example, the inputs 104 include a first image or first set of images 112 that are obtained at a first time using an image sensor 207. The first image or first set of images 112 are of a user 110 that is healthy, or relatively healthy. The inputs 104 also show a second image or second set of images 114 that are obtained using the image sensor 207. The second image or second set of images 114 are of a user 110 that is unhealthy / ill. The inputs 104 also include an audio / sound measurement obtained by a microphone. The audio / sound measurement 116 may have been obtained at the same or a similar time to the second image or second set of images 114.

[0063] The inputs 104 are provided to the LLM 108, optionally after one or more of the inputs 104 have been pre-processed as described below, and the LLM 108 generates textual data describing the user 110. The generated textual data may describe the face of the user 110, the voice / breathing of the user and / or may relate to health aspects of the user.

[0064] Fig. 1 shows three exemplary textual outputs 106 for each of the inputs 104. If the first image 112 was input, the LLM may generate textual data 118 comprising the text: “skin have natural glow; lips are free of significant discolouration; eyes are bright and clear; nose free of persistent redness; mouth normal symmetry”. If the second image 114 was input, the LLM may generate textual data 120 comprising the text: “paler skin tone; pale lips; redness around the eyes; sunken eyes; redness around the ala of the nose; droopy mouth”. For the audio signal 116, the LLM generates textual data 122 comprising text selected from: “dry / wet / “whooping” cough; hoarse, weak, or nasal voice; rapid shallow breathing; sneezing and sniffing”. It will be appreciated that where the audio signal 116 and one of the images / sets of images 112 / 114 is input to the LLM, the LLM can provide a textual data output based on both of those types of input 104.

[0065] The images 112 / 114 input to the LLM 108 can be static or dynamic (e.g. a video stream / sequence of images). The LLM may be able to directly receive the image(s) 112 / 114 and generate the textual data from the image(s) 112 / 114. Alternatively, the image(s) 112 / 114 may be pre-processed, and the output of the pre-processing provided to the LLM for use in generating the textual data. In this case, the output of the pre-processing and the original image(s) 112 / 114 may be provided to the LLM, and both used to generate the textual data. For example, as described with reference to Fig. 3 below, an encoding module and / or perceiving module may be used to encode the image(s) and / or perceive the content of the image(s), with the LLM generating the textual data based on the output of the encoding module and / or perceiving module. Alternatively, the pre-processing of the image(s) 112 / 114 may comprise any of:

[0066] • An object detection algorithm that is used to identify the objects (i.e. body part(s)) within the image(s). The object detection algorithm may provide a bounding box around the detected object(s).

[0067] • A semantic segmentation algorithm that can be used to segment the image into pixel level information.

[0068] • A facial classification algorithm, such as a Haar cascade facial classifier, which can be used to identify the entire face region in an image.

[0069] • A facial feature classifier, e.g. a facial landmark detector, that is used to identify facial features, such as corner of the eye, tip of the nose, corner of the mouth, etc. Examples of suitable facial feature classifiers include the FaceNet or VGGFace models. This type of classifier / model is also known as “Facial Landmark Detection”, with one example being FPENet. The paper “Recombinator networks: Learning coarse-to-fine feature aggregation” by Honari, S., Yosinski, J., Vincent, P., & Pal, C., in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 5743-5752) 2016 provides an example facial landmark detection model.

[0070] In some embodiments, the image(s) 112 / 114 can be processed to extract biometric information for the user. The biometric information can relate to facial features or face shape. The biometric information may be used as a unique identifier for the user in biometric systems.

[0071] In some embodiments, the generated textual data should be associated with a particular user 110 of the device 102. For example, multiple users 110 may be able to use the device 102 (e.g. different members of a family), and it is useful to ensure that the generated textual data is associated with previously-generated textual data for the same user 110. One approach is to create a face database or gallery containing facial templates for each of the different users 110 that can use the device 102. This database can serve as a reference for the processing circuitry 201 to compare newly acquired image(s) to and find a matching face. Thus, when the generated textual content is shared (e.g. sent to a server where it is used as training data for a ML model), the textual data is only associated with the previous textual data for that user 110, and it is not mixed in with the textual data for the other users.

[0072] In some embodiments it is helpful if the textual data comprises or is accompanied by population profile information for the user, as this can provide useful context for the textual data. The population profile information is information that would not allow the user to be identified, but provides general information about the user’s demographics or background. For example, the population profile information can comprise any of information about the user’s age, gender, ethnicity, etc. The population profile information for a user can provide an indication of aspects of their visual appearance. Factors such as age, gender, height and ethnicity can play a significant role in how a person is perceived visually. Providing information about these basic characteristics in or with the textual data enables population characteristics to be accounted for when training a ML model, rule-based system or expert system using the textual data and / or to enable the ML model, rule-based system or expert system to provide an output that is more relevant to a patient’s population profile.

[0073] As noted above, to generate textual data, LLM is used. The LLM may be a multi-modal LLM, meaning that it can process image(s) and one or more other types of input data (e.g. an audio signal) process them together to generate the textual data. The LLM may evaluate the different input modalities separately (e.g. resulting in textual data comprising textual data outputs 118, 120 and 122 in Fig. 1). Alternatively, the LLM may evaluate the different input modalities together and generate textual data from correlations and interactions between the input modalities.

[0074] For the LLM or multi-modal LLM, the input data, e.g. images, audio signals, sensor measurements, sensor measurement signals, etc. may need to be converted or encoded into a format that the LLM can understand. For images, techniques like Convolutional Neural Networks (CNNs) can be used to extract visual features from the images. Similarly, for audio signals, methods like a CNN adapted for sound data, a recurrent network such as Long short-term memory (LSTM), spectrogram representation or audio embeddings can be employed to represent the audio data. Each encoder / model can extract features / feature vectors from the respective modalities (images, audio, etc.). In some cases, the extracted visual and audio features (and features from other modalities, if present) can be combined (e.g. concatenated or merged) or fused together to create a unified representation / vector that represents all modalities. This merged feature vector can be further processed using fully connected layers. The fused representation can then be input into the LLM. The LLM processes the input representation(s) and determines the textual descriptions and phrases corresponding to the visual and audio features in the fused representation.

[0075] The LLM can be based on a commercially-available LLM, with some modifications (e.g. so- called handrails or guardrails) and / or additional training to enable the LLM to generate appropriate textual data, e.g. relating to health aspects of users. For the training, custom loss functions may be used, particularly if the training dataset is imbalanced. The LLM can be trained to map inputs (e.g. images, sounds) to outputs (e.g. textual descriptions) using a labelled dataset. In the case of supervised learning, the ground truth should ideally be supervised by healthcare professionals with relevant expertise. It should be appreciated that the training data used to train the LLM to generate the appropriate textual data does not have to be generated in the manner described in this disclosure.

[0076] It will be appreciated that the quality or performance of the LLM in describing facial details or shades of colour / skin tone, etc. is largely dependent on its training data and the algorithmic techniques employed.

[0077] The processing circuitry 201 / memory 203 can store the weights and / or other parameters of the LLM. The inference (i.e. evaluation of the images and / or other input modalities) by the LLM is performed on the device 102. The textual data output by the LLM is stored in the device 102 and / or sent to a remote server. The original image(s) and / or other sensor measurements used to generate the textual data are deleted.

[0078] The block diagram in Fig. 3 shows an exemplary implementation of a textual data generator 108 that is able to process multi-modal inputs to generate textual data. The textual data generator 108 provides a unified multi-modal model that leverages a language model 140 (e.g.an LLM) as a catalyst to interpret and correlate different sensor modalities.

[0079] Fig. 3 shows multi-modal inputs 104 to the textual data generator 108 that include raw images 112, sound recordings 116, and optionally additional sensor inputs 150, 152, as described above.

[0080] The textual data generator 108 may comprise a set of encoding modules that are used to encode respective input data 104. That is, an image encoding module 154 encodes the input images 112, a sound encoding module 156 encodes the input sound 116, and sensor encoding blocks 158, 160 respectively encode the sensor inputs 150, 152. In general, the encoding modules can determine an embedding or latent space vector from the input data.

[0081] The textual data generator 108 may also comprise a set of perceiving modules that are used to perceive respective input data. That is, an image perceiving module 162 can perceive the encoded images 112, a sound perceiving module 164 can perceive the encoded sound 116, and sensor perceiving blocks 166, 168 can respectively perceive the encoded sensor inputs 150, 152. In general, the perceiving modules can project embeddings of the input to a semantic space of the language model.

[0082] The output of the perceiving modules 162, 164, 166, 168 are provided to the language model 140 which generates a suitable language response 106.

[0083] As shown in Fig. 3, the different input modalities are encoded, respectively, where the encoders 154, 156, 158, 160 are typically pretrained. Current state-of-the-art models for vision encoders (and video) include, e.g., EVA-ViT-G (as described in “EVA: Exploring the Limits of Masked Visual Representation Learning at Scale”, by Fang et al., CVPR 2023) and for audio BEAT (as described in “BEATs: Audio Pre-Training with Acoustic Tokenizers”, by Chen et al. ICML, 2023). Other modalities will require different encoder modules. Regardless of the encoder (and modality), the encoder transforms the raw input into a variable-length embedding. Typically, this can be performed using a deep neural network. Many of the current state-of-the-art networks for various detection and recognition tasks build upon previous work. One example is ResNet (as discussed in “Deep residual learning for image recognition” by He et al., CVPR 2016) which has been used as a backbone (meaning the feature extracting part of the network) for localisation and semantic segmentation networks, e.g., Mask R-CNN (as discussed in “Mask R-CNN" by He et al., CVPR 2017). It should be noted that the backbone is not the final output of the network, but rather an intermediate output from a layer that encodes relevant information. This is then passed to the network head for the specific task, e.g. bounding-box recognition (classification and regression) and mask prediction, as in the example of Mask R-CNN. The paper “Proper Reuse of Image Classification Features Improves Object Detection” by Vasconcelos et al., CVPR 2022 shows that freezing the backbone and only re-training the head can improve many different detection models.

[0084] The perceiver modules 162, 164, 166, 168 bridge the encoders 154, 156, 158, 160 and the language model 140. Such perceiver modules have been successfully used in, for example, Flamingo (as described in “Flamingo: a Visual Language Model for Few-Shot Learning” by Alayrac et al., NeurlPS 2022) and BLIP-2 (as described in “BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models” by Li et al., 2023), where they are referred to as Perceiver Resampler and Q-Former (Querying Transformer), respectively. These are modal-specific; however, the purpose of the perceiver module 162, 164, 166, 168 is to project the embeddings from the different sensor encoders 154, 156, 158, 160 to the semantic space of the language model 140, i.e. to convert non-text-based data to a language representation. The output is typically a lot smaller than the size of the frozen features of the encoder, and hence creates a bottleneck architecture to force queries to extract sensory information that is most relevant to the text.

[0085] In order to align the dimensions of these quasi-linguistic embeddings and the embedding dimension of the language model 140, a fully-connected layer can be used to project the output query embeddings into the same dimension as the text embedding of the language model 140. The projection can be linear, e.g., CLIP (as described in “Learning Transferable Visual Models From Natural Language Supervision” by Radford et al., ICML 2021) or non-linear, e.g., AMDIM (as described in “Learning representations by maximizing mutual information across views” by Bachman et al. In Advances in Neural Information Processing Systems, pp. 15535-15545, 2019) and SimCLR (as described in “A simple framework for contrastive learning of visual representations" by Chen et al. ICML 2020). The projected query embeddings can then be prepended to the input text embeddings, similarly to BLIP-2. Similar concepts are discussed in “X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages” by Chen et al., 2023 (https: / / arxiv.org / abs / 2305.04160) and “ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst” by Zhao et al., 2023 (https: / / arxiv.org / pdf / 2305.16103.pdf).

[0086] Further details of the training of the (large) language model 140 for medical purposes is provided below. There are many variations in how to train the network. Due to the high model complexity, it is common to freeze parts of the model and only fine-tune certain parts of the network, for example as described in the above-referenced papers by Alayrac et al. (2022), Li et al. (2023), Chen et al. (2023) and Zhao et al. (2023), and also the papers “Grounding Language Models to Images for Multimodal Inputs and Outputs” by Koh et al., ICML 2023, and “LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation” by Lee et al., 2023 (https: / / arxiv.org / abs / 2305.11490). The perceivers 162, 164, 166, 168 and their learnable query tokens can be trained while the encoders 154, 156, 158, 160 and the language model 140 are frozen during the entire training process, for example due to limited computational resources.

[0087] In order to add image tokens to a language model 140 that is pre-trained only with text, the extra sensory (image, sound, etc.) tokens should be put into the same embedding space as the text tokens. This has been treated differently in different frameworks. For example, DALL-E (as described in “Zero-Shot Text-to-lmage Generation” by Ramesh et al., ICML 2021) treats images as sequences of discrete tokens using a transformer-based text-to-image model. Flamingo referenced above proposes a visual language model for text generation. The paper by Koh et al. proposes a way to ground pre-trained text-only language models to the visual domain, enabling them to process arbitrarily interleaved image-and-text data, and generate text interleaved with retrieved images.

[0088] The above-referenced paper by Lee et al. shows that clinical information-preserving chest X-rays (CXR) tokenization leads to performance improvement in both the report-to-CXR task and the CXR-to-report task, by utilizing the work by Koh et al. and applying fine-tuning techniques. These are standard fine-tuning processes for LLMs, and no changes are made other than the expansion of the token embedding table, and as such the template used by Alpaca (as discussed in “Stanford Alpaca: An Instruction-following LLaMA model” by Taori et al., 2023, https: / / github.com / tatsu-lab / stanford_alpaca) could be used. In this case, the textual data generation as proposed in this disclosure can consist of tokenized sensory data input to the language model 140 and the output textual data is a corresponding medical expert report based on said data input.

[0089] The mirror / device 102 can have additional functionality to that described above. For example, camera and display functionality of the mirror can be used to connect patients with healthcare professionals for the purpose of remote consultations. The LLM can provide new textual data for the user 110 during this consultation, and / or the sensor(s) 211 in or associated with the mirror / device 102 can be used to provide new or live health data for the user 110.

[0090] While the above description indicates that the textual data is sent to a remote server, for example for use as training data for a ML model, rules-based system or expert system, it will be appreciated that the textual data may not be sent as soon as it is generated. Instead, textual data may be generated at a number of different instances (e.g. over the course of several days or weeks) and stored in the device 102. The stored textual data can be sent to the remote server at a suitable time, e.g. when a diagnosis for the user 110 has been made, if it is detected that the health of the user 110 has degraded, once a sufficient amount of textual data has been stored, periodically, etc.

[0091] Once textual data for a user 110 has been sent to the remote server, it is possible that further textual data relating to the same user 110 is generated from further image(s) etc. obtained on subsequent days or weeks. To improve the use of the textual data at the remote server, it may be useful for the new textual data to be associated with the textual data for the same user 110 previously shared with the server.

[0092] In some cases this need for an association can be avoided by the device 102 storing and sending a full set of textual data for the user 110 each time that textual data is sent to the server. That is, the new textual data can be sent to the server, along with previous textual data that has already been sent to the server. In some cases the set of textual data sent to the server can cover a period of time in which the user 110 went from healthy, to ill, and then recovered.

[0093] If the association between the new textual data and previously-sent textual data is required at the server, the device 102 and server can make use of a unique private key, e.g. a Secure Hash Algorithm-ID (SHA-ID), to securely transfer textual data. The use of this ID, together with the anonymised textual data, would be considered difficult to trace back to the user 110. A similar solution could be used if the user 110 wants to share new textual data and / or accumulated textual data from the device 102 with a medical professional.

[0094] For a user 110 with early signs of a disease, as a user’s sickness is developing, notes (e.g. contents of a log file in addition to the generated textual data) can be logged on the device 102. At some point the symptoms are clear and the patient may be given a diagnosis by a healthcare professional, and this diagnosis can be added to the log file. The log file may be sent to the server along with, or as part of, the textual data.

[0095] In some embodiments, federated learning (FL) techniques can be applied to the training or fine-tuning (updating) of the textual data generator 108, or to the training or fine-tuning of elements of the textual data generator 108 (e.g. the encoders or perceivers). Federated learning is a decentralised machine learning approach where a model (e.g. a language model) is trained across multiple edge devices (in this case devices / mirrors 102) or decentralised servers. By using data from multiple clients, a language model can learn from a diverse range of examples, enhancing its overall understanding and performance, which leads to a more powerful and versatile language models / LLMs. Applying FL techniques to various aspects of language models, including (pre-)training and fine-tuning, offers numerous advantages.

[0096] The paper “Federated Large Language Model: A Position Paper” by Chen et al., 2023 (https: / / arxiv.org / pdf / 2307.08925.pdf) discusses the concept of a federated LLM that integrates FL methods and comprises federated LLM pre-training and federated LLM fine-tuning. Compared with traditional LLM training methods relying exclusively on centralised public datasets, federated learning LLM pre-training incorporates a combination of both centralised public data and decentralised private data sources. This integration of diverse data sources serves the purpose of enhancing model generalisation, enables access to a broader spectrum of medical knowledge, and facilitates future model scalability and expansion in terms of both size and scope.

[0097] After pre-training, LLMs can often be fine-tuned on specific tasks or domains. Thus, federated language model fine-tuning can improve generalisation across tasks and domains. The conventional fine-tuning approach relies on individual institutions using their proprietary datasets, often limiting collaboration and facing challenges related to data availability and generalisation. In contrast, federated language model (e.g. LLM) fine-tuning promotes collaboration among multiple institutions, considers specific task requirements, and employs data from various clients for joint multi-task training. Federated learning and LLMs are not mutually exclusive, but reinforce and complement each other.

[0098] A FL-based approach can be used to fine-tune a deployed language model (i.e. language model 140 in textual data generator 108). In this case, instead of needing to send raw data (e.g. images, audio, etc.) to a central server for training / updating the language model / textual data generator 108, updates to the textual data generator 108 are computed locally (i.e. in a mirror 102), and only these updates are shared and aggregated at a central server to improve the global language model / textual data generator 108. This minimises the need for transferring large volumes of data to a central server, thus reducing bandwidth usage and associated costs.

[0099] While federated learning provides several benefits, there can be concerns over user privacy in other FL use-cases. However, in the present case, the data is impersonal text strings (textual data) stored in the mirror 102, and hence a user’s privacy is preserved even if that data is compromised. In particular, each mirror / device 102 can independently compute updates to the language model 140 locally using their datasets. These updates represent the incremental changes needed to improve the global language model. The updates are then transmitted to a central server (which may be the same or a different server to the server that receives the textual data 106 and uses it for some further purpose). The central server aggregates these updates to create an improved global language model 140. This global language model 140 encapsulates the collective knowledge and insights from all mirrors / devices 102 that are participating in the federated learning. After the central server aggregates the updates and creates an improved global language model 140, this updated global language model 140 is sent back to the participating mirrors / devices 102. Each mirror / device 102 then receives the updated global language model 140 and further fine-tunes its local model 140 based on the improvements to the global model.

[0100] Recent research suggests tackling one of the main security concerns about data privacy by fine-tuning pre-trained language models on a client with minimal communication to the central server (e.g. just the sharing of the model weights). This research is found in “FedPETuning: When Federated Learning Meets the Parameter-Efficient Tuning Methods of Pre-trained Language Models” by Zhang et al., Findings of the Association for Computational Linguistics: ACL, 2023 (https: / / arxiv.org / pdf / 2212.10025.pdf).

[0101] The device / mirror 102 will primarily rely on private data; however, this does not exclude public datasets. The paper “Can Public Large Language Models Help Private Cross-device Federated Learning?” by Wang et al. 2023 (https: / / arxiv.org / pdf / 2305.12132.pdf) proposes a way to improve on-device LLM training using public tokenizers with distillation before private crossdevice federated learning. This approach provides significant improvements, and demonstrated strong private learning accuracy. This approach can be used to improve private federated learning with public LLMs, which could further to improve the performance of the techniques described herein. The paper “FedJudge: Federated legal large language model” by Yue et al., 2023 (https: / / arxiv.org / pdf / 2309.08173.pdf) demonstrate that it is feasible to consider a federated LLM approach for a specific task (here they use legal intelligence as a use case).

[0102] Fig. 4 is a flow chart illustrating a method of operating a device / mirror 102 according to various embodiments. The device / mirror 102 may perform the method in response to executing suitably formulated computer readable code. The computer readable code may be embodied or stored on a computer readable medium, such as a memory chip, optical disc, or other storage medium. The computer readable medium may be part of a computer program product.

[0103] In step 401 , the device 102 obtains one or more images of a user 110 of the device 102. Step 401 can comprise the device 102 controlling an image sensor 207 to generate the one or more images. Alternatively, step 401 can comprise the device 102 / processing circuitry 201 receiving the one or more images from a separate image sensor 207, or retrieving previously- generated and stored images from a storage location (e.g. the memory 203).

[0104] Step 401 may also comprise controlling one or more light sources 209 to illuminate the user 110 and / or the environment around the user 110 when the image(s) are generated or obtained by the image sensor 207. In particular, the light source(s) 209 can be controlled such that the lighting conditions on the user 110 and / or in the environment are consistent with lighting conditions when one or more previous images were generated. In this way, variability in the user’s appearance in the images due to different lighting conditions can be reduced. For example, the ambient light in the environment around the device 102 may be different in the morning and the evening, leading to apparent differences in images obtained at those times, and suitable adjustment of the light source(s) 209 can provide consistent lighting when images are obtained.

[0105] In step 403, the device 102 uses a language model, for example a LLM, to generate textual data describing the user 110 shown in the one or more images. The textual data may describe the face of the user 110 and / or relate to health aspects of the user 110. The language model may be a multi-modal language model that is able to generate the textual data at least from the images. That is, the language model is not restricted to receiving text-based inputs.

[0106] The textual data is generated such that it is not possible to identify the specific user 110 from the textual data itself. Thus, the security of the user’s data is protected in the event that a third party has access to the textual data (e.g. in accordance with step 305 below).

[0107] In step 405, the device 102 sends the generated textual data to a server. As the textual data is generated such that it is not possible to identify the specific user 110 from the textual data itself, the sending of the textual data to the server does not violate the security of the user’s data. For completeness, it should be noted that the image(s) used by the language model to generate the textual data are not sent to the server. The obtained image(s) may be deleted by the device 102, e.g. deleted from the memory 203, after they are used to generate the textual data. The device 102 may provide a confirmation message to the user 110 to indicate that the images have been deleted.

[0108] In some embodiments, prior to using the language model to generate the textual data, or as part of the operation of the language model in generating the textual data, the image(s) and other types of input data can be processed using one or more of the following algorithms or models. For example, as described with reference to Fig. 3 above, a respective encoding module and / or respective perceiving module may be used to encode the input data and / or perceive the content of the input data, with the language model 140 generating the textual data based on the encoded / perceived output.

[0109] In another example, an object detection algorithm can be used to detect a body part (e.g. face) in the image(s). As another example, a facial classifier can be used to identify a face of the user 110 in the image(s). A facial feature classifier can be used to identify facial features of the face of the user 110 (e.g. eyes, nose, mouth, eyebrows, etc.). A semantic segmentation algorithm can be used to segment the image(s) into pixel-level information.

[0110] The information generated by these algorithm(s) and / or model(s) is input to the language model, and the language model generates the textual data based on the image(s) and the input information.

[0111] In some embodiments, the language model generates the textual data based on the image(s) and measurements / information from one or more additional sensors. Any of these sensors 211 may be part of the device 102, or separate from the device 102.

[0112] The sensors 211 can include any of a microphone, thermometer, heart rate sensor, breathing rate sensor, blood pressure sensor and weighing scales. Thus, the measurements / measurement signals can comprise any of an audio recording of the user 110, a temperature of the user 110, heart rate information for the user 110, breathing rate of the user 110, blood pressure of the user 110, and weight of the user 110.

[0113] Thus, the device 102 obtains measurements relating to the user 110 from one or sensors 211 associated with the device 102. This step can comprise the device 102 controlling the sensor(s) 211 to perform measures and generate respective sensor measurements / sensor measurement signals. Alternatively, this step can comprise the device 102 / processing circuitry 201 receiving the measurements / measurement signals from separate sensors 211 , or retrieving previously-generated and stored sensor measurements from a storage location (e.g. the memory 203). The sensor measurements / measurement signals are used by the language model 108 in generating the textual data in step 403. In some embodiments any or all of the sensor measurements / measurement signals can be pre-processed so that they are in a suitable format for inputting to the language model. Alternatively, a multi-modal language model may be able to directly receive the sensor measurements / measurement signals.

[0114] The device 102 may store a data record for the user 110 (e.g. in memory 203). This data record may include previously-generated textual data for the user (i.e. generated from previous image(s)), and / or other information about the user, such as biometric information that can be used by the device 102 to identify the user 110. The device 102 may process the image(s) to extract biometric information for the user 110 (e.g. biometrics relating to face shape / structure), and can use the extracted biometric information to identify a data record for the user 110 that is stored at the device 102. The textual data generated in step 403 can be added to the stored data record.

[0115] In some embodiments, population profile information for the user 110 is associated with the generated textual data, and is sent to the server. This population profile information can provide useful context for the textual data at the server, e.g. the type of person that the textual data was generated for. The population profile information may have been manually entered into the device 102 by the user 110, or the device 102 may derive the population profile information from the image(s) of the user 110. The population profile information may comprise any of age, gender, height and ethnicity.

[0116] Although the computing devices described herein may include the illustrated combination of hardware components, other embodiments may comprise computing devices with different combinations of components. It is to be understood that these computing devices may comprise any suitable combination of hardware and / or software needed to perform the tasks, features, functions and methods disclosed herein. Determining, calculating, obtaining or similar operations described herein may be performed by processing circuitry, which may process information by, for example, converting the obtained information into other information, comparing the obtained information or converted information to information stored in the device, and / or performing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Moreover, while components are depicted as single boxes located within a larger box, or nested within multiple boxes, in practice, computing devices may comprise multiple different physical components that make up a single illustrated component, and functionality may be partitioned between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of the components may be partitioned between the processing circuitry and the communication interface. In another example, non-computationally intensive functions of any of such components may be implemented in software or firmware and computationally intensive functions may be implemented in hardware.

[0117] In certain embodiments, some or all of the functionality described herein may be provided by processing circuitry executing instructions stored on in memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuitry without executing instructions stored on a separate or discrete device- readable storage medium, such as in a hard-wired manner. In any of those particular embodiments, whether executing instructions stored on a non-transitory computer-readable storage medium or not, the processing circuitry can be configured to perform the described functionality.

[0118] The foregoing merely illustrates the principles of the disclosure. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous systems, arrangements, and procedures that, although not explicitly shown or described herein, embody the principles of the disclosure and can be thus within the scope of the disclosure. Various exemplary embodiments can be used together with one another, as well as interchangeably therewith, as should be understood by those having ordinary skill in the art.

Claims

Claims1. A computer-implemented method for operating a device (102), the method in the device (102) comprising: obtaining (401) one or more images (112, 114) of a user of the device (102); using (403) a language model (140) to generate textual data (118; 120; 122) describing the user (110) shown in the one or more images (112, 114); and sending (405) the generated textual data (118; 120; 122) to a server.

2. The method as claimed in claim 1 , wherein the method further comprises deleting the obtained image(s) (112, 114).

3. The method as claimed in claim 1 or 2, wherein the generated textual data (118; 120; 122) describes a face of the user (110) .

4. The method as claimed in any of claims 1 -3, wherein the generated textual data (118; 120; 122) relates to health aspects of the user (110) .

5. The method as claimed in any of claims 1-4, wherein the method further comprises projecting embeddings of the one or more images (112, 114) to a semantic space of the language model (140); and wherein the language model (140) uses the projected embeddings to generate the textual data (118; 120; 122).

6. The method as claimed in any of claims 1-4, wherein the method further comprises: processing the one or more images (112, 114) using one or more of: an object detection algorithm to detect a body part in the one or more images (112, 114); a facial classifier to identify a face of the user in the one or more images (112, 114); a facial feature classifier to identify facial features of the face of the user (110) in the one or more images (112, 114); and a semantic segmentation algorithm to segment the one or more images (112, 114) into pixel-level information; and inputting information output by any of the object detection algorithm, facial classifier, facial feature classifier and semantic segmentation algorithm into the language model (140), whereinthe language model (140) generates the textual data (118; 120; 122) based on the one or more images (112, 114) and the input information.

7. The method as claimed in any of claims 1-6, wherein the method further comprises: processing the one or more images (112, 114) to extract biometric information for the user(110); and identifying a data record for the user that is stored at the device (102), wherein the data record for the user (110) is identified based on the extracted biometric information.

8. The method as claimed in claim 7, wherein the data record comprises previously-generated textual data (118; 120; 122) for the user (110), and wherein the method further comprises: storing the generated textual data (118; 120; 122) in the data record.

9. The method as claimed in any of claims 1-8, wherein population profile information for the user (110) is associated with the generated textual data (118; 120; 122) and sent to the server.

10. The method as claimed in claim 9, wherein the population profile information comprises any of: age, gender, height and ethnicity of the user (110).

11. The method as claimed in any of claims 1-10, wherein the language model (140) is a multimodal language model (140) that is able to generate the textual data (118; 120; 122) at least from the one or more images (112, 114).

12. The method as claimed in any of claims 1-11 , wherein the method further comprises: obtaining measurements relating to the user (110) from one or more sensors (211) associated with the device (102); wherein the sensor measurements and one or more images (112, 114) are used by the language model (140) to generate the textual data (118; 120; 122) describing the user (110).

13. The method as claimed in claim 12, wherein the obtained sensor measurements is any one or more of: an audio recording of the user (110), a temperature of the user (110), heart rate information for the user (110), breathing rate of the user (110), blood pressure of the user (110), and weight of the user (110).

14. The method as claimed in any of claims 1-13, wherein the step of obtaining (401) one or more images (112, 114) comprises controlling an image sensor (207) to generate the one or more images (112, 114) of the user (110).

15. The method as claimed in any of claims 1-14, wherein the step of obtaining (401) one or more images (112, 114) comprises controlling one or more light sources (209) to illuminate the user (110) and / or the environment around the user (110) when the one or more images (112, 114) are generated.

16. The method as claimed in claim 15, wherein the one or more light sources (209) are controlled such that lighting conditions on the user (110) and / or the environment are consistent with lighting conditions when one or more previous images (112, 114) were generated.

17. The method as claimed in any of claims 1-16, wherein the device (102) is part of, or is, a surface.

18. The method as claimed in claim 17, wherein the surface is a reflective surface.

19. The method as claimed in claim 17 or 18, wherein the surface is a mirror or a smart mirror.

20. A computer program product comprising a computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform the method of any of claims 1-19.

21. A device (102) configured to perform the method of any of claims 1-19.

22. A device comprising a processor and a memory, said memory containing instructions executable by said processor whereby said device is operative to perform the method of any of claims 1-19.

23. A health monitoring system, comprising: a device as claimed in claim 21 or 22; and a server configured or operative to receive the textual data from the device.

24. The health monitoring system as claimed in claim 23, wherein the server is further configured or operative to train and / or test a machine learning, ML, model using the received textual data.

25. The health monitoring system as claimed in claim 24, wherein the health monitoring system is configured to generate health monitoring outputs based on the ML model.