System for training machine learning models to image process
By distributing a global explorer model to user devices, training local models, evaluating their performance, and combining the best models, the method addresses inefficiencies in existing trajectory prediction technologies, resulting in a more accurate and generalizable explorer model for all devices.
Patent Information
- Application Number
- PCT/KR2024/018188
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-30
AI Technical Summary
Existing methods for training machine learning models to predict user trajectories in 3D environments are inefficient and do not effectively leverage diverse user data across different devices.
A computer-implemented method where a global explorer model is distributed to user devices, and local explorer models are trained using local user data. A challenge is defined to evaluate the performance of these models, and the best models are combined to generate a current best explorer model, which is then updated and distributed to all devices.
This approach enables the determination of a current best explorer model that is more accurate, faster, or more generalizable than existing local models, allowing all user devices to benefit from superior performance based on diverse local user data.
Smart Images

Figure KR2024018188_30052025_PF_FP_ABST
Abstract
Description
SYSTEM FOR TRAINING MACHINE LEARNING MODELS TO IMAGE PROCESS
[0001] The present application generally relates to training of a machine learning, ML, model for image processing tasks and to the use of the trained model.
[0002] A three-dimensional (3D) representation of an environment can be constructed from a set of two-dimensional images of the environment. A user is able to remotely explore the environment using the 3D representation. Each user will explore the environment in a different manner, following their own subjective goal, for example to follow an object which has attracted their attention or to focus on a specific area of the environment.
[0003] The applicant has therefore identified the need for improved method for training models to predict a trajectory that will be followed through the environment by the user.
[0004] In a first approach of the present techniques, there is provided a computer-implemented method for training a global explorer model to output an image from an input digital representation of a 3D scene, the method comprising: sending a global explorer model from a server to a plurality of user devices; obtaining, at the server, a plurality of explorer models based on local explorer models which have been generated by training the global explorer model on each of the plurality of user devices using local user data; defining a challenge which is used to assess performance of each of the plurality of explorer models, wherein the challenge comprises a test digital representation of a 3D scene to be input into each explorer model to obtain an output image and a score threshold to be met by each optimal output image; evaluating each of the plurality of explorer models, by inputting the test digital representation of the 3D scene into each of the plurality of explorer models to obtain an output image which meets the score threshold; and obtaining an evaluation score for each of the plurality of explorer models, wherein the evaluation score is indicative of time taken to obtain the output image; selecting, using the server, a predetermined number of the plurality of explorer models based on the evaluation score; combining, by the server, the selected models to generate a combined model; determining, by the server, the current best explorer model using the combined model; and outputting, from the server to the plurality of user devices, the current best explorer model as an updated global explorer model.
[0005] The output image may be an optimal image which is determined with respect to any number of metrics or scores. For example, a given image may be assigned an aesthetic score. Therefore, the optimal image may be an image with a maximal aesthetic score (and is therefore optimal). It will be appreciated that an optimal image may be optimal with respect to any number of appropriate metrics or scores. For example, the semantic relevance (i.e. how relevant an object depicted in an image is with respect to a user prompt), scene relevance (i.e. how relevant a scene depicted in an image is with respect to a user prompt), fidelity, peak signal-to-noise ratio, PSNR and / or structural similarity index, SSIM.
[0006] A digital representation of a 3D scene may depict an environment which has been captured by the user. For example, a collection of 2D images may be captured by a user device and used to generate the digital representation of the 3D scene. In other words, a collection of 2D images, wherein each 2D image may depict the same user environment from different camera angles or poses, captured by a user device may be used to build or generate the digital representation of user environment. The digital representation may be digital reconstruction of an environment which allows a user to digitally move through the digital representation to view the representation of the environment from different viewpoints and angles, including ones for which no 2D images were captured. The 3D reconstruction may be generated using any number of known techniques. Specifically, using known techniques such as NeRF, a pixel-by-pixel map of the volume density and RGB colour values of the user environment depicted in each 2D image is output. In other words, the digital representation of the 3D scene (which may be termed a 3D scene for shorthand) corresponds to a 3D reconstruction of a user environment and may be defined as a map of the volume density and RGB colour values of the user environment depicted in each 2D image. Novel 2D views or images, e.g. along a desired axis, may be rendered from the 3D scene by integrating the volume density and colour values along the desired axis.
[0007] An explorer model according to the present techniques is a machine learning, ML, model which predicts at least one plausible future trajectory through the digital representation of the 3D scene. An explorer model may be trained to imitate a pattern of movement for each user from known user data. When presented with a particular digital representation of an environment, an explorer model may determine where a user will go. The output image may be an image from an end point of the predicted trajectory.
[0008] Advantageously, the present techniques enable the determination of a current best explorer model to be distributed to a plurality of user devices. The current best explorer model is based on an aggregation of existing local explorer models which have been trained locally on the user devices. The determination of the best explorer is made based on a challenge which is defined to assess the performance of the local explorer models whose performance will have been determined by the characteristics of each user device. That is, rather than continuing to operate with their existing local explorer models, which may have inferior performance compared to alternative existing local explorer model on another, independent user device, the present techniques enable other user devices to benefit from the superior local explorer model on other user devices. The current best explorer model may be more accurate, faster, or more generalizable than the existing local explorer model on a given device. For instance, the characteristics of local user data, used to train the local explorer models, may vary greatly from device to device. That is, the local user data on one device may be larger and more diverse than the local user data of an alternative device, resulting in better performance of the local explorer model. The present techniques allow each user device to benefit from local user data with more favourable characteristics in the context of training explorer models.
[0009] The method may further comprise: tracking time taken to output each image when evaluating each of the plurality of explorer models using the challenge; comparing the tracked time to a timeout threshold; and when it is determined that the timeout threshold is reached for a particular model, stopping evaluation of the particular model. Advantageously, the present techniques enable an efficient means of evaluating the performance of each explorer mode during training. By tracking the time taken to output each image, and comparing the time to a timeout threshold, the present techniques enable the efficient allocation of resources used to evaluate each model. By stopping the evaluation when the timeout threshold is reached, any resources used in the evaluation are inevitably made available for other purposes.
[0010] The method may further comprise iterating the steps of evaluating, selecting, combining and determining for each updated global explorer model. Therefore, the present techniques can be used to recursively assess the performance of each updated global explorer model.
[0011] In some cases, the plurality of obtained explorer models may be the plurality of local explorer models. In other words, obtaining the plurality of explorer models may be relatively simple and there may be a one-to-one correspondence between the obtained explorer models and each of the local explorer models in the plurality of explorer models.
[0012] In other cases, the plurality of obtained explorer models may be based on the plurality of local explorer models and may be obtained by defining a population of binary vectors with each binary vector having a value selected from zero or one for each of the plurality of explorer models, wherein a value of one in a binary vector for a local explorer model indicates that the local explorer model is selected. That is, rather than obtaining the plurality of explorer models in a direct fashion from the local explorer models, a more sophisticated approach may be taken with a view to efficiently combine the local explorer models.
[0013] For example, in cases where the plurality of obtained explorer models are the plurality of local explorer models, the step of selecting a predetermined number may comprise: ranking each of the plurality of local explorer models and the current best explorer model based on their evaluation score; and selecting a predetermined number of the ranked explorer models with highest rankings as the models to be combined. That is, the present techniques enable the selection of the highest ranking local explorer models as the models to be combined, thereby ensuring the efficiency of the combined model. In this case, the step of combining may comprise using a weighted sum of the selected explorer models.
[0014] In such an example, the step of determining the current best explorer may comprise: evaluating the combined model using the challenge; comparing the evaluation score of the combined model with the evaluation score of a previously stored best explorer model; and when the evaluation score of the combined model is higher than the evaluation score of the previously stored explorer model, storing the combined model as the current best explorer model. In this case, the step of combining may also comprise using a weighted sum of the selected explorer models.
[0015] In another example where the plurality of obtained explorer models are based on the plurality of local explorer models and are obtained by defining a population of binary vectors, each binary vector may further comprise values selected from zero or one for each of a plurality of weight values, wherein a value of one in a binary vector for a weight value indicates that the weight value is to be applied when combining the local explorer models.
[0016] Additionally or alternatively, the selecting, using the server, a predetermined number of the plurality of explorer models based on the evaluation score may comprise: ranking each of the binary vectors in the population based on their evaluation score; and selecting a predetermined number of the highest ranked binary vectors as the models to be combined. In either case, the method may further comprise creating a new population of binary vectors by: deleting, from the population, the lower ranked binary vectors which were not selected; and refilling the population of binary vectors using a binary vector representing the combined model and binary vectors which are mutations of the binary vector representing the combined model. That is, the plurality of local explorer models may be effectively combined by constructing a binary representation of the plurality of local explorer models and thereby avoiding a computationally intensive, brute force approach of exploring the performance of explicit combinations of the local explorer models.
[0017] The binary vectors which are mutations of the binary vector representing the combined model are modified versions of the binary vector representing the combined model. The modified versions may be versions of the binary vector representing the combined model, in which at least one of the components is randomly varied or mutated.
[0018] In examples where the plurality of explorer models based on local explorer models comprises defining a population of binary vectors, the step of determining the current best explorer model may comprise: evaluating the binary vectors in the population using the challenge; determining the binary vector with the highest evaluation score; and storing the combined model represented by the binary vector with the highest evaluation score as the current best explorer model.
[0019] Each explorer model may comprise inputs expressed as colour values and depth values for images within the 3D scene and outputs expressed as camera pose parameters, including a spatial location and viewing direction for a camera which would generate the output image within the 3D scene.
[0020] The challenge may further comprise at least one of a plurality of starting points for a trajectory within the 3D scene and a proximity area around an end point of the trajectory from which the image is output.
[0021] The evaluating using the challenge may comprise: evaluating, at each user device, the local explorer model using the challenge; and sending, from the user device to the server, the evaluation score for the local explorer model.
[0022] In a second approach of the present techniques, there is provided a computer-implemented method for using a trained global explorer model to output an optimal image from an input 3D scene. The method comprises: obtaining a trained global explorer model which has been trained as described above, obtaining a digital representation of a 3D scene, obtaining an output image by inputting the obtained digital representation into the trained global explorer model; and outputting the image to the user.
[0023] Obtaining the digital representation may comprise capturing, using a camera of the user device, a plurality of images of a user environment; and generating the digital representation using the plurality of images. The techniques described above may be used to generate the digital representation. Alternatively, the digital representation may be obtained from a database on the user device or by downloading from a remote server.
[0024] Obtaining an output image may comprise obtaining an optimal output image, wherein an optimal image is an image having a maximal aesthetic score. Alternatively, other scores may be used as explained above. Obtaining an output image may comprise predicting a trajectory through the input digital representation of a 3D scene and obtaining an output image at or near an endpoint of the predicted trajectory.
[0025] In a third approach of the present techniques, there is provided a user device comprising a memory and a processor coupled to the memory, the processor for obtaining a trained global explorer model which has been trained as described above, obtaining a digital representation of a 3D scene, obtaining an output image by inputting the obtained digital representation into the trained global explorer model; and outputting the image to the user. The image may be output on a display screen to the user.
[0026] The user device may be a constrained-resource device, but which has the minimum hardware capabilities to train a ML model and which communicates via a network. The UE may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge). It will be understood that this is a non-exhaustive and non-limiting list of example UEs.
[0027] In a fourth approach of the present techniques, there is provided a server for training a global explorer model to output an image from an input digital representation of a 3D scene. For example, the server may comprise a memory; a processor coupled to the memory, the processor configured to: send a global explorer model from the server to a plurality of user devices; obtain a plurality of explorer models based on local explorer models which have been generated by training the global explorer model on each of the plurality of user devices using local user data; define a challenge which is used to assess performance of each of the plurality of explorer models, wherein the challenge comprises a test digital representation of a 3D scene to be input into each explorer model to obtain an output image and a score threshold to be met by each output image; evaluate each of the plurality of explorer models, by inputting the test digital representation of the 3D scene into each of the plurality of explorer models to obtain an output image which meets the score threshold; and obtaining an evaluation score for each of the plurality of explorer models, wherein the evaluation score is indicative of time taken to obtain the output image; select a predetermined number of the plurality of explorer models based on the evaluation score; combine the selected models to generate a combined model; determine the current best explorer model using the combined model; and output to the plurality of user devices, the current best explorer model as an updated global explorer model.
[0028] Features described with respect to the method in the first approach apply equally to the second, third and fourth approaches, and are therefore not repeated.
[0029] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
[0030] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise sub-components which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.
[0031] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.
[0032] The techniques further provide processor control code to implement the above-described methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.
[0033] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.
[0034] In an embodiment, the present techniques may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.
[0035] The method described above may be wholly or partly performed on an apparatus, i.e. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.
[0036] As mentioned above, the present techniques may be implemented using an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / o may be implemented through a separate server / system.
[0037] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
[0038] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0039] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0040] Figure 1 is a block diagram of an example implementation of the present techniques;
[0041] Figure 2 is an example of user data which may be collected by the system of Figure 1 and different images extracted from the user data;
[0042] Figure 3 shows a plurality of user devices which may be used in the system of Figure 1;
[0043] Figure 4 shows an example of the local ML model which may be used to predict a trajectory through the user data;
[0044] Figure 5 is a schematic 2D representation of a 3D map showing a trajectory which may be predicted by the ML model of Figure 4;
[0045] Figure 6 is a flowchart of the steps for using the system of Figure 1;
[0046] Figure 7 is a schematic representation of combining local ML models as part of the method of Figure 6;
[0047] Figure 8 is a flowchart for one method of combining local ML models as part of the method of Figure 6;
[0048] Figure 9 is a flowchart for an alternative method for combining local ML models as part of the method of Figure 6; and
[0049] Figures 10 shows populations of vectors used in the method of Figure 9.
[0050] Broadly speaking, embodiments of the present techniques provide a method for performing image processing on a user device using a trained explorer model and for training the explorer model on a central server. When training the explorer model, the server defines a challenge to assess the performance of a plurality of local explorer models which have been received. Advantageously, the present techniques enable the outputting of optimal images based on a predicted trajectory through the input digital representation of the 3D scene.
[0051] Figure 1 is a block diagram of a system 10 comprising a server 100 for training a global machine learning, ML, model 106 and a plurality of user devices 150, each of which is used to generate a local ML model 180. For convenience only one user device but it will be appreciated that any number of devices may be connected to the server for training the models as explained in more detail below.
[0052] Each user device 150 may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, a drone, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge). It will be understood that this is a non-exhaustive and non-limiting list of example devices. The device 150 comprises the standard components, for example at least one processor 152 coupled to memory 154. There may also be a camera(s) 156 for capturing images and a user interface 158 to capture / output other user input. It will be appreciated that there may be other standard components which are not shown for simplicity. The at least one processor 152 may comprise one or more of: a microprocessor, a microcontroller, and an integrated circuit. The memory 154 may comprise volatile memory, such as random access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example.
[0053] The device 150 may collect user data 164 using a data collector 182 which collects data from any other module on the device (e.g. the camera (s)). The collected user data 164 is stored in storage 162 on the user device. The user data may be used in a training module 170 on the user device to train a local ML model 180 which predicts a trajectory that a user of the user device will take through a 3D environment and outputs an optimal image at the endpoint of the trajectory. Users typically move through a 3D environment in a subjective manner, perhaps following an object which has caught their attention. Each user may thus be considered to be maximising an unknown subjective goal.
[0054] The local ML model may be termed an explorer model. Such an explorer model predicts future localizations so as to predict plausible future trajectories. Typically, there is no objective function in an explorer model, but the explorer model may be trained to imitate a pattern of movement for each user from known user data. A goal of such an explorer model may be defined as determining where will the user go in this situation (i.e. when presented with a particular environment). In other words, each explorer model on a different user device will have a different goal which is determined by the user. The explorer model may be trained with 2D or 3D data. The explorer model may be any known model, e.g. a neural network.
[0055] The training module 170 may optionally comprise an evaluation module 172 which evaluates the local ML model 180, for example as explained in more detail below by evaluating a score for the output image . Local model parameters 168 for the local ML model may be stored in the storage 662.
[0056] The server 100 comprises a training module 110 which may be used to generate a global ML model based on the local ML models received from a plurality of user devices. The local training module 110 may comprise an evaluation module 112, a competition module 116 and a combination module 114. The evaluation module 112 may provide the same or similar function to an evaluation module on the user device and is used to evaluate each local ML model. A competition module 116 and a combination module 118 may be used to combine local ML models to form the global ML model and / or to select the optimal local ML model to form the global ML model.
[0057] Like the user device, the server 100 comprises standard components such as at least one processor 122 coupled to memory 124. Global model parameters 128 for the global ML model may be stored in storage 126 which is this example is shown as being on the server 100 but it will be appreciated that the storage 126 may be remote (or separate) from the server, e.g. a remote database. The global ML model which is generated on the server may be sent to the user device to be stored as the local ML model 180 and fine-tuned as needed on the user data. As explained in more detail below, only the local ML models and not any user data are sent to the server to generate the global ML model. The user data is also not stored on the server or database connected to the server. In this way, the user data can be kept confidential.
[0058] Figure 2 is an example of user data which may be collected by the system of Figure 1.
[0059] An example of user data which may be collected by the data collector is shown in 2a of Figure 2, which shows an image from a 3D reconstruction of an environment. In this example, for simplicity there is a male, a female and a dog in the environment. As illustrated by the boxes, the camera(s) on the user device takes a plurality of pictures along a trajectory 210 through the environment. 2b to 2d of Figure 2 illustrate different images which may have captured to form the 3D reconstruction in 2a of Figure 2. For example, 2b of Figure 2 shows an image captured using a camera close to a starting point 212 of the trajectory. Similarly, 2c of Figure 2 shows an image captured using a camera close to an endpoint of the trajectory 210. 2d of Figure 2 shows an image collected from approximately mid-point along the trajectory 210. Usually, the camera parameters are stored together with the image data. The camera parameters may be represented in a 3 Х 4 projection matrix called the camera matrix. The extrinsic parameters define the camera pose (position and orientation) while the intrinsic parameters specify the camera image format (focal length, pixel size, and image origin).
[0060] The collected images may be then combined using a data collector which can implement any suitable technique, e.g. NeRF as described in "Nerf: Representing scenes as neural radiance fields for view synthesis" by Mildenhall et al published in Communications of the ACM 65.1 (2021): 99-106 [1] or PlenOctrees as described in "PlenOctrees for real-time rendering of neural radiance fields" by Yu et al published in Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021. to create a 3D reconstruction of the environment by creating a plurality of new 2D images (i.e. views) of the environment. NeRF uses a neural network which takes as an input a spatial location (for example defined by 3D co-ordinates x, y, z) and viewing angle ( , ). The output is the volume density ( ) and colour (RGB) at each spatial location. A novel view may be rendered by shooting rays from each pixel and several sampled points along each ray. The colours along a ray are integrated to produce the final colour for each pixel. Other suitable techniques for creating the digital representation of the 3D scene are described in "Instant neural graphics primitives with a multiresolution hash encoding". By Muller et al published in ACM Transactions on Graphics (ToG) 41.4 (2022): 1-15, "Zip-NeRF: Anti-aliased grid-based neural radiance fields" by Barron et al published in arXiv preprint arXiv:2304.06706 (2023) and "3D gaussian splatting for real-time radiance field rendering" by Kerbl et al published in ACM Transactions on Graphics (ToG) 42.4 (2023): 1-14
[0061] In other words, a digital representation of 3D scene may be defined as a map of the volume density and RGB colour values of the user environment depicted in each collected image, wherein the collected images may be 2D images depicting the same user environment. As noted above, the collected images may be images collected along a trajectory 210. Novel 2D views or images, e.g. along a desired axis, may be rendered from the 3D scene by integrating the volume density and colour values along the desired axis.
[0062] For training, we use ground-truth images with known poses, and compare ground-truth images with new images which have been rendered and minimise the loss in the rendered images by any suitable loss function. Once trained for a scene, we choose a new camera pose and render the novel image.
[0063] It will be appreciated that for each environment there are multiple images that can be generated for each 3D reconstruction. Even when camera movements are limited to certain ranges, the camera pose may still have six degrees of freedom and there they may be over 46 million combinations for each camera pose when the camera angle is defined as 360 degrees without decimals. Each image (both collected and newly generated) may be scored with a score value which may represent a target goal. The score value may be an aesthetic value, for example the images shown in Figures 2b to 2d may each have scores of score B, score C and score D, respectively. Even though there are a large number of images, there will be only one image with a highest score value at the end of a user trajectory and this can be used to improve model prediction as explained below.
[0064] Each score may be a score for the 2D image rather than the 3D reconstruction. There are several methods used in ML models to quantify aesthetics using scores, such as the techniques described in one of "Image composition assessment with saliency-augmented multi-pattern pooling" by Zhang et al arXiv preprint arXiv:2104.03133 (2021), "Blind image quality assessment via cross-view consistency" by Zhu et al published in IEEE Transactions on Multimedia (2022) and "Lifelong blind image quality assessment" by Liu et al published in IEEE Transactions on Multimedia (2022). For example, a single aesthetics score may be given in a fixed range (e.g. from 1 to 10) or multiple N scores may be given (e.g. to mimic a user survey of N people or to score different aspects such as composition, lighting, rule of thirds, symmetry and so on). The score may be represented by a caption rather than a numeric value (e.g. this image has good colours and composition). The same or different aesthetic scores may be used for different local ML models to reflect the different aims of the users. As an alternative (or in addition to) using an aesthetic score, other scores may be used, for example a classification score (which is a measure of the confidence of an object in the image), a retrieval score (which is a measure of the retrieved image's similarity to a given image), and a counting score (which counts the number of elements of the desired class, for example faces in the retrieved image).
[0065] Figure 3 shows a plurality of user devices 150a, 150b, ... , 150n and for simplicity only shows the detail of the local storage 190a, 190b, ... , 190n on each user device. Each user device comprises different 3D scenes together with a trajectory database 190a, 190b, ... , 190n which includes a trajectory for each image. In the example of 2a to 2d of Figure 2, the collection of 2D images may be captured by the user device and used to build one of the 3D scenes which is shown in Figure 3. It will be appreciated that the 3D scenes may be obtained in different ways, e.g. by downloading from a third party service.
[0066] Every time a user selects one of the stored 3D scenes (i.e. loads a 3D scene) and selects a trajectory through the 3D environment, the trajectory database may be updated. As mentioned above, there are almost countless numbers of images for each 3D scene. Thus, the trajectory database may store only key frames for the selected 3D scene. Keyframes may be selected based on one or more of multiple different factors, for example one factor may be that a distance between camera poses must be above a distance threshold, another factor may be that a difference between camera views must be above a view threshold, and so on. The stored data for each key frame comprises the camera parameters, e.g. RGB (red, green, blue) information and camera pose. The stored data may be in vector format.
[0067] Figure 4 shows an example of the local ML model 480 which may be used to predict a trajectory through a particular scene and select a camera pose at the end of the trajectory which generates the camera view with the best score. The prediction and selection are done in the minimum amount of time possible. The local ML model may be termed an exploration module because the problem solved by the ML model may be considered an exploration problem. One objective is to minimise the camera trajectory length. Gradient descent is a possible solution but may have local minimum problems because of the large number of camera poses which give a high score value. In the exploration process, the ML model may decide the next camera pose movement.
[0068] As shown, the local ML model 480 may comprise a plurality of convolutional LSTM layers 412 and a plurality of fully connected (FC) layers 450,452. One of the FC layers 452 may compile or flatten the data from the convolutional layers to generate the data for the output layer 450. In this example, there are two of each type of layer, but it will be appreciated that this is merely indicative, and any suitable number of layers may be used.
[0069] As schematically illustrated in Figure 4, as a first input, the starting image 410 from the camera at the start of the trajectory is input to the ML model. The inputs may be expressed as RGB values and depth values and the outputs from the output layer 450 may be the camera pose parameters, namely the spatial location (x,y,z) and the viewing direction (θ,φ). The second input is the next image 420 from the next camera along the trajectory. The same inputs are used and the outputs for this image are also generated. This is repeated for each camera view along the trajectory until the final image 430 from the camera at the end of the trajectory is input to the ML model. The flattening FC layers 452 may be used to compile all the outputs from the various images.
[0070] During training, the model will be trained to minimize the residual of produced output, compared to the user location and angles. After this model is trained, the model can be used to predict the whole trajectory from start to end. In other words, the final image may be predicted based on the starting image. It will be appreciated that this is just one approach and other suitable training approaches can be used.
[0071] Figure 5 is a schematic 2D representation 500 of a 3D map showing a trajectory 510 which may be followed through the environment by a user. There is a starting position 512 and a best camera pose 514 at the end position of the trajectory 510. A proximity area 520 may be defined around the end position of the trajectory.
[0072] Figure 6 is a flowchart illustrating how the system of Figure 1 may be used to train local ML models to output an image and to use the locally trained models to generate a global model. In a first step S600, user data is collected on a first device. Simultaneously, user data is collected on a second device at step S610. The collection of user data is done as explained above. Using only the user data which is collected and stored locally on a particular device, the next step is to train the local ML model. For example, at step S602, the first local ML model is trained using the user data on the first device and at step S612, the second local ML model is trained using the user data on the second device. It will be appreciated that there may be any number of user devices and the collection of user data and subsequent training may be done for all devices.
[0073] At step S620, a target goal may be defined, and this is typically defined by the server. Although the definition of the target goal is shown subsequent to the collection and training steps, it may be done before or simultaneously with these steps. The target goal may define the number of scene(s) which are to be evaluated. The target goal may also define multiple (N) different starting locations for a particular scene. The target goal may also define the score for the best camera pose at the end location. The score may use aesthetics as described above. As shown in Figure 5, a proximity area may also be defined around the end point of a trajectory through each 3D reconstructed scene. The target goal may be termed a challenge.
[0074] Based on the target goal, each of the first and second trained models are used at steps S622 and S624 to infer the trajectories for each scene and each starting point. The inference may be done on each local device and thus the target goal may be sent to each local device before inference. Alternatively, the inference may be done at the server and each local model may be sent to the server from each local user device so that the server infers the trajectories.
[0075] Step S626 includes an optional evaluation of a timeout for each trajectory which may be defined in the target goal. Each of the local ML model may take different amounts of time to infer a trajectory to the best camera pose, for example the first trained model may use a sparse exploration method whereas the second trained model may use a more exhaustive exploration method. During inference, the time taken to reach the best camera pose at the end position for trajectory may thus be tracked. Alternatively, the time taken to reach the proximity area around the end position may be tracked. The tracked time may be compared to a timeout threshold. Any suitable timeout threshold may be used for example previous quickest time, previous mean time, previous maximum time for an individual trajectory. If the timeout threshold has been reached, the particular trajectory is timed out and the method loops back to determine the next trajectory within the target goal.
[0076] At step S628, the next step is to evaluate any trajectories which have been inferred, for example using the tracked time. The time to reach the best camera pose at the end position can be used as the evaluation score and the fastest time is considered to be the optimal score. The evaluation may be done on the user devices when the inference is done on the user devices. Alternatively, the evaluation may be done on the server, either after the inference has been done by the server or after the inference has been done on each device. When the inference is done on each device, the tracked time for each trajectory may be sent from the user device to the server for evaluation.
[0077] At step S632, some or all of the local ML models may be combined to generate a global model, for example based on the evaluation in step S628. There are various ways to combine the local ML models as described in more detail below. Each user has been exposed to different data which remains on the user device for data security. We can benefit from the variety of data, without exposing the data, by including an aggregation stage in which two or more explorer models are combined in order to improve their individual target goal score.
[0078] At step S634, the combined model is then output and may be shared at step S636 as initial pre-training weights with the first device and at step S638 with the second device. It will be appreciated that the combined model may be shared with all user devices. Alternatively, the combined model may only be shared with the user devices which generated the models which have been combined or user devices which share the target goal. In other words, the sharing step may be optional because some explorer models might have more than one goal and the wrong pretraining can prevent such explorer models from getting their best possible score for a specific target goal. When sharing takes place, the method then loops back to the local training steps S602, S612. The user devices then fine-tune the received combined ML model and the other steps can also optionally be repeated to keep updating the combined ML model.
[0079] There are various methods for combining the local ML models. Figure 7 schematically illustrates one such method which may be termed a "model soup" in line with the approach in "Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time" by Wortsman, Mitchell, et al published in International Conference on Machine Learning. PMLR, 2022 Model soups are a recent discovery in which an ensemble of models is formed by averaging the weights of the models instead of combining each of their individual outputs. The result is a single model, which is the average of many models with many different hyperparameter configurations. In other words, the method is a way of accumulating the learning from all the devices to obtain one unique model that outperforms them all.
[0080] As shown in Figure 7, at a certain time T, a central node requests model updates to a random number of devices (device a, device b, ... , device x). Each of these devices sends a local ML model back to the central node (server). In Figure 7, the local ML models are each represented by a star. The models are now temporarily saved in the centralized model as model ingredients together with the previous best model (if this is available). Another ingredient is the system dataset (individual performance for each model) which is used to individually validate each saved model.
[0081] The goal is then to produce a recipe (algorithm) that combines the ingredients, so the combination obtains better performance than the best-evaluated model (including the previous best). In the end, the resulting model is distributed back to the user devices (device a, device b, ... , device x) so that they can use the updated best model as pretraining in the next explorer training stage.
[0082] As a first example recipe shown in Figure 8, we propose a recipe which may be termed "greedy with discrete weights" and which may be carried out at the server (central node). In this algorithm, the models are not averaged but weighted in order to obtain a better combination. In a first step S800, a plurality of local models are obtained at the server and in step S802, these models are evaluated to obtain a score value which is indicative of their performance, e.g. speed of outputting the best camera pose. As explained above, as an alternative to the evaluation at the server, the evaluation may be carried out at the individual user devices and thus the evaluations may be received at the server. However, it will also be appreciated that the evaluation may also be carried out at the server. That is, the evaluation step S802 may be carried out at the user devices or at the server. At step S804, the models are sorted (ranked) based on the score values. At step S806, a predetermined number (say four or five) of the highest ranked models are combined with the highest performing model using a weighted average (for example using the weights: 0.2, 0.4, 0.5, 0.6, 0.8 in the sum). In the first iteration, the highest performing model will be the model determined from the previous best model.
[0083] At step S808, the combined model is evaluated. At step S810, the score value of the combined model is compared with the current highest performing model. If the score value of the combined model is lower, there is a determination as to whether there are more models to be considered at step S812. If so, the process loops back to step S806 to select the next group of models to be combined. If the score value of the combined model is higher, the combined model replaces the highest performing model at step S814. There is then the determination at step S812 and the process loops back to step S806 to select the next highest ranking group of models to be combined. If there are no more models (ingredients) to consider, the best performing model is then output at step S816. An example pseudo-code of the proposed algorithm is described below:
[0084] 1 Input: Potential soup ingredients {a,...,x, prev_best} {sorted in decreasing order of ValAcc
[0085] 2 ingredients {}
[0086] 3 w {0.2, 0.4, 0.5, 0.6, 0.8}
[0087] 4 best prev_best
[0088] 5 for each i in ingredients do
[0089] 6 if ValAcc(i*w best) best then
[0090] best i*w best
[0091] As a second example recipe shown in Figure 9, we propose a recipe which may be termed "genetic algorithms", and which may be carried out at the server (central node). In the algorithm of Figure 8, there is an assumption that a model that outperforms the best can be obtained by starting from the best model. However, this assumption might not be always true and Figure 9 proposes an alternative combination method. In a first step S900, a plurality of local models {a,... ,x,prev_best} are obtained at the server. In step S902, a population of binary vectors are defined based on the plurality of models. Each binary vector has the same length as the received set of local models and each binary vector is initialized with random values of 0 and 1. A zero at a particular point in the vector indicates that that particular model is not selected and a one indicates that the particular model at that location is selected.
[0092] Figures 10 shows populations of vectors used in the method of Figure 9.
[0093] 1001 of Figure 10 gives an example population. As explained below this starting population of vectors will evolve to obtain the best possible configuration.
[0094] The selected models are combined by averaging to give a single model which can be evaluated. Each of the vectors is then evaluated at step S904 to obtain a score value S1, S2, ... , Sn which is indicative of their performance, e.g. time to output best camera pose. Each of the vectors can then be ranked based on their score values and at step S906, the lowest ranking vectors can be deleted. The number of vectors to be deleted may be denoted by Nd. For example, looking at 1001 of Figure 10, all of the vectors between the top listed and the lowest listed vector have low scores and are deleted. Deletion of low ranking vectors leaves gaps in the population of vectors.
[0095] At step S908, a predetermined number Nc of the highest ranking models are combined. The combination may be simply by multiplying or adding. For example as shown in 1002 of Figure 10, the top and bottom listed vector are combined to provide a new vector at the second listed location in the population. The combinations of models may not be sufficient to fill all the gaps and thus the combinations may also be mutated at step S910 to create further new vectors. Each vector position has a probability P of being mutated. For example, as shown in 1003 of Figure 10, a mutated version of the second listed vector is used as the third listed vector.
[0096] At step S912 there is a determination to see if there are any more cycles to be carried out and if so, the method loops back to evaluating each binary vector in the current population. The number of generations (cycles) may be denoted by Ng. The deleting, combining and mutating steps are then repeated to create a new population. Once all cycles are completed, the highest performing model is then output at step S914. The model may be sent to the user devices as described above.
[0097] A variation of the process of Figure 9 is to include floating weight values to the vectors when defining the population at step S902. The float values describe the weight to be assigned to each of the models when combining the models at step S908. Such a recipe may be termed a "weighted genetic algorithm". For the averaging process described in Figure 9, the weights will be 0.5 for each model when two models are combined.
[0098] In summary, as described above we propose a pipeline to train trajectory prediction models using an unsupervised goal. The trained models are not aware of the objective. Data privacy is maintained by not sharing user data with any other device and the model is trained on device and later on validated on a centralized server. It will be appreciated that the possible use case (image aesthetics) which is described above is just one example and any other method that provides a relative score can benefit from this pipeline.
[0099] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.
Claims
1.A computer-implemented method for training a global explorer model to output an image from an input digital representation of a 3D scene, the method comprising:sending a global explorer model from a server to a plurality of user devices;obtaining, at the server, a plurality of explorer models based on local explorer models which have been generated by training the global explorer model on each of the plurality of user devices using local user data;defining a challenge which is used to assess performance of each of the plurality of explorer models, wherein the challenge comprises a test digital representation of a 3D scene to be input into each explorer model to obtain an output image and a score threshold to be met by each output image;evaluating each of the plurality of explorer models, byinputting the test digital representation into each of the plurality of explorer models to obtain an output image which meets the score threshold; andobtaining an evaluation score for each of the plurality of explorer models, wherein the evaluation score is indicative of time taken to obtain the output image;selecting, using the server, a predetermined number of the plurality of explorer models based on the evaluation score;combining, by the server, the selected models to generate a combined model;determining, by the server, the current best explorer model using the combined model; andoutputting, from the server to the plurality of user devices, the current best explorer model as an updated global explorer model.2.The method of claim 1, further comprising:tracking time taken to output each image when evaluating each of the plurality of explorer models using the challenge;comparing the tracked time to a timeout threshold; andwhen it is determined that the timeout threshold is reached for a particular model, stopping evaluation of the particular model.3.The method of claim 1 or claim 2, further comprising iterating the steps of evaluating, selecting, combining and determining for each updated global explorer model.4.The method of any preceding claim, wherein the plurality of obtained explorer models are the plurality of local explorer models and wherein selecting a predetermined number comprises:ranking each of the plurality of local explorer models and the current best explorer model based on their evaluation score; andselecting a predetermined number of the ranked explorer models with highest rankings as the models to be combined.5.The method of claim 4, further comprising, determining the current best explorer model by:evaluating the combined model using the challenge;comparing the evaluation score of the combined model with the evaluation score of a previously stored best explorer model; andwhen the evaluation score of the combined model is higher than the evaluation score of the previously stored explorer model, storing the combined model as the current best explorer model.6.The method of claim 4 or claim 5, wherein combining comprises using a weighted sum of the selected explorer models.7.The method of any one of claims 1 to 3, wherein obtaining the plurality of explorer models based on local explorer models comprises:defining a population of binary vectors with each binary vector having a value selected from zero or one for each of the plurality of explorer models, wherein a value of one in a binary vector for a local explorer model indicates that the local explorer model is selected.8.The method of claim 7, wherein each binary vector further comprises values selected from zero or one for each of a plurality of weight values, wherein a value of one in a binary vector for a weight value indicates that the weight value is to be applied when combining the local explorer models.9.The method of claim 7 or claim 8, wherein selecting, using the server, a predetermined number of the plurality of explorer models based on the evaluation score comprises:ranking each of the binary vectors in the population based on their evaluation score; andselecting a predetermined number of the highest ranked binary vectors as the models to be combined.10.The method of claim 9, further comprising creating a new population of binary vectors by:deleting, from the population, the lower ranked binary vectors which were not selected; andrefilling the population of binary vectors using a binary vector representing the combined model and binary vectors which are mutations of the binary vector representing the combined model.11.The method of any one of claims 7 to 10, further comprising determining the current best explorer model by:evaluating the binary vectors in the population using the challenge;determining the binary vector with the highest evaluation score; andstoring the combined model represented by the binary vector with the highest evaluation score as the current best explorer model.12.The method of any one of the preceding claims, wherein each explorer model comprises inputs expressed as colour values and depth values for images within the test digital representation of the 3D scene and outputs expressed as camera pose parameters, including a spatial location and viewing direction for a camera from which an output image is obtained.13.The method of any one of the preceding claims, wherein the challenge further comprises at least one of a plurality of starting points for a trajectory within the test digital representation of the 3D scene and a proximity area around an end point of the trajectory from which an output image is obtained. l.14.A computer-implemented method for using a trained global explorer model to output an image from an input digital representation of a 3D scene, the method comprising:obtaining a trained global explorer model from a server, wherein the global explorer model has been trained as set out in any one of the preceding claims;obtaining a digital representation of a 3D scene;obtaining an output image by inputting the obtained digital representation into the trained global explorer model; andoutputting the image.15.A server for training a global explorer model to output an image from an input digital representation of a 3D scene, the server comprising:a memory;a processor coupled to the memory, the processor configured to:send a global explorer model from the server to a plurality of user devices;obtain a plurality of explorer models based on local explorer models which have been generated by training the global explorer model on each of the plurality of user devices using local user data;define a challenge which is used to assess performance of each of the plurality of explorer models, wherein the challenge comprises a test digital representation of a 3D scene to be input into each explorer model to obtain an output image and a score threshold to be met by each output image;evaluate each of the plurality of explorer models, byinputting the test digital representation of the 3D scene into each of the plurality of explorer models to obtain an output image which meets the score threshold; andobtaining an evaluation score for each of the plurality of explorer models, wherein the evaluation score is indicative of time taken to obtain the output image;select a predetermined number of the plurality of explorer models based on the evaluation score;combine the selected models to generate a combined model;determine the current best explorer model using the combined model; andoutput to the plurality of user devices, the current best explorer model as an updated global explorer model.
Citation Information
Patent Citations
Image classification scheduling method based on edge calculation in embedded scene
CN109961097A
Image synthesis
EP4246988A1
Systems and Methods of Distributed Optimization
US20210382962A1
Parallel cross validation in collaborative machine learning
US20220343219A1
System and method for scalable machine learning in a communication network
WO2022066089A1