Data processing device on board a spacecraft
The spacecraft data processing device uses fast, deterministic operations to reduce resource consumption for neural networks, enabling efficient on-board execution and ground-based replication for analysis and supervision.
Patent Information
- Application Number
- FR2023007722
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-07-21
AI Technical Summary
Current neural networks require significant computing power and electrical energy, making them unsuitable for on-board space applications due to resource constraints, and their internal functioning is often opaque, complicating integration and performance analysis.
A data processing device on board a spacecraft implements a neural network using fast, deterministic multiplication and division operations that consume less computing power and energy, allowing for resource-efficient execution and enabling reproducibility on the ground, where identical neural networks can be trained and analyzed.
This approach reduces resource requirements for neural network operations in space applications while ensuring identical behavior can be replicated on the ground for analysis and supervision, enhancing autonomy and performance monitoring.
Smart Images

Figure 00000031_0000 
Figure 00000031_0001 
Figure 00000032_0000
Abstract
Description
Title of the invention: Data processing device on board a spacecraft 1. TECHNICAL DOMAIN
[0001] The field of the invention is that of neural networks implemented in spacecraft, in particular satellites and space probes.
[0002] The invention relates more particularly to a data processing device, on board a spacecraft and comprising a dedicated (for example an FPGA, "Field-Programmable Gate Array" or an ASIC "application-specific integrated circuit") or reprogrammable (for example a processor) computing machine, configured to implement a neural network.
[0003] The invention applies in particular, but not exclusively, to the case where a neural network implemented in a spacecraft is used for a data compression application.
[0004] The present invention is not limited to this particular application and is of interest for all applications embedded in spacecraft such as satellites (also called "embedded space applications") and based on the use of a neural network. 2. Technological background
[0005] Neural networks are becoming increasingly widespread and used in increasingly varied applications, for example under the names of artificial intelligence or machine learning. Data compression and image analysis are two examples of applications that can be implemented using neural networks. Neural networks can be, for example, deep neural networks (also referred to in English as DNN, for "Deep Neuronal Network") or convolutional neural networks (also referred to in English as CNN for "Convolutional Neuronal Network").
[0006] Due to their complexity, current neural networks require more and more resources (computing power and electrical energy in particular), in particular for the learning phases which involve numerous calculations, which generally excludes their use in on-board space applications.
[0007] US patent 2020 / 082269 entitled “Memory efficient neural networks” and filed in the name of NVIDIA CORPORATION teaches a method in which weight parameters are converted from a first floating point value representation to a second floating point value representation having a reduced number of bits. However, such a method only allows a marginal reduction in computing power requirements.
[0008] Furthermore, the available neural networks appear, for their users, as "black boxes", i.e. modules whose internal functioning is either inaccessible or deliberately hidden by the supplier, so that only their interactions can be studied. It is therefore very problematic to integrate into a spacecraft a neural network whose behavior and performance can only be analyzed from data received on the ground from this spacecraft. 3. SUMMARY
[0009] The present invention aims to remedy all or part of the drawbacks of the prior art.
[0010] To this end, according to a first aspect, the present invention relates to a data processing device, on board a spacecraft and comprising a means of communication with a device on the ground and a calculation machine, configured to implement a neural network receiving, as input, input data and generating, as output, at least one output data item, the neural network comprising a plurality of layers, a neuron of a given layer being determined as a function of at least one upstream value multiplied by a weight, the calculation machine comprising a processing unit and a memory space storing data representative of the neural network.
[0011] The computing machine is configured to perform so-called "slow" multiplication and division operations, consuming a first computing power and a first quantity of electrical energy, and so-called "fast" deterministic multiplication and division operations, consuming a second computing power less than the first computing power and / or a second quantity of electrical energy less than the first quantity of electrical energy, fast operations whose result has, depending on the numbers on which a fast operation is carried out, a precision less than or equal to the corresponding slow operation.
[0012] The computing machine is configured to perform at least one training phase of the neural network, each training phase comprising an adjustment of the data representative of the neural network by an iterative algorithm implementing training data, each iteration of the iterative algorithm comprising inference calculations followed by gradient descent calculations, said inference calculations using at least one fast deterministic division operation.
[0013] The means for communicating with the ground device is configured to transmit to the ground device and / or receive from the ground device training data.
[0014] The data processing device thus makes it possible, on the one hand, to adapt the computing power and the electrical energy required for said at least one phase. learning on board the spacecraft and, on the other hand, to implement, in the ground device, another neural network identical to the on-board neural network.
[0015] Thus, by performing at least some of the division operations according to a fast calculation (replacing a classic slow calculation) and, possibly, at least some of the multiplication operations according to a fast calculation (replacing a classic slow calculation), the resources required for these operations are reduced, in particular for the learning phases which involve numerous calculations. In particular, the computing power, i.e. the quantity of calculations performed per unit of time, increases with fixed resource consumption. These resources are generally energy and the number of logic gates (or LUTs on an FPGA). It is in particular the number of gates which is drastically reduced for division. This allows the neural network to be executed in the spacecraft, and therefore to be used in on-board space applications.In other words, the present invention makes it possible to improve the embedding (including for the learning phase(s)) of the neural network in the spacecraft (such as a satellite).
[0016] Furthermore, the fact that the rapid calculation of division operations and, possibly, of multiplication operations are both deterministic makes it possible to implement, in another data processing device on the ground, another neural network identical to the neural network on board the spacecraft.
[0017] In other words, the present invention allows reproducibility on the ground, in particular for each learning phase, of the neural network embedded in the spacecraft. It is therefore possible to reproduce the behavior of the embedded neural network on the ground and to study its behavior and performance. This allows, for example, an exact interpretation, in ground equipment, of telemetry data transmitted by a spacecraft. Thus, if an onboard space application uses a first neural network implemented in a spacecraft, there is equipment on the ground capable of implementing a second neural network identical to the first, with in particular the same learning results. However, the hardware resources available on the ground include greater computing power and electrical resources than those available in the spacecraft.
[0018] For example, in the case where the on-board space application is a data compression application using a first neural network providing a prediction signal to a compression module also executed in the spacecraft and generating compressed data, it is possible to execute on the ground a data decompression application using a second neural network identical to the first and providing the same prediction signal to a decompression module also executed on the ground and generating decompressed data from the data compressed data received from the spacecraft.
[0019] In another example, for an onboard space application, the onboard space application that is executed in the spacecraft is re-executed in ground equipment, to verify on the ground that everything is happening correctly in the spacecraft. This allows a gain in autonomy of the onboard space application, while performing supervision on the ground, for example to ensure that the learning of the neural network executed in the spacecraft does not diverge.
[0020] In embodiments, the computing machine is configured so that the gradient descent computations use fast deterministic multiplication operations and fast deterministic division operations.
[0021] The advantages of implementing the invention, recalled above, thus cover not only inference calculations, but also gradient descent calculations.
[0022] In embodiments, the on-board data processing device further comprises a performance evaluation module configured to, on the one hand, store a configuration of the neural network having a determined measured performance on a determined date and, on the other hand, measure the performance of the neural network obtained by adjusting its representative data by a learning phase, and compare the stored performance with the measured performance, the computing machine being configured to return the neural network to the stored network configuration if the measured performance of the network configuration adjusted by a learning phase is lower than the performance of the stored network configuration.
[0023] Thus, if the performance of the neural network decreases, for example due to the use of fast operations, the most efficient configuration of the neural network is restored.
[0024] In embodiments, the calculating machine is configured so that each of the fast division operations performed, for two stored numbers D and F, corresponds to: ( SD X SF ) X ( 1 + Md - Mf) X SD is the sign of D, SF is the sign of F, Md is the mantissa of D, MF is the mantissa of F, ED is the exponent of D, and EF is the exponent of F, mantissa and exponent being defined in a given base.
[0025] This fast calculation is deterministic in that it allows working on integer arithmetic to perform a fast division. Furthermore, in the case of a large number of classic slow divisions performed on floating-point arithmetic, the management of rounding may prove to be non-deterministic, for example when using a commercial calculator of the CPU or GPU type.
[0026] In embodiments, the computing machine is configured so that each rapid multiplication operations performed, for two memorized numbers A and B, correspond to: ( SAX SB ) X ( 1 + M 4 + MB ) X 9e'^A^lcs two memorized numbers A and B being decomposed as follows: A = S4 x ( 1 + Ma ) x 2E^CB = Sbx(ï + Mb)k 2^°* Sa is the sign of A, SB is the sign of B, MA is the mantissa of A, MB is the mantissa of B, EA is the exponent of A, Eb is the exponent of B, and C is a bias used in the sign decomposition, mantissa and exponent being defined in a given base.
[0027] This fast calculation is deterministic in that it allows working on integer arithmetic to perform a fast multiplication. Furthermore, in the case of a large number of classic slow multiplications performed on floating-point arithmetic, the management of rounding may prove to be non-deterministic, for example when using a commercial calculator of the CPU or GPU type.
[0028] For example, the base is determined by the IEEE 754-2008 standard.
[0029] In embodiments, the computing machine is configured so that all division operations included in said inference and gradient descent computations are fast division operations, except division operations included in the execution of the Softmax function for a final layer of the neural network.
[0030] Thus, one benefits as much as possible from the advantages of the implementation of the present invention, with fast calculations. On the other hand, for the division operations included in the execution of the Softmax function, one uses the classic slow calculation to avoid resulting in a learning, or training, which is rapidly unstable in most cases. This instability is due to the fact that the error is back-propagated through the last layer. In this last layer of the neural network, the Softmax function is performed, the output of which is a probability distribution on a finite set of classes, which covers classification problems but also the statistical modeling of data, for example in the context of a compression algorithm.Coupled with the fact that this last layer uses exponential functions, the induced error is non-linear and does not average out when back-propagated in the other layers, unlike the fast multiplication used during inference calculations, whose errors average out very well.
[0031] In embodiments, the computing machine is configured so that all multiplication operations included in said inference and gradient descent computations are fast multiplication operations.
[0032] These embodiments use fast computation instead of conventional slow computation for multiplication operations as much as possible.
[0033] In embodiments, said input data is raw data consisting of symbols and wherein said at least one output data is representative of a prediction of an occurrence of the symbols, said at least one output data being provided to a module, implemented on board the spacecraft, for entropic compression of said raw data.
[0034] This particular implementation corresponds to the case where the neural network implemented in the spacecraft is used for a data compression application.
[0035] It is recalled, however, that the present invention is not limited to this particular on-board space application and is of interest for all on-board space applications which are based on the use of a neural network.
[0036] In embodiments, the computing machine is configured to operate according to a first operating mode corresponding to the learning phases and according to a second operating mode corresponding to inference calculations without learning, fast multiplications and fast divisions being used in these two operating modes.
[0037] The advantages of implementing the invention thus extend to both operating modes.
[0038] In embodiments, the computing machine is configured so that each value carried by one of the neurons and each weighting factor associated with a transmission of a value from one neuron to another, are coded on the same number of binary data, in the same format in the form of a floating point number.
[0039] For example, this floating point number can comprise 32 binary data (“bits”) with floating point or 16 bits for example for embedded applications (notably of the calculator type for robotic applications).
[0040] In embodiments, the neural network is configured to be tuned by initial training prior to launch of the spacecraft.
[0041] According to a second aspect, the present invention aims at a data processing system for a space application comprising a data processing device on board a spacecraft which is the subject of the present invention and as succinctly set out above and a ground processing device in a ground installation comprising a means of communication with the communication means of the spacecraft, the ground data processing device comprising a computing machine identical to the onboard computing machine.
[0042] In embodiments, the ground computing machine further comprises a performance evaluation module configured to, on the one hand, store a configuration of the ground neural network having a determined measured performance on a determined date and, on the other hand, measure the performance of the ground neural network obtained by adjusting its representative data by a learning phase, and comparing the stored performance with the measured performance, the computing machine being configured to return the ground-based neural network and the on-board neural network to the stored network configuration if the measured performance of the network configuration set by a learning phase is lower than the performance of the stored network configuration.
[0043] Thus, the performance evaluation of the neural network can be performed on the ground and, if the performance of the neural network decreases, for example due to the use of fast operations, the best performing configuration of the ground neural network and the on-board neural network is restored.
[0044] According to a third aspect, the present invention relates to a method for processing data on board a spacecraft comprising a step of communication with a device on the ground and a calculation step implementing a neural network receiving, as input, input data and generating, as output, at least one output data, the neural network comprising a plurality of layers, a neuron of a given layer being determined as a function of at least one upstream value multiplied by a weight, the calculation step also implementing a memory space storing data representative of the neural network.
[0045] The calculations include so-called “slow” multiplication and division operations, consuming a first computing power and a first quantity of electrical energy, and so-called “fast” deterministic multiplication and division operations, consuming a second computing power less than the first computing power and / or a second quantity of electrical energy less than the first quantity of electrical energy, fast operations whose result has, depending on the numbers on which a fast operation is carried out, a precision less than or equal to the corresponding slow operation.
[0046] The calculation step comprises at least one phase of learning the neural network, each learning phase comprising an adjustment of the data representative of the neural network by an iterative algorithm implementing learning data, each iteration of the iterative algorithm comprising inference calculations followed by gradient descent calculations, said inference calculations using at least one fast deterministic division operation.
[0047] The step of communicating with the ground device transmits to the ground device and / or receives from the ground device learning data.
[0048] The data processing method thus makes it possible, on the one hand, to adapt the computing power and the electrical energy necessary for said at least one learning phase on board the spacecraft and, on the other hand, to implement, in the ground device, another neural network identical to the on-board neural network.
[0049] The advantages, aims and particular characteristics of this method being similar to those of the device which is the subject of the invention, they are not recalled here. 4. LIST OF FIGURES
[0050] Other aims, characteristics and advantages of the invention will appear on reading the following description, given by way of illustrative and non-limiting example, in relation to the appended drawings, in which:
[0051] [Fig.l] represents an example of a neural network according to a first known structure;
[0052] [Fig.2] represents another example of a neural network according to a second known structure;
[0053] [Fig.3] illustrates an example of a spacecraft according to a particular embodiment of the invention, incorporating a data processing device implementing a compression algorithm composed of an entropic coding module and a statistical model of the symbols in the form of a neural network;
[0054] [Fig.4] illustrates an example of the structure of the data processing device of the [Fig.3], according to a particular embodiment of the invention;
[0055] [Fig.5] illustrates an example of the implementation of neural network learning implemented by the data processing device of [Fig.3];
[0056] [Fig.6] illustrates an entropic coding;
[0057] [Fig.7a] represents an example of a method according to the invention;
[0058] [Fig.7b] illustrates an example of processing implemented in the first function combination according to the invention;
[0059] [Fig.7c] illustrates the properties of continuity, monotonicity and differentiability of the first combination function according to the invention;
[0060] [Fig.8a] illustrates a photo of the surface of the moon;
[0061] [Fig.8b] illustrates areas of the photo of [Fig.8a] that may be used to the moon landing of a spacecraft, the areas in question having been classified as such by a network trained according to the invention;
[0062] [Fig.8c] illustrates an example of classification performance obtained in the context of the detection of characteristic areas on a photo with a neural network implemented according to the invention;
[0063] [Fig.9a] illustrates a noisy photo of a sports field;
[0064] [Fig.9b] illustrates the result of the denoising of the photo of [Fig.9a] by a network of neurons trained according to the invention, where the inference is also done according to the invention;
[0065] [Fig.9c] illustrates the result of the denoising of the photo of [Fig.9a] by a network of neurons trained half according to the invention and half using classic slow multiplication, the inference being done according to the invention;
[0066] [Fig.9d] illustrates the result of the denoising of the photo of [Fig.9a] by a network of neurons trained using classical slow multiplication, inference also being done using classical slow multiplication (control case);
[0067] [Fig. 10a] represents a detail of [Fig.9a];
[0068] [Fig. 10b] represents a detail of [Fig.9b];
[0069] [Fig. 10c] represents a detail of [Fig.9c];
[0070] [Fig.lOd] represents a detail of [Fig.9d]; [0071 ] [Fig.lia] illustrates a noisy photo of a road;
[0072] [Fig. 11b] illustrates the result of denoising the photo of [Fig.11a] by a network of neurons using classical slow multiplication for learning and inference (control case);
[0073] [Fig. 1 le] illustrates the result of the denoising of the photo of [Fig. 11a] by a neural network implemented according to the invention, for learning and for inference;
[0074] [Fig. 1 Id] illustrates the result of the denoising of the photo of [Fig.11a] by a neural network trained half according to the invention and half using a classic slow multiplication, the inference being done according to the invention;
[0075] [Fig. 12a] represents a detail of [Fig. 11a];
[0076] [Fig. 12b] represents a detail of [Fig. 11b];
[0077] [Fig. 12c] represents a detail of [Fig. 1 le];
[0078] [Fig. 12d] represents a detail of [Fig. 1 Id];
[0079] [Fig. 13] illustrates an example of the evolution of the loss function obtained during the learning according to the invention of a neural network, in the context of image denoising; and
[0080] [Fig. 14] represents a system of data processing devices for a space application comprising a processing device on board a spacecraft and a processing device on the ground. 5. DETAILED DESCRIPTION
[0081] In all the figures of this document, identical elements and steps are designated by the same numerical reference.
[0082] [Fig.l] represents a network 100 of neurons of the “feedforward” type, which can for example be implemented by the device and the method which are the subject of the invention. The scope of the invention is in no way limited to this type of neural network, the invention can for example be applied to deep neural networks, also designated by DNN (acronym for Deep Neuronal Network), convolutional neural networks, also designated by CNN (acronym for Convolution Neuronal Network) in which the layers define at least a first block, as input, for extracting characteristics and a second block, as output.
[0083] The neural network may also comprise dense layers. The neural network may also be of the deep convolutional network type, also designated by DCN (acronym for Deep Convolution Network). The neural network can also be of the recurrent type, also designated by RNN (acronym for Recurrent Neural Network), or of the restricted Boltzmann machine type, also designated by RBM (acronym for Restricted Boltzmann Machine), long-short-term memory unit, also designated by LSTM (acronym for Long Short-Term Memory), gated recurrent unit, also designated by GRU (acronym for Gated Recurrent Unit), or self-adaptive map, also designated by SOM (acronym for Self Organizing Maps).
[0084] The neural network may notably use a number of forward or backward links with a weight and an activation function. In another example, the neural network may include functionality to perform clustering, principal component analysis, also referred to as PCA (acronym for Principal Component Analysis), latent semantic analysis, also referred to as LSA (acronym for Latent Semantic Analysis), and / or another unsupervised learning technique. The neural network may also implement the functionality of a regression model, a support vector machine, a decision tree, a random forest, a gradient boosted tree, a naive Bayes classifier, a Bayesian network, a hierarchical model and / or an ensemble model.
[0085] The network 100 is a directed graph of nodes, called neurons 110, arranged in successive layers L1, L2, L3 in [Fig.l]. Each neuron 110 of a given layer receives information from one or more neurons 110, for example from the previous layer, and combines this information according to weightings identified in [Fig.l] by the terms viX' P (if we refer to the j-th neuron of the z-th layer, the term h'V P weights the k-th input value of the j-th neuron in question). ^9J Each neuron 110 has for example an activation threshold bj. In [Fig.l], the output of the j-th neuron 110 of the z-th layer is expressed as a function A0 ] of its input The parameters that are the weights and the thresholds j \ j / J activation are optimized in practice during a training phase of the network 100 on a training data set.
[0086] For example, for the inputs of a given training set, the theoretical output results of the network are known. An optimization consists, for example, in minimizing the sum of the squares of the differences between the calculated outputs and the expected outputs for given input values xk. For applications in particular of classification and compression, cross-entropy is used, for example.
[0087] The neural network 100 of [Fig.l] is of the “feed-forward” type according to the ter Anglo-Saxon network, where each layer feeds the next, while a recurrent network allows loops. Such a network propagates the network's input to the following layers without ever going back.
[0088] The neural network 120 of [Fig.2] comprises feedback loops between the different layers L1, ..., L7. More particularly, certain neurons 110 here receive information from their successors (i.e. neurons 110 of the following layer according to the direction of propagation of the information from the input of the network 120 to the output of the network 120) and combine this information according to corresponding weightings. Furthermore, certain neurons 110 here also receive information from predecessors belonging to a previous layer (for example, a neuron of the layer L6 here receives information as delivered by a neuron of the layer L3) and combine this information according to corresponding weightings.
[0089] More generally, the method according to the invention can be applied to numerous types of neural networks.
[0090] As illustrated in [Fig. 3], for the implementation of the invention, a spacecraft 30 comprises a data processing device 32 comprising a means of communication 37 with a device on the ground (see [Fig. 14]) and a computing machine 38 illustrated in [Fig. 4]. In the example of [Fig. 3], the neural network 33 implements, in a non-limiting manner, a statistical model of the symbols used in combination with the entropic coding module 35, in order to carry out data compression.
[0091] We now present, in relation to [Fig. 4] an example of a computing machine 38 making it possible to implement a neural network with at least one fast division operation and generating outputs as a function of its inputs. The computing machine comprises, for example, a random access memory 323 (for example a RAM memory), a processing unit 321, equipped for example with one (or more) processor(s), and driven by a computer program stored in a read-only memory 322 (for example a ROM memory or a hard disk). Upon initialization, the code instructions of the computer program are for example loaded into the random access memory 323 before being executed by the processor of the processing unit 321. The computing machine 38 comprises for example, without limitation, an interconnection (bus) which connects one or more processing units, one or more input / output interfaces coupled to one or more inputs / outputs and the memory.The instructions can be executed, for example, using a reprogrammable computing machine.
[0092] The processing unit or units may be any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU), a signal processing processor (DSP), a microcontroller, a application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).
[0093] In the case where the device 32 is produced at least in part with a reprogrammable computing machine, the corresponding program (i.e. the sequence of instructions) may be stored in a removable storage medium (such as for example a CD-ROM, a DVD-ROM, a USB key) or not, this storage medium being partially or totally readable by a computer or a processor. This program may include in particular a set of instructions for the implementation of a neural network according to the invention where the execution of a division operator of two stored numbers D and F corresponds to:
[0094] (5^xSF) x (1 + MD-MF) x2ZW
[0095] where:
[0096] - SD is the sign of D;
[0097] - SF is the sign of F;
[0098] - Md is the mantissa of D;
[0099] - Mf is the mantissa of F;
[0100] - Ed is the exponent of D and
[0101] - Ef is the exponent of F.
[0102] In certain embodiments, according to the instruction set, the execution of a multiplication operator of the two stored numbers A and B corresponds to: [°103] (S4 x SB) x (1 + Ma + Mb) x 2Ea+e^2c
[0104] the two stored numbers A and B being decomposed as follows:
[0105] A = S x (1 + M4) x 2E^C
[0106] B = x ( 1 + Mb ) x
[0107] where:
[0108] - SA is the sign of A;
[0109] - SB is the sign of B;
[0110] - Ma is the mantissa of A;
[0111] - Mb is the mantissa of B;
[0112] - Ea is the exponent of A;
[0113] - Eb is the exponent of B; and
[0114] - It is a bias used in the decomposition into sign, mantissa and exponent in a determined basis.
[0115] The instructions can be standard, such as the C language, or specific to the microprocessor or computer that executes them.
[0116] By executing a computer program stored in the memory space of memories 322 and 323, the computing machine 38 implements the neural network 33 receiving, as input, input data 31 and generating, as output, at least one output data 34. As explained with regard to [Fig.l] and [Fig.2], the neural network 33 comprises a plurality of layers, a neuron 110 of a given layer being determined as a function of at least one upstream value multiplied by a weight. The RAM 323 stores data representative of the neural network 33.
[0117] In the particular case illustrated in [Fig.l], a data compression module 35 is also on board the spacecraft 30. The computing machine 38 is configured to perform so-called “slow” multiplication and division operations, consuming a first computing power and a first quantity of electrical energy. These operations are those conventionally used in neural networks known in the prior art. The computing machine 38 is also configured to perform so-called “fast” deterministic multiplication and division operations, consuming a second computing power less than the first computing power and / or a second quantity of electrical energy less than the first quantity of electrical energy.These fast multiplication and division operations give results which, depending on the numbers they are dealing with, have a precision less than or equal to the corresponding slow operation, multiplication or division respectively. Examples of fast operations are given below.
[0118] The means of communication 37 with the ground device transmits to the ground device and / or receives from the ground device learning data.
[0119] As illustrated in [Fig.5], after a pre-learning step 41, or initial learning, adjusting the neural network 33, preferably before the launch of the spacecraft, the computing machine 38 carries out at least one other learning step 42 of the neural network 33. Each learning phase 42 comprises an adjustment of the data representative of the neural network 33 by an iterative algorithm implementing learning data. The term adjustment here means an adjustment by learning of the data representative of the weightings and / or the activation threshold values defining the neural network 33.
[0120] Each of the iterations 423 of the iterative algorithm comprises inference calculations 421 followed by gradient descent calculations 422. The inference calculations 421 use at least one fast deterministic division operation.
[0121] By using at least one fast division operation, the onboard data processing device 32 makes it possible to adapt the computing power and electrical energy required for each learning phase on board the spacecraft. In addition, the onboard data processing device 32 makes it possible to implement, in the ground device, another neural network identical to the onboard neural network.
[0122] In the example illustrated in [Fig.l], an on-board space application is a data compression application performed by the compression module 35. The on-board neural network 33 provides a prediction signal to the compression module 35 which generates compressed data which are transmitted to the ground device by the communication means 37. With the duplication of the neural network 33, in the ground processing device, a data decompression application is executed on the ground using a neural network identical to the on-board neural network 33 and therefore the same prediction signal is obtained used by a decompression module also executed on the ground, which generates decompressed data from the compressed data received from the spacecraft 30.
[0123] An example of an entropic data compression algorithm 50 is described with regard to [Fig.6]. Symbols 51 to be coded are obtained as inputs. These symbols 51, which are, for example, bytes representing values obtained by sensors, or pixel values, are processed by a statistical model 52, which divides them into classes 53, for example according to a histogram. An update 55 is carried out from the symbols 51 received as input. This update of the statistical model implements the embedded neural network 33.
[0124] From this distribution into classes, we determine a number of binary data (bits) used to code each symbol, according to the following rule: the more probable a symbol is, that is to say the more populated the class in which it is found, the fewer bits we use to code it.
[0125] In the example illustrated in [Fig.6], the input data 51 are raw data consisting of symbols and at least one output data is representative of a prediction of an occurrence of the symbols, this output data being supplied to a module 35, implemented on board the spacecraft 30, for entropic compression of the raw input data 51.
[0126] In this data compression, multiplications are used during inference and when calculating the gradient by the error backpropagation method. Divisions are used in all statistical normalization steps (e.g. Batch Normalization, or Softmax) but also in calculating the descent direction from the gradient during neural network training (see, for example, the definition of classical optimization algorithms such as Adam or RMSProp).
[0127] Preferably, a fast division operation is used for all operations involving division except Softmax. In mathematics, the softmax function, or normalized exponential function, is a generalization of the logistic function which takes as input a vector z = (zi, ... , zK) of K real numbers and which outputs a vector o (z) of K strictly positive real numbers and of sum 1. The The j-component of vector o(z) is equal to the exponential of the j-component of vector z divided by the sum of the exponentials of all components of z. In probability theory, the output of the softmax function can be used to represent a categorical distribution - that is, a probability distribution over K different possible outcomes.
[0128] This Softmax step is generally the last layer of a neural network whose output is a probability distribution over a finite set of classes. This covers classification problems but also statistical modeling of data, for example in the context of a compression algorithm. Using fast division for Softmax would result in rapidly unstable learning in most cases. This is due to the fact that the error is back-propagated through this layer, the last of the network, before reaching all the other layers. Coupled with the fact that this step uses exponential functions, the induced error is non-linear and is not compensated by averaging when it is back-propagated in the other layers (unlike fast multiplication used during inference whose errors compensate each other on average very well).
[0129] In variants, the computing machine 38 is configured to operate according to a first operating mode corresponding to the learning phases and according to a second operating mode corresponding to inference calculations without learning, fast multiplications and fast divisions being used in these two operating modes.
[0130] Preferably, the computing machine 38 is configured so that the gradient descent calculations use fast deterministic multiplication operations and fast deterministic division operations.
[0131] In the application to data compression, and in other applications of the invention, preferably, the network is trained by an algorithm aiming to adjust at least the weights from an error calculated at the output and in which the error at the output of each neuron is propagated upstream by a sum of the multiplications of each error, according to the fast multiplication, with at least one weight specific to the connection with an upstream neuron, the weights being adjusted iteratively. The gradient is thus obtained, the correction applied to the weights is calculated from this gradient, generally by normalizing it by an estimate of its variance. A fast term-by-term division is therefore used by the estimate of the standard deviation of each term. Each term of this vector thus obtained from the gradient is the correction applied to the corresponding weight of the neural network.
[0132] As shown in [Fig.7a], the method comprises at least one step E300, where for at least one layer of neurons, the output of each neuron 110 of the layer is determined as a function of at least one upstream value (for example xk if we refer to the layer L1 of the network 100 of [Fig.l]) associated with a weight (for example if we refer to the j-th neuron of the first layer L1 in the network 100 of [Fig.l] ) to be combined. For this purpose, each upstream value and its weight are decomposed into sign, mantissa and exponent to be combined, in a fast multiplication, in the form:
[0133] Fill (upstream value, weight) = (SAxSB) x (1 + MA + MB) x2Ea+e^2c (1)
[0134] the upstream value and the weight being broken down as follows:
[0135] upstream value = SA x ( 1 + MA ) x 2Ea+c [0i36] weight = Sb*(\ + Mb)k 2E^C (2)
[0137] where:
[0138] - Combl is a first combination function;
[0139] - SA is the sign of the upstream value;
[0140] - SB is the sign of the weight associated with the upstream value;
[0141] - Ma is the mantissa of the upstream value;
[0142] - Mb is the mantissa of the weight associated with this upstream value;
[0143] - Ea is the exponent of the upstream value;
[0144] - Eb is the exponent of the weight associated with this upstream value; and
[0145] - It is a bias used in the decomposition into sign, mantissa and exponent in a determined basis.
[0146] Depending on the embodiment considered, the bias C can be positive or negative. The bias depends on the coding used. More particularly, the bias makes it possible to code negative exponents, which corresponds to the coding used by the IEEE-754 standard, but it is also possible to use a “2AN complement” coding which would make it possible to eliminate the subtractions or additions linked to the bias. Thus, the bias can be zero in equation (1).
[0147] The mantissas MA and MB can each be coded on a binary word interpreted as an integer when adding the mantissas. In the event of an overflow resulting from the sum of the mantissas MA + MB, the overflow bit is reported as an increment bit of the exponent.
[0148] Except for saturation management, this combination is obtained using a simple addition on the binary writing of the numbers interpreted as integers.
[0149] In a degraded embodiment, the sum of the mantissas is calculated using the addition of floating point numbers, for example to facilitate implementation on already existing tools such as TensorFlow (registered trademark) or PyTorch (registered trademark). The result of this calculation is however not continuous. The discontinuity negatively impacts performance in certain applications such as for example for image denoising or signal denoising. If the continuity is not verified over the entire definition interval, however, this does not prevent the calculation of a derivative and therefore does not hinder learning. Indeed, in the event of discontinuity at certain points, it is necessary to choose between the right and left derivative, but these are sufficiently similar and statistically rare not to significantly impact the performance of a stochastic learning method.
[0150] As illustrated in [Fig.7b], the exponents EA and EB are each coded on a binary word of N bits, the signs SA and SB are each coded on one bit, the mantissas are coded on K bits. The carryover of the exponent increment bit is symbolized in [Fig.7b] by the dotted arrow. For example, the first combination function Combl can be written in C language according to the following routine when the bias C corresponds to the value 127 and the mantissas are coded on K = 23 bits (case of [Fig.7b]):
[0151] float Comb 1 (float a, float b)
[0152] {
[0153] union
[0154] {
[0155] float f;
[0156] int32_ti;
[0157] } A, B, C;
[0158] Af = a;
[0159] Bf = b;
[0160] Ci = Ai + Bi - (127 “23);
[0161] return Cf;
[0162] }
[0163] Furthermore, the combination function Combl (comprising the addition of the mantissas with carry-over to the exponent in the event of overflow) is continuous, monotonic and differentiable as illustrated in [Fig.7c]. More particularly, curve 300 represents the result of the combination function Combl applied to a constant value equal to 1.5 and to the variable x. For comparison, curve 310 represents the result of a conventional slow multiplication of the constant value equal to 1.5 by the variable x. Advantageously, curve 300 of the first combination according to the invention has properties of continuity, monotonicity and differentiability. These properties mean that the combination function Combl can be used in the neural network 100 applied in particular to image denoising or signal denoising.
[0164] In certain exemplary embodiments, each weight associated with each upstream value for the determination of a neuron 110 is specific to the connection between this upstream value and this neuron 110. Such an implementation corresponds for example to the network 100 of [Fig.l].
[0165] In certain exemplary embodiments, each neuron 110 of a layer, for example a neuron 110 of layer L2 in [Fig.l], is determined as a function of several neurons 110 of a preceding layer, for example several neurons 110 of layer L1 in [Fig.l], each carrying an upstream value, by a sum of the combinations, according to the first combination function Combl, of the neurons 110 of the second preceding layer with their weight, corresponding to fast multiplications. The results of the fast multiplications, according to the first combination function Combl, are simply summed together.
[0166] In certain exemplary embodiments, the neurons of a previous layer are each associated with an activation function (for example JfA if we refer to j -i \ / 7 th neuron 110 of the z-th layer of the network 100 of [Fig.l]) beyond a determined activation threshold to calculate the neurons in the next layer.
[0167] In certain exemplary embodiments, each neuron 110 of a layer can be determined as a function of a combined bias, according to a rapid multiplication corresponding to the first combination function Combl, with a weight specific to the link between this bias and the neuron.
[0168] In certain exemplary embodiments, each neuron 110 of a layer (for example layer L3 in [Fig.2]) is also determined as a function of at least one value, carried by a neuron 110 of a downstream layer (for example layer L7 in [Fig.2]), combined, according to the first combination function Combl, with an own weight.
[0169] In certain exemplary embodiments, the network is trained by an algorithm aimed at adjusting the weights, as well as possibly the activation thresholds, from an error calculated at the output, the error at the output of each neuron 110 being propagated backward by a sum of the combinations, using the first function according to the invention, of each error. The weights are thus adjusted iteratively.
[0170] In certain exemplary embodiments, the weight is adjusted iteratively according to a learning rate a combined, according to the first Combl function according to the invention, with the gradient of an error function (for example the error calculated at the output of the network 100), 3) with 6 a vector comprising the parameters to be adjusted. The learning rate, combined according to the invention, advantageously makes it possible to adjust the speed of convergence of the learning.
[0171] Advantageously, the parameters of the network 100 can converge towards a state minimizing the output error.
[0172] The gradient is for example used in a normalized form according to the invention and calculated using a second combination function:
[0173] Comb2 ^gradient, RACiConM (gradient norm, gradient norm}^- (5^5^^(1 + ^-71 / ^^2^°^ (3)
[0174] where:
[0175] - RAC is the known square root function
[0176] - Comb2 is the second combination function;
[0177] - SD is the sign of the gradient value;
[0178] - SF is the sign of RAC (Combl(gradient norm, gradient norm));
[0179] - Md is the mantissa of the gradient value;
[0180] - Mf is the mantissa of RAC (Combl(gradient norm, gradient norm));
[0181] - Ed is the exponent of the gradient value and
[0182] - Ep is the exponent of RAC (Combl(gradient norm, gradient norm)).
[0183] Thus, the second combination function Comb2 represents an alternative to a classic normalization and thus makes it possible to reduce the computing power, the computing time and the energy required. The second combination function Comb2 according to the invention, which corresponds for example to a fast division, in fact implements a simple subtraction of the mantissas.
[0184] In the aforementioned exemplary embodiments, the first combination function Combl according to the invention and respectively the second combination function Comb2 according to the invention are used for specific calculations, when implementing a neural network to perform fast multiplications or respectively fast divisions.
[0185] Alternatively, in certain exemplary embodiments, any multiplication and / or division can be systematically and automatically replaced by a set of specific instructions compiled and executed by a conventional computer or by a specific computer. Thus, the first combination function Combl according to the invention can replace each slow multiplication. The second combination function Comb2 according to the invention can replace each slow division.
[0186] First combination functions Combl according to the invention can thus replace the classic slow multiplications of a classic implementation of a neural network where the multiplications intervened in almost all the operations (inference, gradient calculation, step calculation, gradient normalization). Second combination functions Comb2 according to the invention can replace the slow divisions which intervened for example during normalization operations, even if the divisions represent a small volume of calculations compared to the quantity of multiplications. It is then possible to produce a calculator completely devoid of complex operations and capable of carrying out the training of a neural network. This approach is advantageous because classic slow division is expensive to implement in circuit form.
[0187] In some cases, some complex slow activation functions may still remain but they nevertheless represent a small volume of calculations compared to the quantity of multiplications.
[0188] Alternatively, the first combination function according to the invention can replace a part of the classic slow multiplications, for example 20%, 30%, 50% or even 80%. This results in particular in a saving of energy and calculation time.
[0189] As understood from reading the preceding description, in embodiments, each of the fast division operations performed, for two stored numbers D and F, corresponds to: ( SD X SF ) X ( 1 + Md - Mf) X SD is the sign of D, SF is the sign of F, Md is the mantissa of D, MF is the mantissa of F, ED is the exponent of D, and EF is the exponent of F, mantissa and exponent being defined in a given base.
[0190] In embodiments implementing fast multiplication operations, each of the fast multiplication operations performed, for two stored numbers A and B, corresponds to: ( X SB ) x ( 1 + Ma + Mb ) x two memorized numbers A and B being broken down as follows: A = x (1 + Ma ) x 2EB = SB x ( 1 + Mb ) x 2^C°Û Sa is the sign of A, SB is the sign of B, MA is the mantissa of A, MB is the mantissa of B, EA is the exponent of A, EB is the exponent of B, and C is a bias used in the sign decomposition, mantissa and exponent being defined in a given base.
[0191] Preferably, all multiplication operations included in said inference and gradient descent calculations are fast multiplication operations.
[0192] Preferably, each value carried by one of the neurons 110 and each weighting factor associated with a transmission of a value from one neuron 110 to the other, is coded on the same number of binary data (“bits”), in the same format in the form of a floating point number. For example, this format is a 32-bit floating point number.
[0193] In [Fig.14], we observe a data processing system 60 for a space application comprising an onboard data processing device 32 in a spacecraft 30 and a ground processing device 82 in a ground installation 80. The spacecraft 30 comprises for example a neural network implemented by an onboard computing machine using fast multiplications and fast divisions. The ground installation 80 comprises for example a neural network implemented by a computing machine also using fast multiplications and fast divisions. The ground installation 80 also comprises a means of commu communication 87 with the communication means 37 of the spacecraft 30 and the ground data processing device 82. This ground data processing device 82 comprises a computing machine 88 identical to that on board the satellite but operating as a decompressor, while the onboard computing machine operates as a compressor. In the ground computing machine, a data input 81 provides data to a neural network 83 implemented by the computing machine 88. The neural network 83 provides data at output 84.
[0194] The communication means 37 is configured to transmit data to the communication means 87, in particular compressed training data. The transmitted data being compressed, they would be recovered at the output 84 of the decompressor. Indeed, the decompressor is configured according to the past data. Therefore, the decoder has also received and decoded all the necessary information when it decodes a byte. This principle can be used for compression, it is also valid for video with motion compensation. The encoder on board the satellite is thus synchronized with the decoder on the ground. The operation of the calculation machines 38 and 88 being identical, the data representative of the networks obtained by adjustment implementing the same iterative algorithm and the same number of iterations are identical.In particular, since the fast division and, possibly, fast multiplication operations are deterministic, their results are identical in the spacecraft 30 and in the ground installation 80.
[0195] In embodiments, at least one of the data processing devices 32 or 82 further comprises a performance evaluation module. In [Fig. 14], it is the ground data processing device 82, which comprises a performance evaluation module 89.
[0196] The performance evaluation module 89 (or its equivalent in the spacecraft 30) stores, on the one hand, a configuration of the neural network (including its weights and its activation threshold values) having a determined measured performance at a determined date and compares, on the other hand, the stored performance with a measured performance of another configuration of the network adjusted by a learning phase. The computing machine 88 is configured to return the neural network 83 to the configuration of the network stored by the performance evaluation module 89, if the performance of the configuration of the network adjusted by a learning phase is lower than the performance of the stored network configuration. In order for the neural network 33 to remain identical to the neural network 83, the onboard data processing device 32 comprises a module 39 for storing the same neural network configuration as the module 89.
[0197] Alternatively, the on-board data processing device 32 comprises a performance evaluation module 39 identical to the module 89.
[0198] Alternatively, when the computing machine 88 returns the neural network 83 to the network configuration stored by the performance evaluation module 89, the ground communication means 87 transmits to the on-board communication means 37 the configuration data of the neural network 33. 6. EXAMPLES OF IMPLEMENTATION OF THE INVENTION
[0199] We now present, in relation to [Fig.8a], [Fig.8b] and [Fig.8c] an example of classification performance obtained with a network 100 of neurons 110 implemented according to the invention.
[0200] More particularly, [Fig.8a] shows a photo of the surface of the moon on which we seek to classify areas that can be used for the landing of a spacecraft. [Fig.8b] shows a photo resulting from the processing of the photo of [Fig.8a] by a network 100 (for example, a DNN type network) trained for this purpose according to the invention. The inference is implemented using a classic slow multiplication. The black areas on the photo at the output of the neural network, as shown in [Fig.8b], represent the areas classified as being able to be used for a moon landing.
[0201] In this context, [Fig.8c] presents the efficiency function of the network 100, or ROC curve (from the English "Receiver Operating Characteristic"), also called sensitivity / specificity curve, applied to the classification of characteristic zones on a photo. Such a function is a measure of the classification performance of the network 100 in the form of a curve which gives the true positive rate (fraction of positives which are actually detected), or TPR (from the English "True Positive Rate"), as a function of the false positive rate (fraction of negatives which are incorrectly detected), or FPR (from the English "False Positive Rate").
[0202] For comparison, curve 400 represents the ROC curve obtained by conventional learning using slow operations of the network 100. Curve 410 represents the ROC curve obtained by learning the network 100 according to the invention, by implementing fast operations. It can thus be seen that the implementation using the Combl combination function according to the invention does not significantly degrade the classification performance of the network 100, despite a remarkably reduced computational load.
[0203] Examples of image denoising performance are now presented in relation to [Fig.9a], [Fig.9b], [Fig.9c], [Fig.9d], [Fig.10a], [Fig.10b], [Fig.10c] and [Fig.10d].
[0204] More specifically, [Fig.9a] shows a noisy photo of a sports field, provided as input to a neural network and which we are trying to denoise. [Fig.9d] which illustrates the result of denoising by a neural network trained using a classic slow multiplication, where the inference is also done using a multi classical slow multiplication, is used as a control case. [Fig.9b] illustrates the result of denoising by a neural network trained according to the invention, where the inference is also done according to the invention while [Fig.9c] illustrates the result of denoising by a neural network trained half according to the invention and half using classical slow multiplication, where the inference is done according to the invention, with fast operations. [Fig.9d] illustrates the result of denoising the photo in [Fig.9a] by a neural network trained using classical slow multiplication, where the inference is also done using classical slow multiplication (control case). It is observed that the implementation of the invention does not significantly degrade the denoising performance despite a remarkably reduced computational load. The noisy [Fig.9a] notably presents grainy areas and broken lines.
[0205] If we refer to the details, in relation to [Fig. 10a], [Fig. 10b], [Fig. 10c] and [Fig.10d], the granular area G1 is transformed into a homogeneous surface G2 on the control image. On the image obtained halfway according to the invention, G3 is also homogeneous. On the image obtained entirely according to the invention, the area G4 has been homogenized and has a very low granulometry, making it usable for certain applications. The initial lines H1 can be irregular or even discontinuous, these lines H2 being made more regular and continuous on the control image. In both cases according to the invention H3 and H4 are also made more regular and continuous.
[0206] According to another example of application, [Fig. 11a] represents a noisy photo of a road on which two vehicles are traveling, one, referenced W2, in a shaded area and the other, referenced W1, in a sunny area. [Fig. 11b] illustrates the result of denoising by a neural network using a classic slow multiplication for learning and for inference and is used as a control case. [Fig. 11c] illustrates the result of denoising the photo by a neural network implemented according to the invention, for learning and for inference, while [Fig.11d] illustrates the result of denoising the photo by a neural network trained half according to the invention and half using a classic slow multiplication, the inference being done according to the invention with fast multiplications. It is observed that the denoising results are remarkably close.
[0207] By zooming in with a clockwise rotation of a quarter turn, on the vehicle in the sunshine zone Wl, we see in [Fig. 12a] that initially, a shadow zone Z1 and a light zone Y1 each have a significant granulometry. In the case of the control image of [Fig. 12b] these zones Z2 and Y2 are homogenized. Similarly, in the results of the treatments according to the invention illustrated in [Fig.12c] and in [Fig.12d], these zones Z3, Z4, Y3 and Y4 have been homogenized and make it possible to clearly distinguish a sunny zone or a shaded zone. Thus, the implementation of a neural network tuned according to the invention does not significantly degrade the noise reduction performance despite a remarkably reduced computational load thanks to the implementation of the invention.
[0208] In this context of image denoising, [Fig. 13] shows the evolution of the loss function L obtained as a function of the number of iterations during the training of a neural network 100. Curve 500 represents the control loss function L obtained by the conventional training of the network 100 entirely using a conventional slow multiplication. Curve 510 represents the loss function L obtained by training according to the invention of the network 100 with a fast multiplication. It can be seen that the speed of learning of the network 100 is not significantly degraded, despite a significantly reduced computational load thanks to the implementation of the invention. Curve 520 represents the loss function L obtained by conventional training of the network 100 up to the label SP, then a switch to training according to the invention using in particular the first combination function according to the invention. The performance is then almost not degraded.
[0209] The method according to the invention is not limited to training or inference calculation, in the field of image processing. In general, the invention can be applied to training or inference calculation, in various fields such as:
[0210] - image recognition, the input data corresponding to the set of pixels of an image and output data corresponding to object classes, an indicator of regions of interest and / or object coordinates;
[0211] - data compression applications, the input data corresponding to images, sounds, signals, videos and output data corresponding to compressed or more easily compressible data;
[0212] - RF data processing (detection, tracking, decoding and / or separation of signals), input data corresponding to RF signals and output data corresponding to output signals and / or data associated with RF signals;
[0213] - image processing for a navigation system (space appointment, landers, rovers), the input data corresponding to one or more hnage(s) and the output data corresponding to parameters characterizing the state of a target or the vessel;
[0214] - image processing for robotics applications (satellite maintenance, etc.) orbital assembly), the input data corresponding to one or more image(s) and the output data corresponding to parameters characterizing the state of a target, or command to be sent to an actuator (for example a robot arm);
[0215] - image processing for metrology applications (measurement of geometry of a deployable reflector, measurements for ground tests), the input data cor corresponding to one or more image(s) or image-derived data and the output data corresponding to 3D information about the target;
[0216] - object detection, the input data corresponding to one or more image(s) and output data corresponding to a map showing detected objects (e.g. clouds or cars);
[0217] - recognition of flat areas, the input data corresponding to the set of pixels in an image and the output data corresponding to binary results for each pixel; and
[0218] - denoising, or “denoising” in English, the input data corresponding to the set of pixels of a noisy image and the output data corresponding to the set of pixels of a denoised image.
[0219] In applications of the present invention for an on-board space application, the on-board space application that is executed in the spacecraft is re-executed in ground equipment, to verify on the ground that everything is happening correctly in the spacecraft. This allows a gain in autonomy of the on-board space application, while performing supervision on the ground, for example to ensure that the learning of the neural network executed in the spacecraft does not diverge.
[0220] The saving of material and energy resources favors the implementation of a neural network according to the invention in the space domain, for example on board a spacecraft such as a satellite. The neural networks according to the invention can also be applied to autonomous piloting.
Claims
1. Claims Data processing device (32), on board a spacecraft (30) and comprising a means of communication with a ground device and a computing machine (38, 321, 322, 323), configured to implement a neural network (100, 120) receiving, as input, input data (31) and generating, as output, at least one output data (34), the neural network comprising a plurality of layers (L1 to L7), a neuron (110) of a given layer being determined as a function of at least one upstream value multiplied by a weight, the computing machine comprising a processing unit (321) and a memory space (323) storing data representative of the neural network, characterized in that the computing machine is configured to carry out so-called "slow" multiplication and division operations, consuming a first computing power and a first quantity of electrical energy,and so-called "fast" deterministic multiplication and division operations, consuming a second computing power less than the first computing power and / or a second quantity of electrical energy less than the first quantity of electrical energy, fast operations the result of which has, depending on the numbers on which a fast operation is carried out, a precision less than or equal to the corresponding slow operation, and in that the computing machine is configured to carry out at least one learning phase (42) of the neural network, each learning phase comprising an adjustment of the data representative of the neural network by an iterative algorithm implementing learning data, each iteration (423) of the iterative algorithm comprising inference calculations (421) followed by gradient descent calculations (422), said inference calculations using at least one fast deterministic division operation,the means of communication with the ground device being configured to transmit to the ground device and / or receive from the ground device learning data, the data processing device thus making it possible, on the one hand, to adapt the computing power and the electrical energy necessary for said at least one learning phase on board the spacecraft and, on the other hand, to implement, in the ground device, another, neural network identical to the embedded neural network.
2. The data processing device (32) of claim 1, wherein the computing machine (38) is configured so that the gradient descent calculations use fast deterministic multiplication operations and fast deterministic division operations.
3. Data processing device (32) according to one of claims 1 or 2, which further comprises a performance evaluation module (39) configured to, on the one hand, store a configuration of the neural network (33) having a determined measured performance on a determined date and, on the other hand, measure the performance of the neural network obtained by adjusting its representative data by a learning phase (42), and compare the stored performance with the measured performance, the computing machine being configured to return the neural network to the stored network configuration if the measured performance of the network configuration adjusted by a learning phase is lower than the performance of the stored network configuration.
4. Data processing device (32) according to one of claims 1 to 3, wherein the calculating machine (38) is configured so that each of the fast division operations carried out, for two stored numbers D and F, corresponds to: ( SD x SF) x ( 1 + Md - Mf) x 2Ed~Ef where Sd is the sign of D, SF is the sign of F, MD is the mantissa of D, MF is the mantissa of F, ED is the exponent of D, and EF is the exponent of F, mantissa and exponent being defined in a determined base.
5. Data processing device (32) according to one of claims 1 to 4, wherein the calculating machine (38) is configured so that each of the fast multiplication operations carried out, for two stored numbers A and B, corresponds to: (SA xSB) x(ï + MA+MB)x 2£A+^2Qes the two stored numbers A and B being decomposed as follows: A = SA x ( 1 + Ma ) x 2eB = SB x ( 1 + Mb ) x 2£flTC°û sa is the sign of A, SB is the sign of B, MA is the mantissa of A, MB is the mantissa of B, EA is the exponent of A, EB is the exponent of B, and C is a bias used in the sign decomposition, mantissa and exponent being defined in a determined base.
6. Data processing device (32) according to one of the claims 1 to 5, wherein the computing machine (38) is configured so that all division operations included in said inference and gradient descent calculations are fast division operations, except division operations included in the execution of the Softmax function for a final layer of the neural network.
7. A data processing device (32) according to one of claims 1 to 6, wherein the computing machine (38) is configured so that all multiplication operations included in said inference and gradient descent calculations are fast multiplication operations.
8. Data processing device (32) according to one of claims 1 to 7, wherein said input data (31) are raw data consisting of symbols and wherein said at least one output data (34) is representative of a prediction of an occurrence of the symbols, said at least one output data being supplied to a module (35), implemented on board the spacecraft (30), for entropic compression of said raw data.
9. Data processing device (32) according to one of claims 1 to 8, in which the computing machine (38) is configured to operate according to a first operating mode corresponding to the learning phases and according to a second operating mode corresponding to inference calculations without learning, fast multiplications and fast divisions being used in these two operating modes.
10. Data processing device (32) according to one of claims 1 to 9, in which the calculation machine (38) is configured so that each value carried by one of the neurons and each weighting factor associated with a transmission of a value from one neuron to another, are coded on the same number of binary data, in the same format in the form of a floating point number.
11. Data processing device (32) according to one of claims 1 to 10, wherein the neural network (33) is configured to be adjusted by an initial training (41) before the launch of the spacecraft.
12. Data processing system (60) for a space application comprising a data processing device (32) on board a spacecraft (30) according to one of claims 1 to 11 and a ground processing device (82) in a ground installation (80) comprising a means of communication (87) with the means of communication (37) of the spacecraft (30), the ground data processing device comprising a computing machine (88) identical to the on-board computing machine (38).
13. Data processing system (60) according to claim 12, wherein the ground computing machine (88) further comprises a performance evaluation module (89) configured to, on the one hand, store a configuration of the ground neural network (83) having a determined measured performance on a determined date and, on the other hand, measure the performance of the ground neural network obtained by adjusting its representative data by a learning phase (42), and compare the stored performance with the measured performance, the computing machine being configured to return the ground neural network and the onboard neural network (33) to the stored network configuration if the measured performance of the network configuration adjusted by a learning phase is lower than the performance of the stored network configuration.
14. Method for processing data on board a spacecraft (30) comprising a step of communication with a ground device and a calculation step implementing a neural network receiving, as input, input data (31) and generating, as output, at least one output data (34), the neural network comprising a plurality of layers (L1 to L7), a neuron (11) of a given layer being determined as a function of at least one upstream value multiplied by a weight, the calculation step also implementing a memory space (323) storing data representative of the neural network, characterized in that the calculations comprise so-called “slow” multiplication and division operations, consuming a first computing power and a first quantity of electrical energy, and so-called “fast” deterministic multiplication and division operations,consuming a second computing power lower than the first computing power and / or a second quantity of electrical energy lower than the first quantity of electrical energy, fast operations whose result has, depending on the numbers on which a fast operation is carried out, a precision lower than or equal to the corresponding slow operation, and in that the calculation step comprises at least one learning phase (42) of the neural network, each phase, learning comprising an adjustment of the data representative of the neural network by an iterative algorithm implementing learning data, each iteration (423) of the iterative algorithm comprising inference calculations (421) followed by gradient descent calculations (422), said inference calculations using at least one fast deterministic division operation, the step of communicating with the ground device transmitting to the ground device and / or receiving from the ground device learning data, the data processing method thus making it possible, on the one hand, to adapt the computing power and the electrical energy necessary for said at least one learning phase on board the spacecraft and, on the other hand, to implement, in the ground device, another neural network identical to the on-board neural network.