Computer system and method for real-time evaluation of matrix expressions of a code
The method addresses the challenge of ensuring bounded execution time in real-time systems by compiling source code into object code with pre-allocated temporary memory spaces, optimizing computational performance and compliance with time constraints.
Patent Information
- Application Number
- PCT/EP2024/088137
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Real-time systems, such as robotic control systems, face challenges in ensuring bounded execution time for matrix expressions due to automatic and implicit dynamic memory allocation, which can lead to compliance issues with time constraints.
A method for compiling source code into object code that involves determining matrix operations from syntactic analysis, generating machine instructions, and executing these instructions with pre-allocated temporary memory spaces, avoiding automatic dynamic allocation in the Heap.
This approach ensures optimal computational performance and bounded execution time, reducing memory footprint and implementation errors, while maintaining readability and maintainability of complex matrix expressions.
Smart Images

Figure EP2024088137_26062025_PF_FP_ABST
Abstract
Description
DESCRIPTION Method and computer system for real-time evaluation of matrix expressions of a code Technical field
[0001] The present invention relates generally to computer systems, and in particular to a method and system for real-time evaluation of matrix code expressions.
[0002] Time-constrained systems (called "real-time systems") operate under the control of one or more computer programs, and take into account execution time constraints to be respected. In particular, real-time systems use a computing device to execute an object code, containing instructions to be executed, which governs the operation of the system under time constraints.
[0003] For some real-time systems, such as robotic control systems, the object code to be executed is derived from the compilation of source code containing one or more matrix expressions to be evaluated (i.e., to be calculated). In the case of such real-time systems, compliance with the system's execution time constraints is as important as the conformity of the results of the evaluation of the matrix expressions in the source code.
[0004] Computer devices for executing software code requiring the evaluation of dynamic matrix expressions typically use automatic and implicit instantiation of dynamic matrices to store the results of intermediate computations of these matrix expressions. In particular, the implemented automatic instantiation induces dynamic memory allocation and deallocation operations. Such operations, when performed through the language's default allocator, typically take place in the "Heap".
[0005] However, in systems subject to real-time constraints, such automatic instantiation, inducing dynamic memory allocation via The default allocator of the language, generally cannot guarantee a bounded execution time, and therefore cannot guarantee compliance with the time constraints essential to the proper functioning of such systems.
[0006] To allow the processing of a matrix expression within a bounded execution time, a commonly used method consists of decomposing complex matrix expressions into simple unary or binary subexpressions for arithmetic operators, as well as an explicit pre-instantiation, outside of real-time context, of dynamic matrices intended for storing the intermediate results of the decomposed subexpressions. However, such a decomposition places a heavy burden on the programmer, making the writing of complex matrix expressions tedious, difficult to read and subject to implementation errors. This also requires the definition of numerous intermediate matrices, thus increasing the memory footprint of the computer program at runtime.
[0007] There is thus a need for a system and a method capable of improving the processing of complex matrix expressions using dynamic matrices, in particular in real time. Summary of the invention
[0008] For this purpose, a method is proposed for compiling a source code into an object code, implemented in a compilation tool. The compilation method comprises the steps of: - receive a source code including at least one matrix expression, - determine, for each matrix expression of the source code, N matrix operations, N being an integer greater than or equal to 1, from a syntactic analysis of the matrix expression, the matrix operations being ordered according to an order relative to operational priority rules, - generating an object code from the matrix operations determined from the at least one matrix expression, the object code comprising machine instructions executable by at least one processor, the object code comprising at least one thread of execution associated with one or more matrix expressions among the at least one matrix expression, and - delivering the object code to an object code execution device capable of executing the object code, and comprising a storage unit, the storage unit being capable of comprising at least one temporary memory space, each temporary memory space being capable of comprising temporary memory sub-spaces.
[0009] The object code instructions include, for each thread of execution: - instructions for pre-allocating temporary memory space in the storage unit of the execution device, and - for each matrix expression associated with the execution thread, instructions for successive execution of N groups of instructions according to the order relative to operational priority rules, each group of instructions being associated with one of the N matrix operations and comprising: - an instruction for allocating an intermediate result matrix corresponding to the result of the evaluation of the matrix operation, in a temporary memory subspace of the temporary memory space, and - an instruction for evaluating the matrix operation and storing the corresponding intermediate result matrix, - an instruction to deliver the Nth intermediate result matrix as a result of the execution of the at least one matrix expression associated with the thread of execution.
[0010] The present invention further proposes a computer method comprising a phase of compiling a source code into an object code and a phase of executing the object code, the method being implemented in a computer system comprising a compilation tool and code execution means, the code execution means comprising a storage unit.
[0011] The method comprises, in the compilation phase, the steps of: - receive a source code including at least one matrix expression, - determine, from a syntactic analysis of the at least one matrix expression, N matrix operations, N being an integer greater than or equal to 1, the matrix operations being ordered according to an order relative to operational priority rules, and - generating an object code from the N matrix operations, comprising machine instructions executable by at least one processor, the object code comprising at least one thread of execution associated with one or more matrix expressions among the at least one matrix expression,
[0012] The method comprises, in the object code execution phase, for each thread of execution, the steps of: - pre-allocate temporary memory space in the storage unit of the execution device, and - for each matrix expression associated with the execution thread, successively executing the N matrix operations determined for the matrix expression, according to the order relative to operational priority rules, the execution of one of the matrix operations comprising the steps consisting of: - allocate an intermediate result matrix corresponding to the result of the evaluation of the matrix operation, in a temporary memory subspace of the temporary memory space, and - evaluate the matrix operation and store the intermediate result matrix in the temporary memory subspace, - output the Nth intermediate result matrix as a result of executing the at least one matrix expression associated with the thread of execution.
[0013] In embodiments, for each matrix expression associated with the thread of execution, the temporary memory subspaces, associated with the provided intermediate result matrices, can be allocated according to a dynamic allocation strategy of the "stack" type in the at least one temporary memory space pre-allocated in the storage unit and associated with the thread of execution.
[0014] Advantageously, for each thread of execution, the temporary memory space can be pre-allocated in a "Heap" memory segment, the storage unit corresponding to the default dynamic memory allocator.
[0015] In some embodiments, the method may comprise, in the execution phase, a step of pre-allocating at least one persistent memory space in the storage unit, and for each thread of execution, a step of assigning, for each matrix expression associated with the thread of execution, the Nth intermediate result matrix of the Nth temporary memory subspace to a matrix of the persistent memory space.
[0016] The method may comprise, in the execution phase, for each thread of execution, a step consisting of deallocating, for each matrix expression associated with the thread of execution, the N temporary memory subspaces of the temporary memory space associated with the thread of execution.
[0017] Advantageously, the method may comprise, in the execution phase, for each thread of execution, a step consisting of deallocating the temporary memory space of the storage unit.
[0018] The size of the pre-allocated temporary memory space for each thread can be determined based on the number N of intermediate result matrices provided and the size of the intermediate result matrices associated with a thread's matrix expression.
[0019] A compilation tool is also provided that is configured to implement the compilation process.
[0020] In embodiments, the tool may include a computer library and a compiler, the source code being compiled by the compiler from a set of processing instructions predefined in the computer library.
[0021] Advantageously, the computer library can comprise a definition of at least one class model of different data used during the execution of the source code, a definition of at least one matrix operation, and a definition of at least one matrix assignment operation.
[0022] The at least one class template may include a single class template including a definition of a persistent matrix and a temporary matrix as different instances relative to the single class template, the single class template defining the set of operations associated with the matrices.
[0023] In embodiments, the at least one class template may include a first class template and a second class template, the second class template including a definition of a persistent matrix and of a temporary matrix as different instances relative to the second class template, and the first class template comprising a definition of a parent matrix common to the persistent matrix and the temporary matrix, the first class template defining arithmetic operations associated with the matrices and the second class template defining maintenance operations associated with the matrices.
[0024] Embodiments of the invention thus provide a computer system configured to implement the compilation method.
[0025] Advantageously, the code execution means can be subject to real-time constraints.
[0026] Embodiments of the invention thus provide a system and method capable of improving the processing of matrix expressions using dynamic matrices, by optimizing computational performance and ensuring a bounded execution time.
[0027] In particular, they ensure the absence of automatic and implicit dynamic allocation of dynamic matrices through the language's default allocator (i.e. in the Heap) to store the results of intermediate calculations of matrix expressions, and guarantee a minimal memory footprint and optimal execution performance due to preserved co-locality of data via a specific allocation strategy. This results in a solution compatible with computing devices (or computer systems) subject to time constraints.
[0028] They also ensure the use of writing so-called literal matrix expressions, in a conventional format (i.e. in a format conforming to the mathematical writing of the expression), readable and easy to maintain, whatever their complexity, without requiring the definition of intermediate matrices. The programmer's effort is greatly reduced, the associated computer code gains in readability, maintainability and is less subject to implementation errors. Description of the figures
[0029] Other characteristics, details and advantages of the invention will emerge from reading the description given with reference to the appended drawings given by way of example.
[0030] [Fig. 1] Figure 1 is a diagram showing a system comprising a compilation computing device and an execution computing device, according to embodiments of the invention.
[0031] [Fig.2] Figure 2 shows two diagrams (a) and (b) representing the states of a so-called temporary memory space respectively before and during the evaluation of a matrix expression by an executing computer device, according to embodiments of the invention.
[0032] [Fig.3] Figure 3 is an example of implementation of lines of code of a computing library used by a compilation computing device, according to embodiments of the invention.
[0033] [Fig.4] Figure 4 is an example of implementation of a computing library used by a compilation computing device, according to embodiments of the invention.
[0034] [Fig.5] Figure 5 is a flowchart representing a method of processing matrix expressions, by a compiling computer device, according to embodiments of the invention.
[0035] [Fig.6] Figure 6 is a flowchart representing a method of processing an object code resulting from the compilation of matrix expressions, by an executing computer device, according to embodiments of the invention.
[0036] Like references are used in figures to designate identical or similar elements. For clarity, the elements shown are not to scale. Detailed description
[0037] Figure 1 schematically represents a system 1 comprising a compilation computing device 10 and an execution computing device 20, according to embodiments of the invention.
[0038] The compilation computing device 10 is configured to compile a source code 121 into an object code 221 intended to be executed (i.e., the object code is executable) by the execution computing device 20.
[0039] The execution computing device 20 can be used in numerous applications requiring calculations on dynamic matrices and subject to real-time constraints. The computing device may in particular be a control-command device (or 'process control' according to the English expression). For example and without limitation, such a computing device may be a robotic control-command device, an aircraft navigation control-command device, a control-command device for a complex cyber-physical device, a device for digital simulation of complex devices, or even a signal processing device (in particular image processing).
[0040] As shown in Figure 1, the executing computing device 20 may be different from the compiling computing device 10. Alternatively, the compiling device 10 and the executing device 20 may be one and the same computing device.
[0041] The compilation computer device 10 (also called 'compilation tool' or 'source code compilation device') comprises a processor 160, the source code 121 to be compiled and a computer library 123. The computer device 10 further comprises a compilation computer program, also called compiler 125, intended to be loaded into a RAM 141 of the compilation computer device 10 in order to be executed by the processor 160.
[0042] The source code 121 and the computer library 123 comprise lines of code written in a computer language. The computer language may be a compilable object-oriented computer language, such as, for example and without limitation, the C++ language. The source code 121 comprises one or more matrix expressions E to be processed (i.e., to be evaluated, to be calculated). The computer library 123 comprises predefined processing instructions comprising definitions and declarations of properties in the form of source code. The computer library may also comprise executable machine code. The lines of source code 121 are intended to be compiled into an object code 221 by the compiler 125 from language compilation rules and predefined processing instructions from the 123 computer library.
[0043] In embodiments, the source code 121, the computer library 123 and the compiler 125 may be included in the mass memory (not shown in the figures) of the computing device 10.
[0044] The compiler 125 is used to compile the source code 121 into an object code 221. Advantageously, the compiler 125 can be a standard compiler, using the computer library 123 defining the predefined processing (i.e. compilation) instructions necessary to compile the source code 121.
[0045] As used herein, the term "matrix expression" refers to a literal matrix algebraic expression of any form, composed of a finite combination of symbols and obeying the rules of precedence for arithmetic operators.
[0046] For example and without limitation, a matrix expression, denoted E, can be defined as the following equation (01):
[0047] E: {R = A * X + B * Y} (01)
[0048] In equation (01), the terms R, A, X, B and Y are symbols representing matrices. The matrices A, X, B and Y correspond to the input matrices from which the matrix expression E is evaluated (i.e. calculated, determined). The matrix R is the result matrix resulting from the evaluation of the matrix expression E.
[0049] In equation (01) also, the terms "*" and "+" are operator symbols, representing respectively the arithmetic operators of matrix composition (or multiplication, or product) and addition (or sum).
[0050] It should be noted that a matrix expression E is composed of N elementary matrix operations (also called 'elementary expressions'), denoted O n, associated with an index 'n' which is an integer between 1 and N and corresponds to the nth matrix operation. The value of N is an integer greater than or equal to 1. Each matrix expression E from the source code 121 can be characterized by a number N of ordered matrix operations distinct and / or specific to the matrix expression E. The result of the execution (or evaluation) of a matrix operation O n of the matrix expression E produces an intermediate result matrix, denoted T n .
[0051] Each matrix operation O nof a matrix expression E includes a single operation symbol. A matrix operation can be a unary operator (relating to the treatment of a single matrix operand), a binary operator (relating to the treatment of two operands), or a function of any number of operands. For example, and without limitation, such a function can be a unary function, such as the 'determinant' function det(A) or the trace function tr(B), or a binary function, such as the 'Kronecker product' function kron(C, D), where the terms A, B, C, and D represent matrices.
[0052] A matrix operand of matrix operation O n can be an input matrix of the matrix expression E, or an intermediate result matrix T h associated with another matrix operation of the matrix expression E, denoted O h, whose index 'h' is an integer between 1 and n. The N matrix operations of a matrix expression E can be classified according to an order defined from precedence rules (corresponding to the precedence rules of operations and, in particular, of arithmetic operators, i.e. the rules of order of processing of operation symbols). In certain embodiments, a binary matrix operation can comprise a matrix operand and a scalar operand.
[0053] In the example represented in equation (01), the matrix expression E illustrated comprises three matrix operations Oi, O2 and O3 defined according to the following expressions (02), (03) and (04):
[0054] Oi: { Ti = A * X} (02)
[0055] O2: { T2= B * Y} (03)
[0056] O3: { T3 = Ti + T2} (04)
[0057] The intermediate matrices (resulting from the evaluation of the intermediate matrix operations) are dynamic matrices. As used herein, the term "dynamic matrix" refers to a matrix of dimensions unknown at the compilation of the source code 121. The input matrices of the matrix expression E and / or the result matrix R (resulting from the evaluation of E) may also be dynamic matrices.
[0058] The input matrices of the matrix expression E can be called "persistent matrices". Advantageously, the intermediate result matrices T n associated with matrix operations O nmay be referred to as “temporary arrays.” As used herein, the term “temporary array” refers to an array created in a specific memory space of a computing device during the execution of software code and then deleted during or at the end of execution time, i.e., at the end of the evaluation of the array expression E. Any array that is not a temporary array is a persistent array. Thus, the term “persistent array” refers to an array that remains (i.e., persists) in memory in the computing device beyond the local scope of the software code that created it.
[0059] The executing computing device 20 (also called 'code execution means', 'code execution device') may include a storage unit 240, as well as a processor 260 for executing the object code 221 generated by the compiling computing device 10, as shown in FIG. 1.
[0060] In embodiments, the storage unit 240 of the executing computing device 20 (also referred to as a 'backup unit' or 'storage module') may be the data segment.
[0061] As used herein, the term "data segment" may refer to a segment of RAM dedicated to the data of the object code execution process. The term "data segment" may therefore refer to a memory region that may include: the memory space used by the default dynamic memory allocator of the language, i.e., the "Heap" (or 'stack memory' according to the corresponding English expression), the "Stack" (or 'stack memory' according to the corresponding English expression), and / or the memory space reserved for so-called static data.
[0062] The object code 221 comprises instructions which, when executed by the processor 260, control the executing computing device 20, so that the latter performs the processing of the matrix expressions from the source code 121. In other words, the object code 221 comprises machine instructions executable by at least one processor. In particular, the object code 221 may comprise initialization instructions, processing instructions and termination instructions. The computing device 20 for executing the object code 221 can thus be configured to process some or all of the matrix expressions E included in the source code 121.
[0063] Considering an execution computing device 20 of the robotic system type, by way of non-limiting example, the estimation of the state of such a robotic system can be carried out by Kalman filtering, an algorithm requiring the evaluation of complex matrix expressions. In such a case, a complex matrix expression to be processed by the robotic system can comprise for example up to 7 combined unary and / or binary elementary expressions, for linear Kalman filtering.
[0064] To compile the source code 121 and generate the object code 221, the compilation computing device 10 can be configured to identify (or detect) one or more matrix expressions E in the source code 121 and to determine, for each matrix expression E detected (or found), the N matrix operations O nassociated, from a syntactic analysis of the matrix expression E considered. Advantageously, such a syntactic analysis of the matrix operations can be carried out from the computer library 123 and the rules of the computer language.
[0065] The object code 221 generated by the compilation computer device 10 comprises a set of at least M execution threads, denoted F m , whose index 'm' is an integer between 1 and M corresponding to the m-th thread of execution, the value of M being an integer greater than or equal to 1. Each thread of execution F m is associated with one or more matrix expressions determined from among the set of matrix expressions E detected in the source code 121.
[0066] In the same thread F m, the associated matrix expressions E can be processed for example sequentially. Advantageously, it should be noted that threads of execution F m of an object code 221 can be executed according to a parallel processing according to which several threads of execution F m are executed in parallel by the processor 260 of the executing computing device 20.
[0067] The initialization instructions of the object code 221 include, for each thread F m , the pre-allocation of memory space reserved for matrices temporary 241 (also called 'temporary space' or 'temporary memory space') in the storage unit 240. In other words, such a pre-allocation corresponds to the instruction to pre-allocate a temporary space 241, for each thread of execution F m, in the storage unit 240 of the executing computing device 20 when the initialization instructions are executed by the processor 260 of the system 20. Thus, for an object code 221 comprising M threads of execution associated with matrix expressions to be calculated, the initialization instructions comprise the pre-allocation of M distinct temporary spaces 241 in the storage unit 240.
[0068] The object code processing instructions 221 include, for each matrix expression E associated with a thread of execution F m , the execution of the N matrix operations O n determined for this matrix expression E. In the object code 221, the N matrix operations O n are ordered according to an order relative to the operational priority rules.
[0069] The processing instructions of object code 221 also include, for each execution (or evaluation, calculation) of a matrix operation O n , the allocation of a temporary subspace 241 -n (also called 'temporary memory subspace') included in the allocated temporary space 241, associated with the thread of execution F m . Such a temporary subspace 241 -n constitutes the storage space of the intermediate result matrix T n corresponding to the result of the execution of the matrix operation O n considered.
[0070] In other words, for each matrix expression E associated with a thread F m , the object code 221 may comprise a sequence of grouped instructions for each matrix operation O n . Each group of instructions, denoted G n , thus corresponds to one of the matrix operations O n of an array expression E. Furthermore, in a thread Fm , the G instruction groups n associated with a given matrix expression E are ordered by the compiler in the generated object code according to the operator precedence rules.
[0071] A group of G instructions n , relating to a matrix operation O n , may include: - an instruction to allocate storage space for the result matrix intermediate T n resulting from the evaluation of the matrix operation O n , in a temporary memory subspace 241 -n of the temporary memory space 241 , - an instruction for evaluating the matrix operation O n , and storage (i.e. provision) of the result in the form of the intermediate result matrix T n .
[0072] So, in a thread F m, the evaluation of a matrix expression E can include the successive execution of the N groups of instructions G n (the order of execution of the instruction groups, from Gi to GN, having been determined by the compiler, following the priority of the operations).
[0073] In embodiments, the temporary subspaces 241 -n, relating to the intermediate result matrices T n , can be allocated in the temporary space 241 according to a dynamic allocation strategy of the "stack" type. Such an allocation strategy can be based on a "last in, first out" (LIFO) processing of dynamic matrices. A temporary space 241 therefore forms a data structure whose last element added is the first to leave it during deallocation.
[0074] In embodiments, the object code processing instructions 221 also include, for each matrix expression E associated with a thread of execution F m , the transfer (ie the copy or the movement) of the last intermediate result matrix T N determined at the end of the execution of the N matrix operations O n (i.e. the Nth intermediate result matrix T N stored in the temporary subspace 241 -n associated with the matrix operation O N ), to a result matrix R, temporary or persistent, corresponding to the result of the evaluation (i.e. of the execution) of the matrix expression E considered, associated with the execution thread F m .
[0075] Advantageously, during the execution of the object code 221 by the processor 260, the execution computing device 20 can be configured to perform a pre-allocation operation of a temporary space 241 in the storage unit 240, for each thread of execution F m of the object code 221. The executing computing device 20 may be configured to further perform, for each matrix expression E associated with a thread of execution F m , the successive execution of the N groups of instructions G n . Each group of G instructions n is associated with one of the N matrix operations O n of the matrix expression E. For each group G instructions n (corresponding for example to an iteration), the execution computing device 20 can be configured to determine the intermediate result matrix T n corresponding to the result of the execution of the matrix operation O nconsidered (i.e. the matrix resulting from the calculation of the intermediate matrix operation). The N successive iterations associated with the same matrix expression E can be implemented according to the order of classification of the N matrix operations previously determined.
[0076] Advantageously, for a group of instructions G n , the executing computing device 20 may be configured to allocate, to the intermediate result matrix T n , a temporary subspace 241 -n in the temporary space 241 , and to evaluate (i.e., perform or compute) the matrix operation O n , which provides a result stored in the intermediate result matrix T n The allocation operation corresponds to an operation of “reserving” a storage space (241 -n) for the intermediate result matrix T nThe evaluation operation produces the result of the evaluation and allows saving the intermediate result matrix T n in the allocated temporary subspace 241-n.
[0077] The executing computing device 20 may further be configured to return the Nth intermediate result matrix T N of the temporary subspace 241 -n as a result of the execution of the matrix expression E considered.
[0078] Figure 2 schematically illustrates the states of a temporary space 241 before or after (Figure 2(a)), and during (Figure 2(b)) the evaluation of a matrix expression of a thread of execution F m, by the execution computing device 20. In particular, figure 2(b) illustrates the states of a temporary memory space 241 following the application of successive iterations associated with the example of the matrix expression E represented in equation (01) comprising three matrix operations Oi, O2 and O3 defined according to formulas (02), (03) and (04).
[0079] In embodiments, the size of a pre-allocated temporary space 241 associated with a thread F m , can be determined as a function of the maximum number N of intermediate matrices T n to be evaluated (noted N ma x) for the or each of the matrix expressions E associated with the thread of execution F m considered and the size of the intermediate matrices T n to be evaluated (designated by Size(T n )). In particular, the size of a pre-allocated temporary space 241 may be determined by the product of numbers N maxand Size(T n ). For example, the number Size(T n ) can correspond to the maximum size of the intermediate matrices T n to be evaluated.
[0080] In embodiments, the size of a temporary space 241 may be predefined. For example and without limitation, the size of a temporary space 241 may be equal to a few kilobytes.
[0081] Advantageously, the size of a temporary subspace 241 -n allocated, during the execution of an nth iteration (i.e. group of instructions G n ) associated with a matrix expression E, corresponds to the size of the intermediate result matrix T n for the matrix operation O n considered. The intermediate result matrices T n can be stored in the pre-allocated temporary space 241 associated with a thread F m from a stack top pointer 241 -i, as shown in Figure 2.
[0082] Using pre-allocated temporary space 241 associated with a thread F m , using a dynamic allocation strategy of the "stack" type allows to generate a storage area in the 'Heap' in which the intermediate result matrices T n stored are positioned in memory one after the other, contiguously, thus ensuring optimal co-locality of data storage in the storage unit 140. Such co-locality of data guarantees a minimal memory footprint, allowing the use of a single temporary space 241 per thread of execution F mof the program. This co-locality of the data also makes it possible to ensure optimal execution performance during the evaluation of a matrix expression E. This structuring of the data in the temporary space 241 in fact ensures optimal use of the processor cache (or 'cache-friendly' according to the corresponding English expression), the execution times of the different operations of the execution computing device 20 then being optimal. This optimized execution performance can be illustrated with the matrix expression E represented in equation (01) generating a matrix operation O3 to be determined and defined according to the previous expression (04), using as matrix operands the intermediate matrices T1 and T2, and producing a result in the intermediate matrix T3, all three co-located as represented in figure 2(b).
[0083] Such a temporary space 241, pre-allocated associated with a thread F m, also makes it possible to guarantee a storage area allowing the allocation of matrices of any size, within a limited allocation time of the intermediate result matrices T n , the temporary subspaces 241 -n being positioned at each iteration at the level of the stack top pointer 241 -i.
[0084] In embodiments, the instructions for processing the object code 221 generated by the compilation computing device 10 may comprise, for each matrix expression E associated with a thread of execution F m , the successive deallocation of all the temporary subspaces 241 -n of the temporary space 241 allocated considered, in response to the determination of the result matrix R, obtained at the end of the evaluation of the matrix expression E.
[0085] Thus, during execution of the object code 221 by the processor 260, the execution computing device 20 can be configured to perform, for each matrix expression E associated with a thread of execution F m , after the implementation of the N iterations (i.e. instruction groups Gi to GN), a deallocation operation of the N temporary subspaces 241 -n of the temporary space 241. Such a deallocation operation corresponds to a consequence of the destruction of the set of intermediate result matrices T nevaluated for the matrix expression E. Such destructions take place implicitly and in the reverse order of construction, by principle of operation of the language, the destructions leading to the deallocation of the temporary subspaces 241 -n. The operation of deallocation of the N temporary subspaces 241 -n is therefore carried out in the reverse order of the allocations of the temporary subspaces 241 -n (thus respecting the order of construction of the temporary matrices T n instantiated during the N iterations). Thus, for each matrix expression E processed associated with a thread of execution F m , the temporary space 241 allocated considered is empty (i.e. free) at the end of the processing.
[0086] Advantageously, the use of the temporary space 241 with a dynamic allocation strategy of the “stack” type makes it possible to guarantee a limited deallocation time for each temporary subspace 241 -n, this corresponding to the movement of the stack top pointer 241 -i of the allocated temporary subspaces.
[0087] In embodiments, the termination instructions of the object code 221 generated by the compiling computing device 10 may comprise, for each thread of execution F m , the deallocation of the temporary space 241 in the storage unit 240. In this case, during the execution of the object code 221 by the processor 260, the executing computing device 20 can be configured to perform a deallocation operation of the temporary space(s) 241 in the storage unit 240.
[0088] In some embodiments, the initialization instructions of the object code 221 may comprise a pre-allocation of one or more spaces reserved for persistent matrices 243 (also called 'persistent spaces' or 'persistent memory spaces') in the storage unit 240. In this case, during the execution of the object code 221 by the processor 260, the executing computing device 20 may be configured to perform a pre-allocation operation of one or more persistent memory spaces 243 in the storage unit 240, as shown in FIG. 1. In particular, a pre-allocated persistent memory space 243 may correspond to a given memory segment, in which the allocation of the persistent matrices may be carried out according to a dynamic allocation strategy of the 'pool' type.
[0089] The executing computing device 20 may further be configured to perform an operation of storing some or all of the input matrices of the matrix expression(s) to be processed, in the persistent memory space(s) 243. The dimensions and number of persistent memory spaces 243 to be pre-allocated may for example be predetermined as a function of the sizes and number of input matrices of the matrix expression(s) to be processed. A persistent memory space 243 may also be pre-allocated and adapted to contain the result matrix R corresponding to the result of the evaluation of a matrix expression E.
[0090] The use of such persistent memory spaces 243 can make it possible to instantiate a persistent matrix both in the real-time context of a specific application and outside of such a context, when the allocation strategy guarantees a bounded-time allocation, as in the case of a dynamic allocation strategy of the “pool” type.
[0091] Advantageously, the processing instructions of the object code 221 may comprise, for each matrix expression E associated with a thread of execution F m , the assignment of the Nth intermediate result matrix T Nof the temporary subspace 241 -n to a matrix corresponding to the result matrix R (providing the result of the evaluation of the matrix expression E), the latter being able to be allocated in the storage unit 240. In this case, during the execution of the object code 221 by the processor 260, the executing computing device 20 can be further configured to carry out an assignment operation consisting of assigning the Nth intermediate result matrix T N from temporary space 241 associated with thread F mto the matrix R of this subspace of the storage unit 240. Such a subspace can be included directly in the “Heap”, the “Stack” or the static data area of the storage unit 240, or alternatively in a persistent memory space 243 pre-allocated in this same storage unit. Such a memory subspace can be allocated prior to the implementation of the iterations associated with one or more matrix expressions E of an execution thread F m , For example.
[0092] Such an assignment operation may be included in the processing instructions of the object code 221 during compilation of the source code 121, and determined during the syntactic analysis of the matrix expression E, in response to the detection of the operation symbol "=", as illustrated by equation (01). The symbol "=" represents the operator associated with this matrix assignment operation. Thus, the assignment operation may correspond to a copy from a temporary matrix to a persistent matrix.
[0093] Figure 3 and Figure 4 illustrate possible implementation examples of lines of code written, in the C++ computer language in the form of class template(s), in the computer library 123 used by the computer device 10 to compile the source code 121. These implementation examples, each corresponding to predefined processing instructions, include declarations and definitions of properties, used by the compiler 125 to compile the source code 121 into an object code 221.
[0094] In particular, Figure 3(a) (or part 3(a) of Figure 3) and Figure 4(a) (or part 4(a) of Figure 4) illustrate examples of class template definitions (or 'template class' in English) of the different data used during execution. of the object code 221. The template parameters 'S', 'P' and 'T' represent, respectively, a scalar data type, a class implementing an allocator for persistent matrices, and a class implementing an allocator for temporary matrices. A possible implementation of an allocator for persistent matrices corresponds to the class 'HeapAlloc' (i.e. implementing the allocation and deallocation functions, for example, in the "Heap" of the storage unit 240), and a possible implementation of an allocator for temporary matrices corresponds to the class 'StackAlloc' (i.e. implementing the allocation and deallocation functions in the memory space reserved for temporary matrices 241).
[0095] Figure 3(b) (or part 3(b) of Figure 3) and Figure 4(b) (or part 4(b) of Figure 4) illustrate examples of implementation of a matrix operation using the operation symbol "*" corresponding to the matrix composition (or product, or multiplication) arithmetic operator, taking as parameter two operands of persistent matrix type and / or temporary matrix, and returning a result of temporary matrix type. Any other matrix operation using another arithmetic operator or even a function, in any complex matrix expression E, can be implemented in a manner similar to the example implementation shown in Figure 3(b).Advantageously, the computer library 123 can comprise all the lines of code making it possible to declare and define the properties relating to any other matrix operation associated with an operation symbol, taking as parameters one or more operands of persistent matrix, temporary matrix, or scalar type, and returning a temporary matrix. It should be noted that, in FIG. 3(b), the instantiation of the temporary matrix, resulting from the evaluation of the matrix operation using the operator symbol “*” is ensured via the allocator for temporary matrices according to the use of the template parameter 'T'.
[0096] Figure 3(c) (or part 3(c) of Figure 3) and Figure 4(c) (or part 4(c) of Figure 4) are examples of implementation of an assignment operation ensuring the copying or moving (or more generally the transfer) of the contents of a source matrix of any type (i.e. persistent or temporary) to a destination matrix of any type (i.e. persistent or temporary).
[0097] Thus, in embodiments, as illustrated by the first exemplary implementation of Figure 3, the class template definition may comprise defining a persistent matrix (i.e., MatrixDP in Figure 3) and a temporary matrix (i.e., MatrixDT in Figure 3) as different instances relative to a single common class template (i.e., 'Matrix' in Figure 3), defining the set of operations associated with the matrices (as illustrated by part 3(a) of Figure 3).
[0098] The implementation of a matrix computation operator, as illustrated by part (b) of Figure 3, allows taking as input two matrices of any type (i.e., a persistent matrix and / or a temporary matrix), and generating as output a temporary matrix.
[0099] Furthermore, the implementation of an assignment operator, as illustrated by part (c) of Figure 3, allows taking as input a matrix of any type (i.e. a persistent matrix and / or a temporary matrix), and therefore of a type potentially different from the type of the assigned matrix.
[0100] Thus, the first implementation example, illustrated by Figure 3, allows us to solve the problem of assigning a matrix of any type to a matrix of the same or another type.
[0101] In some embodiments, as illustrated by the second example implementation corresponding to Figure 4, the class template definition may include defining: - a first class model (also called "parent class model"), corresponding to 'MatrixB' in Figure 4, defining the arithmetic operations associated with the matrices, and - a second class model (also called "derived class model"), corresponding to 'MatrixH' in figure 4, defining the so-called maintenance operations, which can include for example the public operations of construction, destruction, or even assignment, responsible for the dynamic allocation and deallocation of memory.
[0102] As used herein, a "public operation" refers to an operation that can be invoked by a user of the computer library 123. Indeed, technically, the first class template (MatrixB) can also define operations of construction, destruction and assignment but, unlike the second class model (MatrixH), these are not public, in the sense that these operations cannot be invoked by the user.
[0103] The second class model (MatrixH) is called "derived" in the sense that it "inherits" from the first class model (MatrixB). As used here, the notion of "inheritance" refers to a process, in an object-oriented language (for example C++), aimed at making a derived class inherit the functionalities of its parent class. This allows in particular to transfer an object by reference (i.e. to "pass by reference") from a derived class to a parent class. It should be noted that the notion of "inheritance" in C++ language uses the operator ":" (possibly followed by the term "public", "protected" or even "private"), as shown in part (a) of figure 4.
[0104] Thus, in embodiments, as illustrated by the second exemplary implementation of Figure 4:
[0105] - the definition of the second class model (MatrixH) may comprise the definition of a persistent matrix (i.e. MatrixDP in Figure 4) and a temporary matrix (i.e. MatrixDT in Figure 4) as different instances relative to the second class model, and
[0106] - the definition of the first class model may include the definition of a parent matrix common to the persistent matrix and the temporary matrix (i.e. MatrixDB in Figure 4) as an instance relative to the first class model.
[0107] The persistent matrix (MatrixDP) and the temporary matrix (MatrixDT) of the second class model are therefore matrices linked by inheritance to the parent matrix (MatrixDB) of the first class model.
[0108] In particular, the first class template (MatrixB) can depend only on template parameters 'S' and 'T', and thus be independent of template parameter 'P' (relative to a class implementing an allocator for persistent matrices).
[0109] Furthermore, the example implementation of an assignment operator shown in part (c) of figure 4 allows taking as input a matrix of any type (i.e. a persistent matrix and / or a temporary matrix), and therefore of a type potentially different from the type of the assigned matrix. Thus, the second implementation example also solves the problem of assigning a matrix of any type to a matrix of the same or another type.
[0110] Furthermore, the second implementation example further allows the separation of maintenance operations (responsible for matrix memory allocation and deallocation via the second class template - MatrixH) from arithmetic operations (via the first class template - MatrixB), by defining two separate class templates linked by "inheritance". Such a separation therefore allows the performance of arithmetic operations on matrices passed by reference to this "parent" type, regardless of their original type (i.e. persistent or temporary), as represented by example (b) in Figure 4.
[0111] Thus, the use of the computer library 123 based on two distinct class models linked by “inheritance” makes it possible to resolve a problem of “passing by reference” to a user function using the computer library 123.
[0112] For example and without limitation, a "pass by reference" problem may correspond to a user function (in an application in the computing device 10) taking as a parameter a persistent matrix (MatrixDP) passed by constant reference, as illustrated by the following line of code
[0001] :
[0113] void user_function(const MatrixDP& m) { ...}
[0001]
[0114] In embodiments using two separate class models linked by "inheritance", as represented by the second exemplary implementation in Figure 4, such a user function may be defined differently by the user, for example by using a parent matrix common to both the persistent matrix and the temporary matrix (MatrixDB), as illustrated by the following line of code
[0002] :
[0115] void user_function(const MatrixDB& m) {...}
[0002]
[0116] The user function, defined according to code line
[0002] , thus uses a pass by reference and can take as parameter a persistent matrix or a temporary matrix indistinctly. The embodiments using two distinct class models linked by "inheritance" thus make it possible to pass to the function an object of type MatrixDP or MatrixDT, without risk of type conversion. Indeed, the compiler can in this case be able to promote (according to a "promotion type upcast) the object from its original type (MatrixDP or MatrixDT) to the MatrixDB type expected by the function.
[0117] Such embodiments using two distinct class models linked by "inheritance" make it possible in particular to avoid an implicit call to a conversion operator, ensuring the conversion of a temporary matrix passed into a persistent matrix. It should be noted that the use of such a conversion operator can induce an allocation of memory in the "Heap", in particular during a conversion of a temporary matrix into a persistent matrix.
[0118] Thus, the implementation of a matrix calculation operator, as detailed in example (b) of Figure 4, allows a persistent matrix to be processed (MatrixH) to be interpreted as a basic matrix (MatrixB).
[0119] In certain embodiments, the computer library 123 can be adapted so that, during the syntactic analysis of a matrix expression E of the source code 121, the compilation computer device 10 is further configured to identify typical mathematical expressions, so as to generate an optimized matrix expression, prior to the determination of the N matrix operations O n . For example and without limitation, a mathematical expression of the type to be identified can be defined according to equation (05):
[0120] F += u * C * D (05)
[0121] In equation (05), the terms C, D, F are symbols representing matrices, the term u is a symbol representing a scalar, and the term “+=” represents the addition-assignment operator.
[0122] Figures 5 and 6 represent the method for the real-time evaluation (or calculation) of matrix code expressions, (or method for processing matrix expressions), implemented by the system 1, according to embodiments of the invention.
[0123] The method comprises a phase P1 of compiling the source code 121 implemented by the compilation computer device 10 (i.e. the compilation tool).
[0124] Advantageously, the compilation phase P1 (or compilation method) can comprise a step consisting of receiving the source code 121 comprising one or more matrix expressions E, and identification of these matrix expressions by analysis of the source code 121.
[0125] The compilation phase P1 further comprises, for each matrix expression E identified in the source code 121, a step 420 consisting of applying a syntactic analysis to determine N elementary matrix operations O n , which can be classified according to the rules of operational priority.
[0126] The compilation phase P1 further comprises a step 440 consisting of generating the executable object code 221, by the execution computer device 20 (i.e. the code execution means), from the elementary matrix operations O n determined and processing instructions, predefined in particular in the computer library 123.
[0127] Advantageously, the compilation phase P1 may comprise a step consisting of delivering the object code 221 to an object code execution device capable of executing the object code 221, and comprising a storage unit 240 capable of comprising at least one temporary memory space 241, each temporary memory space 241 being capable of comprising temporary memory sub-spaces 241 -n.
[0128] The method may further comprise a phase P2 of execution of the object code 221 implemented by the execution computing device 20. The execution phase P2 comprises, for each thread of execution F m relating to the object code 221, a step 406 of pre-allocation of a temporary memory space 241 in the storage unit 140, implemented so that the allocations which will take place in this temporary space are implemented according to a dynamic allocation strategy of the “stack” type.
[0129] The execution phase P2 may further include, for each elementary matrix operation O n , determined for a matrix expression E associated with the thread of execution F m , a step 542 consisting of allocating a temporary subspace 241 -n, in the associated temporary memory space 241, to the intermediate result matrix T n , resulting from the execution (i.e. evaluation or calculation) of the matrix operation O n , and a step 544 consisting of evaluating the matrix operation O n and thus determine the intermediate result matrix T n .
[0130] The last (i.e. N-th) intermediate result matrix T N , obtained at the end of the N matrix operations O n, is transferred (a transfer may correspond for example to a copy or a move) into the result matrix R providing the result of the evaluation of the matrix expression E. The result matrix R can be used by the executing computer device 20 (which may be for example and without limitations a robotic system).
[0131] The execution phase P2 may further comprise a step 562 consisting of assigning the Nth intermediate result matrix T N from the temporary subspace 241 -N to a result matrix R, such a result matrix R being allocated in the storage unit 240 (for example, in the memory segment “Heap” of the storage unit 240, or a persistent space 243 of type “pool”).
[0132] The execution phase P2 can also include, for each matrix expression E associated with the execution thread F m, a step 564 consisting of deallocating all of the temporary subspaces 241 -n allocated, associated with the evaluation of the N elementary matrix operations O n determined for the matrix expression E.
[0133] The execution phase P2 can further include, for each thread F m , a step 580 consisting of deallocating the associated temporary memory space 241 after the evaluation of the matrix expression(s) E associated with the execution thread F m .
[0134] A person skilled in the art will easily understand that certain steps of the method of FIG. 5 can be carried out simultaneously, sequentially, independently or not, and / or in a different order, for example in an order defined by the executing computer device.
[0135] It should be noted that certain features of the invention may have advantages when considered separately.
[0136] Those skilled in the art will understand that the invention may be implemented as a computer program comprising instructions for its execution. The computer program may be recorded on a recording medium readable by a processor. Reference to a computer program which, when executed, performs any of the functions described above, is not limited to an application program running on a single host computer. Rather, the terms computer program and software are used herein in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be used to program one or more processors to implement aspects of the techniques described herein. The computing means or resources may notably be distributed ("Cloud computing"), possibly using peer-to-peer technologies. The software code may be executed on any suitable processor (e.g., a microprocessor) or processor core or a set of processors, whether provided in a single computing device or distributed among multiple computing devices (e.g., as may be accessible within the device environment).The executable code of each program allowing the programmable device to implement the processes according to the invention may be stored, for example, in the hard disk or in read-only memory. Generally speaking, the program(s) may be loaded into one of the storage means of the device before being executed. A central unit may control and direct the execution of the instructions or portions of software code of the program(s) according to the invention, instructions which are stored in the hard disk or in the read-only memory or in the other aforementioned storage elements.
[0137] The invention is not limited to the embodiments described above as a non-limiting example. It encompasses all variant embodiments that may be envisaged by those skilled in the art. In particular, those skilled in the art will understand that the invention is not limited to the various computer units of the device described as a non-limiting example. In particular, certain embodiments of the invention may be combined.
Claims
Claims 1. Method for compiling (P1) a source code (121) into an object code (221), implemented in a compilation tool, the compilation method comprising the steps of: - receive a source code (121) comprising at least one matrix expression (E), - determine, for each matrix expression (E) of said source code (121), N matrix operations (O n ), N being an integer greater than or equal to 1, from a syntactic analysis of said matrix expression (E), said matrix operations (O n ) being ordered according to an order relative to operational priority rules, - compiling said source code (121) and generating an object code (221) from said matrix operations (O n) determined from said at least one matrix expression (E), said object code (221) comprising machine instructions executable by at least one processor, said object code (221) comprising at least one thread of execution (F m ) associated with one or more matrix expressions (E) among said at least one matrix expression, and - delivering said object code (221) to an object code execution device capable of executing the object code (221), and comprising a storage unit (240), the storage unit (240) being capable of comprising at least one temporary memory space (241), each temporary memory space (241) being capable of comprising temporary memory sub-spaces (241-n); the instructions of the object code comprising, for each thread of execution (F m ): - pre-allocation instructions (520) of a temporary memory space (241) in said storage unit (240) of the execution device, and - for each matrix expression (E) associated with said thread of execution (F m ), instructions for successive execution of N groups of instructions (G n ) according to said order, each group of instructions (G n ) being associated with one of said N matrix operations (O n ) and including: o an instruction for allocating an intermediate result matrix (T n ) corresponding to the result of the evaluation of said matrix operation (O n ), in a temporary memory subspace (241 -n) of said temporary memory space (241 ), and o an instruction for evaluating said matrix operation (O n ) and storage of said intermediate result matrix (T n) corresponding, said temporary memory subspaces (241 -n), associated with said intermediate result matrices (Tn) provided, being allocated according to a dynamic allocation strategy of the “stack” type in said pre-allocated temporary memory space (241), - an instruction to deliver the Nth intermediate result matrix (T N ) as a result of the execution of said at least one matrix expression (E) associated with the thread of execution.
2. Computer method comprising a compilation phase (P1) of a source code (121) into an object code (221) and an execution phase (P2) of said object code (221), the method being implemented in a computer system comprising a compilation tool and code execution means, the code execution means comprising a storage unit (240), the method comprising, in said compilation phase (P1), the steps consisting of: - receive a source code (121) comprising at least one matrix expression (E), - determine, from a syntactic analysis of said at least one matrix expression (E), N matrix operations (O n ), N being an integer greater than or equal to 1, said matrix operations (O n ) being ordered according to an order relative to operational priority rules, and - compiling said source code (121) and generating an object code (221) from said N matrix operations (O n ), comprising machine instructions executable by at least one processor, said object code (221) comprising at least one thread of execution (F m ) associated with one or more matrix expressions (E) among said at least one matrix expression, the method comprising, in said execution phase (P2) of said object code (221), for each thread of execution (F m ), the steps consisting of: - pre-allocate (520) a temporary memory space (241) in said storage unit (240) of the execution device, and - for each matrix expression (E) associated with said thread of execution (F m ), successively execute the N matrix operations (O n ) determined for the matrix expression (E), according to said order, the execution of one of said matrix operations (O n ) comprising the steps of: o allocating an intermediate result matrix (T n ) corresponding to the result of the evaluation of said matrix operation (O n ), in a temporary memory subspace (241 -n) of said temporary memory space (241 ), and o evaluate said matrix operation (O n ), and store said intermediate result matrix (T n) in said temporary memory subspace (241 -n), said temporary memory subspaces (241 -n), associated with said intermediate result matrices (Tn) provided, being allocated according to a dynamic allocation strategy of the “stack” type in said pre-allocated temporary memory space (241), - deliver the Nth intermediate result matrix (T N ) as a result of the execution of said at least one matrix expression (E) associated with the thread of execution.
3. Method, according to claim 2, in which, for each thread of execution (F m ), said temporary memory space (241) is pre-allocated (520) in a “Heap” memory segment, said storage unit (240) corresponding to the default dynamic memory allocator.
4. Method according to one of claims 2 to 3, in which the method comprises, in said execution phase (P2), a step consisting of pre-allocating at least one persistent memory space (243) in said storage unit (240), and for each thread of execution (F m ), a step of assigning (562), for each matrix expression (E) associated with said thread of execution (F m ), said Nth intermediate result matrix (T N ) of the Nth temporary memory subspace (241 -N) to a matrix of said persistent memory space (243).
5. Method according to one of claims 2 to 4, in which the method comprises, in said execution phase (P2), for each thread of execution (F m ), a step of deallocating (564), for each matrix expression (E) associated with said thread of execution (F m), said N temporary memory subspaces (241 -n) of said temporary memory space (241) associated with said thread of execution (F m ).
6. Method according to one of claims 2 to 5, in which the method comprises, in said execution phase (P2), for each thread of execution (F m ), a step of deallocating (580) said temporary memory space (241) from said storage unit (240).
7. Method according to one of claims 2 to 6, in which the size of said temporary memory space (241) pre-allocated for each thread of execution (F m ) is determined as a function of the number N of intermediate result matrices (T n ) provided and the size of said intermediate result matrices (T n ) associated with a matrix expression (E) of said thread of execution (F m ).
8. Compilation tool configured to implement the compilation method (P1) according to claim 1.
9. Compilation tool, according to claim 8, wherein the tool comprises a computer library (123) and a compiler (125), said source code (121) being compiled by the compiler (125) from a set of processing instructions predefined in said computer library (123).
10. Compilation tool, according to one of claims 8 or 9, in which said computer library (123) comprises a definition of at least one class model of different data used during the execution of said source code (121), a definition of at least one matrix operation, and a definition of at least one matrix assignment operation.
11. A compilation tool, according to claim 10, wherein said at least one class template comprises a single class template comprising a definition of a persistent matrix and a temporary matrix as different instances relative to said single class template, said single class template defining the set of operations associated with the matrices.
12. A compilation tool, according to claim 10, wherein said at least one class template comprises a first class template and a second class template, said second class template comprising a definition of a persistent matrix and a temporary matrix as different instances relative to said second class template, and said first class template comprising a definition of a parent matrix common to said persistent matrix and said temporary matrix, said first class template defining arithmetic operations associated with the matrices and said second class template defining maintenance operations associated with the matrices.
13. Computer system configured to implement the compilation method (P1) according to one of claims 2 to 7.
14. System according to claim 13, in which the code execution means are subject to real-time constraints.
Citation Information
Patent Citations
Method of compilation, computer program and computing system
FR2986343A1